addPartOfSpeechDetails

Add part-of-speech tags to documents

collapse all in page

Syntax

updatedDocuments = addPartOfSpeechDetails(documents)

updatedDocuments = addPartOfSpeechDetails(documents,Name,Value)

Description

Use addPartOfSpeechDetails to add part-of-speech tags to documents.

The function supports English, Japanese, German, and Korean text.

example

updatedDocuments = addPartOfSpeechDetails(documents) detects parts of speech in documents and updates the token details. The function, by default, retokenizes the text for part-of-speech tagging. For example, the function splits the word "you're" into the tokens "you" and "'re". To get the part-of-speech details from updatedDocuments, use tokenDetails.

updatedDocuments = addPartOfSpeechDetails(documents,Name,Value) specifies additional options using one or more name-value pair arguments.

Tip

Use addPartOfSpeechDetails before using the lower, upper, erasePunctuation, normalizeWords, removeWords, and removeStopWords functions as addPartOfSpeechDetails uses information that is removed by these functions.

Examples

collapse all

Add Part-of-Speech Details to Documents

Open Live Script

Load the example data. The file sonnetsPreprocessed.txt contains preprocessed versions of Shakespeare's sonnets. The file contains one sonnet per line, with words separated by a space. Extract the text from sonnetsPreprocessed.txt, split the text into documents at newline characters, and then tokenize the documents.

filename = "sonnetsPreprocessed.txt";
str = extractFileText(filename);
textData = split(str,newline);
documents = tokenizedDocument(textData);

View the token details of the first few tokens.

tdetails = tokenDetails(documents);
head(tdetails)

ans=8×5 table
       Token       DocumentNumber    LineNumber     Type      Language
    ___________    ______________    __________    _______    ________

    "fairest"            1               1         letters       en   
    "creatures"          1               1         letters       en   
    "desire"             1               1         letters       en   
    "increase"           1               1         letters       en   
    "thereby"            1               1         letters       en   
    "beautys"            1               1         letters       en   
    "rose"               1               1         letters       en   
    "might"              1               1         letters       en

Add part-of-speech details to the documents using the addPartOfSpeechDetails function. This function first adds sentence information to the documents, and then adds the part-of-speech tags to the table returned by tokenDetails. View the updated token details of the first few tokens.

documents = addPartOfSpeechDetails(documents);
tdetails = tokenDetails(documents);
head(tdetails)

ans=8×7 table
       Token       DocumentNumber    SentenceNumber    LineNumber     Type      Language     PartOfSpeech 
    ___________    ______________    ______________    __________    _______    ________    ______________

    "fairest"            1                 1               1         letters       en       adjective     
    "creatures"          1                 1               1         letters       en       noun          
    "desire"             1                 1               1         letters       en       verb          
    "increase"           1                 1               1         letters       en       noun          
    "thereby"            1                 1               1         letters       en       adverb        
    "beautys"            1                 1               1         letters       en       verb          
    "rose"               1                 1               1         letters       en       noun          
    "might"              1                 1               1         letters       en       auxiliary-verb

Get Part of Speech Details of Japanese Text

Open Live Script

Tokenize Japanese text using tokenizedDocument.

str = [
    "恋に悩み、苦しむ。"
    "恋の悩みで 苦しむ。"
    "空に星が輝き、瞬いている。"
    "空の星が輝きを増している。"
    "駅までは遠くて、歩けない。"
    "遠くの駅まで歩けない。"
    "すもももももももものうち。"];
documents = tokenizedDocument(str);

For Japanese text, you can get the part-of-speech details using tokenDetails. For English text, you must first use addPartOfSpeechDetails.

tdetails = tokenDetails(documents);
head(tdetails)

ans=8×8 table
     Token     DocumentNumber    LineNumber       Type        Language    PartOfSpeech     Lemma       Entity  
    _______    ______________    __________    ___________    ________    ____________    _______    __________

    "恋"             1               1         letters           ja       noun            "恋"       non-entity
    "に"             1               1         letters           ja       adposition      "に"       non-entity
    "悩み"           1               1         letters           ja       verb            "悩む"      non-entity
    "、"             1               1         punctuation       ja       punctuation     "、"       non-entity
    "苦しむ"          1               1         letters           ja       verb            "苦しむ"    non-entity
    "。"             1               1         punctuation       ja       punctuation     "。"       non-entity
    "恋"             2               1         letters           ja       noun            "恋"       non-entity
    "の"             2               1         letters           ja       adposition      "の"       non-entity

Get Part of Speech Details of German Text

Open Live Script

Tokenize German text using tokenizedDocument.

str = [
    "Guten Morgen. Wie geht es dir?"
    "Heute wird ein guter Tag."];
documents = tokenizedDocument(str)

documents = 
  2x1 tokenizedDocument:

    8 tokens: Guten Morgen . Wie geht es dir ?
    6 tokens: Heute wird ein guter Tag .

To get the part of speech details for German text, first use addPartOfSpeechDetails.

documents = addPartOfSpeechDetails(documents);

To view the part of speech details, use the tokenDetails function.

tdetails = tokenDetails(documents);
head(tdetails)

ans=8×7 table
     Token      DocumentNumber    SentenceNumber    LineNumber       Type        Language    PartOfSpeech
    ________    ______________    ______________    __________    ___________    ________    ____________

    "Guten"           1                 1               1         letters           de       adjective   
    "Morgen"          1                 1               1         letters           de       noun        
    "."               1                 1               1         punctuation       de       punctuation 
    "Wie"             1                 2               1         letters           de       adverb      
    "geht"            1                 2               1         letters           de       verb        
    "es"              1                 2               1         letters           de       pronoun     
    "dir"             1                 2               1         letters           de       pronoun     
    "?"               1                 2               1         punctuation       de       punctuation

Input Arguments

collapse all

`documents` — Input documents
`tokenizedDocument` array

Input documents, specified as a tokenizedDocument array.

Name-Value Pair Arguments

Specify optional comma-separated pairs of Name,Value arguments. Name is the argument name and Value is the corresponding value. Name must appear inside quotes. You can specify several name and value pair arguments in any order as Name1,Value1,...,NameN,ValueN.

Example: 'DiscardKnownValues',true specifies to discard previously computed details and recompute them.

`'RetokenizeMethod'` — Method to retokenize documents
`'part-of-speech'` (default) | `'none'`

Method to retokenize documents, specified as one of the following:

'part-of-speech' – Transform the tokens for part-of-speech tagging. The function performs these tasks:
- Split compound words. For example, split the compound word "wanna" into the tokens "want" and "to". This includes compound words containing apostrophes. For example, the function splits the word "don't" into the tokens "do" and "n't".
- Merge periods with preceding abbreviations. For example, merge the tokens "Mr" and "." into the token "Mr.".
- Merge runs of periods into ellipses. For example, merge three instances of "." into the single token "...".
'none' – Do not retokenize the documents.

`'DiscardKnownValues'` — Option to discard previously computed details
`false` (default) | `true`

Option to discard previously computed details and recompute them, specified as true or false.

Data Types: logical

Output Arguments

collapse all

`updatedDocuments` — Updated documents
`tokenizedDocument` array

Updated documents, returned as a tokenizedDocument array. To get the token details from updatedDocuments, use tokenDetails.

Algorithms

If the input documents do not contain sentence details, then the function first runs addSentenceDetails.

Documentation

addPartOfSpeechDetails

Syntax

Description

Tip

Examples

Add Part-of-Speech Details to Documents

Get Part of Speech Details of Japanese Text

Get Part of Speech Details of German Text

Input Arguments

`documents` — Input documents
`tokenizedDocument` array

Name-Value Pair Arguments

`'RetokenizeMethod'` — Method to retokenize documents
`'part-of-speech'` (default) | `'none'`

`'DiscardKnownValues'` — Option to discard previously computed details
`false` (default) | `true`

Output Arguments

`updatedDocuments` — Updated documents
`tokenizedDocument` array

Algorithms

See Also

Topics

Introduced in R2018b

Text Analytics Toolbox Documentation

Support

Documentation

addPartOfSpeechDetails

Syntax

Description

Tip

Examples

Add Part-of-Speech Details to Documents

Get Part of Speech Details of Japanese Text

Get Part of Speech Details of German Text

Input Arguments

documents — Input documents tokenizedDocument array

Name-Value Pair Arguments

'RetokenizeMethod' — Method to retokenize documents 'part-of-speech' (default) | 'none'

'DiscardKnownValues' — Option to discard previously computed details false (default) | true

Output Arguments

updatedDocuments — Updated documents tokenizedDocument array

Algorithms

See Also

Topics

Introduced in R2018b

Text Analytics Toolbox Documentation

Support

`documents` — Input documents
`tokenizedDocument` array

`'RetokenizeMethod'` — Method to retokenize documents
`'part-of-speech'` (default) | `'none'`

`'DiscardKnownValues'` — Option to discard previously computed details
`false` (default) | `true`

`updatedDocuments` — Updated documents
`tokenizedDocument` array