Tokenises text via the same tokeniser as term_frequency() (whitespace
splitting, lower-casing, punctuation stripping, and stop-word removal),
then slides a window of n tokens across each response's token vector
and counts how often each resulting n-gram occurs. n = 2 (the default)
gives bigrams, and n = 3 gives trigrams.
Arguments
- text
Character vector of responses (raw or already cleaned by
clean_text_responses()).- n
Integer. N-gram size. Default
2(bigrams).- stop_words
Character vector of words to exclude, or
NULLto use the built-in English list, orcharacter(0)for no filtering.- top_n
Integer. Maximum number of n-grams to return, most frequent first. Default
30.
Details
An n-gram never straddles a removed stop word. Where "but" and "not" are filtered, "clean but not comfortable" yields no bigram at all, because "clean" and "comfortable" were never next to each other. This is what makes the output phrases respondents actually wrote.
Examples
demo <- sframe_demo_data()
cleaned <- clean_text_responses(demo$responses, "comments")
head(ngram_frequency(cleaned, n = 2, top_n = 10))
#> term n pct
#> 1 sustainability details 35 24.1
#> 2 clear online 31 21.4
#> 3 online content 31 21.4
#> 4 service information 24 16.6
#> 5 useful service 24 16.6