Takes the result of sframe_run_stm_topics() (not a raw stm model
object) and, for each topic, pulls the top n_quotes documents by
topic-document probability using stm::findThoughts(), then maps each
one back to its original respondent row via the mapping
sframe_run_stm_topics() stored on result$fit.
Arguments
- model
A
stm_topicsresult list fromsframe_run_stm_topics()(i.e.result, notresult$fitand not the rawstmobject).- text
The response vector the topic model was fit on: either the raw vector (one entry per original data row, indexed 1:1 by row number) or the
clean_text_responses()-cleaned vector, which drops blank/missing rows and is therefore shorter, its positions no longer equal original row numbers once any earlier row was dropped. Both forms work correctly: whentextcarries therespondentattributeclean_text_responses()sets, that mapping is used to find each quote's real position; otherwisetextis assumed to be the raw, 1:1-indexed vector.- n_quotes
Integer. Quotes to return per topic. Default
3L.
Value
A data.frame with columns topic, rank, respondent (the
original row index in the data the model's text argument came from,
not a document-matrix or corpus row index), and quote.
Examples
# \donttest{
if (requireNamespace("stm", quietly = TRUE) &&
requireNamespace("tidytext", quietly = TRUE)) {
demo <- sframe_demo_data()
sf_plan(demo$instrument) <- list(list(
id = "RQ1", research_question = "What themes appear in the comments?",
family = "text_analysis", method = "stm_topics",
roles = list(item = "comments"), options = list(k = 3, seed = 42)
))
res <- run_analysis_plan(demo$responses, demo$instrument)
quotes <- extract_quotes(res$RQ1, demo$responses$comments, n_quotes = 2)
quotes
}
#> topic rank respondent quote
#> 1 1 1 11 Clear online content.
#> 2 1 2 12 Clear online content.
#> 3 2 1 5 Would like more sustainability details.
#> 4 2 2 8 Would like more sustainability details.
#> 5 3 1 2 Useful service information.
#> 6 3 2 3 Useful service information.
# }