Excluding Tokens¶
Sometimes we do not want to include all tokens in our comparison. For example, we may want to remove exceedingly common tokens ("stop words") or exceedingly rare ones if we think that they may not contribute meaningfully to the difference between our corpora.
Note
If you exclude tokens from a comparison, WordLevel automatically renormalizes all normalized frequencies, ensuring that all scores are correctly calculated over the narrower vocabulary.
Example data¶
We will use an example dataset of US presidential speeches from Franklin D. Roosevelt and Joe Biden. Each document is a speech, and each corpus contains all of the speeches by that particular president. See the cookbook for more about working with catalogs.
import wordlevel as wl
speeches = wl.Dataset("presidential_speeches")
cl = wl.Catalog(speeches, corpora=["Joe Biden", "Franklin D. Roosevelt"])
We will start by comparing them with the rank-turbulence divergence. We can look at the most distinguishing tokens (the ones with the highest scores) before we exclude any. Observe how most of the top tokens are function words and contractions. While these are genuine differences, we may want to exclude them because they are not particularly insightful.
cl = cl.with_comparisons(
wl.comp("Franklin D. Roosevelt", "Joe Biden")
.score.rank_divergence(alpha=1/3)
.alias("roosevelt_biden_rank_divergence"),
)
print(
cl.comparisons["roosevelt_biden_rank_divergence"].df
.select("token", "score_rank", "Franklin D. Roosevelt", "Joe Biden")
.sort("score_rank")
.head(n=10)
)
┌────────────┬────────────┬───────────────────────┬───────────┐
│ token ┆ score_rank ┆ Franklin D. Roosevelt ┆ Joe Biden │
│ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ u32 ┆ u16 ┆ u16 │
╞════════════╪════════════╪═══════════════════════╪═══════════╡
│ thats ┆ 1 ┆ 1 ┆ 277 │
│ im ┆ 2 ┆ 0 ┆ 206 │
│ of ┆ 3 ┆ 6440 ┆ 2791 │
│ to ┆ 4 ┆ 4161 ┆ 3842 │
│ which ┆ 5 ┆ 672 ┆ 78 │
│ applause ┆ 6 ┆ 1 ┆ 143 │
│ weve ┆ 7 ┆ 2 ┆ 147 │
│ ukraine ┆ 8 ┆ 0 ┆ 129 │
│ you ┆ 9 ┆ 380 ┆ 1004 │
│ government ┆ 10 ┆ 432 ┆ 45 │
└────────────┴────────────┴───────────────────────┴───────────┘
Excluding specific tokens¶
If we know which specific tokens we want to drop—like a pre-defined list of stop words—we can exclude them from the comparison.
stop_words = ["thats", "im", "of", "to", "which", "weve", "you"]
cl = cl.with_comparisons(
wl.comp("Franklin D. Roosevelt", "Joe Biden")
.exclude.tokens(stop_words)
.score.rank_divergence(alpha=1/3)
.alias("without_stop_words"),
)
Because we excluded several of the previous top tokens, we can see that the new top tokens change, even though the two comparisons are otherwise identical.
print(
cl.comparisons["without_stop_words"].df
.select("token", "score_rank", "Franklin D. Roosevelt", "Joe Biden")
.sort("score_rank")
.head(n=10)
)
┌─────────────┬────────────┬───────────────────────┬───────────┐
│ token ┆ score_rank ┆ Franklin D. Roosevelt ┆ Joe Biden │
│ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ u32 ┆ u16 ┆ u16 │
╞═════════════╪════════════╪═══════════════════════╪═══════════╡
│ applause ┆ 1 ┆ 1 ┆ 143 │
│ government ┆ 2 ┆ 432 ┆ 45 │
│ ukraine ┆ 3 ┆ 0 ┆ 129 │
│ president ┆ 4 ┆ 73 ┆ 438 │
│ men ┆ 5 ┆ 258 ┆ 13 │
│ q ┆ 6 ┆ 0 ┆ 120 │
│ ive ┆ 7 ┆ 1 ┆ 126 │
│ lets ┆ 8 ┆ 0 ┆ 115 │
│ afghanistan ┆ 9 ┆ 0 ┆ 108 │
│ dont ┆ 10 ┆ 2 ┆ 115 │
└─────────────┴────────────┴───────────────────────┴───────────┘
Excluding by frequency¶
We can also exclude tokens by how often they appear. The by_frequency method takes exactly one of three bounds: less_than, greater_than, or between. Below, we exclude all tokens that appear 10 or fewer times.
cl = cl.with_comparisons(
wl.comp("Franklin D. Roosevelt", "Joe Biden")
.exclude.by_frequency(less_than=10)
.score.rank_divergence(alpha=1/3)
.alias("without_rare_tokens"),
)
We can also use the normalized frequency, or proportion, to exclude tokens.
cl = cl.with_comparisons(
wl.comp("Franklin D. Roosevelt", "Joe Biden")
.exclude.by_proportion(greater_than=0.005)
.score.rank_divergence(alpha=1/3)
.alias("without_common_tokens"),
)
Note
By default, a token is only excluded if it meets the criteria for both corpora. For example, exclude.by_frequency(less_than=10) excludes tokens that appear 10 or fewer times in both the reference and comparison corpora. If it appeared 11 times in one corpus and 3 times in the other, then it would not be excluded.
To apply the exclusion criteria against only one corpus, or either corpus, see the corpus parameter for details.
Warning
A token missing from one corpus has a frequency and proportion of zero there, so specifying corpus="either" and the less_than parameter will remove every token that the two corpora do not share, in addition to tokens below the bound.
Excluding by lexicon score¶
When a comparison is scored by a lexicon, tokens can be excluded by the score that the lexicon assigns them.
For example, say that we score tokens with the labMT sentiment lexicon, which assigns scores on a scale from 1 to 9. To focus on tokens with "strong" sentiment, we can exclude those with "neutral" sentiment by filtering out tokens with scores between 4 and 6 (inclusive).
cl = cl.with_comparisons(
wl.comp("Franklin D. Roosevelt", "Joe Biden")
.score.lexicon(wl.lex.labMT("english"))
.exclude.by_lexicon_score(between=(4.0, 6.0))
.alias("without_neutral_words"),
)
Narrowing to a shared vocabulary¶
Some measures, like the Kullback-Leibler divergence, are only well defined if both corpora have the same vocabulary. We can exclude tokens that are not shared between the two corpora.
cl = cl.with_comparisons(
wl.comp("Franklin D. Roosevelt", "Joe Biden")
.exclude.unshared_tokens()
.score.kullback_leibler_divergence()
.alias("shared_vocabulary_only"),
)
Warning
A token that only appears in one corpus but not the other may strongly indicate how the corpora are different. Excluding it removes that signal. It is preferable to use measures that are robust to different vocabularies, instead of excluding tokens to create a common vocabulary.
Chaining multiple exclusions¶
Exclusions can be chained together, each one narrowing the vocabulary further. Below, we drop our list of stop words and tokens that appear more than 40 or more times in both corpora.
cl = cl.with_comparisons(
wl.comp("Franklin D. Roosevelt", "Joe Biden")
.score.rank_divergence(alpha=1/3)
.exclude.tokens(stop_words)
.exclude.by_frequency(greater_than=40)
.alias("without_stop_words_or_common_tokens"),
)
The tokens "government" and "president" were top distinguishing tokens when we only excluded stop words, but they are removed by the additional frequency exclusion.
print(
cl.comparisons["without_stop_words_or_common_tokens"].df
.select("token", "score_rank", "Franklin D. Roosevelt", "Joe Biden")
.sort("score_rank")
.head(n=10)
)
┌──────────┬────────────┬───────────────────────┬───────────┐
│ token ┆ score_rank ┆ Franklin D. Roosevelt ┆ Joe Biden │
│ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ u32 ┆ u16 ┆ u16 │
╞══════════╪════════════╪═══════════════════════╪═══════════╡
│ men ┆ 1 ┆ 258 ┆ 13 │
│ jobs ┆ 2 ┆ 20 ┆ 147 │
│ applause ┆ 3 ┆ 1 ┆ 143 │
│ such ┆ 4 ┆ 160 ┆ 11 │
│ thank ┆ 5 ┆ 7 ┆ 138 │
│ shall ┆ 6 ┆ 153 ┆ 9 │
│ ukraine ┆ 7 ┆ 0 ┆ 129 │
│ ive ┆ 8 ┆ 1 ┆ 126 │
│ q ┆ 9 ┆ 0 ┆ 120 │
│ upon ┆ 10 ┆ 122 ┆ 6 │
└──────────┴────────────┴───────────────────────┴───────────┘