Plotting Word Shifts¶
All comparisons can be visualized as a bar graph. If the comparisons is a difference in weighted averages, or can be represented as one, then it can be plotted as a word shift graph. In addition to showing which tokens contribute to the difference, a word shift graph also shows how they contribute.
Tip
A comparison records whether a weighted average measure was used under its is_weighted_avg attribute.
Example dataset¶
We will use an example dataset of US presidential speeches from Franklin D. Roosevelt and Joe Biden. Each document is a speech, and each corpus contains all of the speeches by that particular president. See the cookbook for more about working with catalogs.
We make two comparisons using sentiment lexicons. The first comparison uses the labMT lexicon for both corpora. The second comparison uses the SocialSent lexicons, which are specific to each decade. We score Roosevelt's speeches with the lexicon from the 1940s, and Biden's speeches with the lexicon from the 2000s, the most recent decade available.
import wordlevel as wl
speeches = wl.Dataset("presidential_speeches")
cl = wl.Catalog(speeches, corpora=["Franklin D. Roosevelt", "Joe Biden"])
cl = cl.with_comparisons(
wl.comp("Franklin D. Roosevelt", "Joe Biden")
.score.lexicon(wl.lex.labMT("english"), reference_score="center")
.normalize(by="total_diff")
.alias("roosevelt_biden_labmt"),
wl.comp("Franklin D. Roosevelt", "Joe Biden")
.score.lexicon(
lexicon_reference=wl.lex.SocialSent("1940"),
lexicon_comparison=wl.lex.SocialSent("2000"),
reference_score="center",
)
.normalize(by="total_diff")
.alias("roosevelt_biden_socialsent"),
)
Word shift graph¶
We can create a word shift graph by passing a comparison's name to Catalog.plot.shift. The name is the alias that we gave it when adding the comparison to the Catalog.
Total contributions¶
Word shift graphs indicate the different ways that tokens contribute to the difference. A token can be used more or less in the comparison corpus versus the reference corpus \(\left( \uparrow / \downarrow \right)\), and it can have a relatively positive or negative score \((+ / -)\). This means a token can contribute in four ways:
- \(\left(+ \uparrow \right)\) A relatively positive token is used more in the comparison corpus.
- \(\left(+ \downarrow \right)\) A relatively positive token is used less in the comparison corpus.
- \(\left(- \uparrow \right)\) A relatively negative token is used more in the comparison corpus.
- \(\left(- \downarrow \right)\) A relatively negative token is used less in the comparison corpus.
If we set show_totals=True, we can see the relative magnitude of the different kinds of contributions.
Single lexicon¶
When both corpora are scored with the same lexicon, tokens contribute in the four ways described above.
The overall difference (the "Total") indicates that Biden's speeches have a higher sentiment than Roosevelt's speeches. This is because Biden's speeches use relatively positive words more \(\left(+ \uparrow \right)\), such as "you" and "america," and relatively negative words less \(\left(- \downarrow \right)\), such as "war". The sentiment difference is partly offset by Biden's speeches using some relatively positive words less \(\left(+ \downarrow \right)\), such as "great" and "peace," and some relatively negative words more \(\left(- \uparrow \right)\), such as "violence." These words offset how much more positive Biden's speeches would have been otherwise.
The bar colors can be configured with a ShiftBarConfig.
chart = cl.plot.shift(
"roosevelt_biden_labmt",
max_rank=20,
show_totals=True,
bar_config={
"more_of_positive": "#009E73",
"less_of_positive": "#A6DBC8",
"more_of_negative": "#7B3294",
"less_of_negative": "#C9A8D9",
},
)
Multiple lexicons¶
When each corpus is scored with its own lexicon, each token can have a different score in each of those lexicons. This means there are two other ways a token can contribute to the corpora's difference:
- \(\left(\bigtriangleup \right)\) A token has a higher score in the comparison corpus's lexicon
- \(\left(\bigtriangledown \right)\) A token has a lower score in the comparison corpus's lexicon
These contributions are additive to the other ways that a token can contribute, and so they are shown as an additional stacked bar.
The total is negative, so Biden's speeches have a lower sentiment than Roosevelt's speeches when each is scored with the lexicon of its time. The totals panels show that much of this difference comes from words having lower scores in the 2000s \(\left(\bigtriangledown \right)\), such as "nation," "families," and "americans."
The different kinds of contributions can partially cancel each other out. When this happens, we show it by fading the canceled out portions. For example, "violence" is a relatively negative word that was used more in Biden's speeches. However, that contribution is partly offset by the score change, because "violence" also has a relatively higher sentiment score in the 2000s lexicon than the 1940s lexicon.
The higher and lower score contribution colors can also be configured with a ShiftBarConfig.
chart = cl.plot.shift(
"roosevelt_biden_socialsent",
max_rank=20,
show_totals=True,
bar_config={
"higher_score": "#D55E00",
"lower_score": "#009E73",
"higher_score_counteracted": "#F2BFA0",
"lower_score_counteracted": "#A6DBC8",
},
)
Saving the figure¶
WordLevel plots are Altair charts, so they can be saved as PNG, SVG, PDF, or HTML files with the save method. Saving requires some extra dependencies, which are installed with the save extra.