Skip to content

Make a Word Shift Graph

This guide shows you how to make a word shift graph with WordLevel. To follow along, make sure that you first install the package, and read the introduction on making comparisons.

We will use an example dataset of US presidential speeches from Franklin D. Roosevelt and Joe Biden. Each document is a speech, and each corpus contains all of the speeches by that particular president.

import wordlevel as wl

speeches = wl.Dataset("presidential_speeches")
cl = wl.Catalog(speeches, corpora=["Franklin D. Roosevelt", "Joe Biden"])

Load a sentiment lexicon

A lexicon assigns a "score" to each token in its vocabulary. For example, the labMT sentiment lexicon maps each token to a score between 1 and 9, where 1 indicates that the token is associated with negativity ("sadness") and 9 indicates the token is associated with positivity ("happiness").

WordLevel provides several sentiment lexicons previously constructed by researchers. We will use the English labMT lexicon as an example. Lexicons are loaded using wl.lex.

labmt = wl.lex.labMT("english")

Note

The first time that you load a lexicon, it is downloaded from Hugging Face. The lexicon is cached locally, so subsequent calls do not download it again.

Score the corpora

Like our first comparison, we compare the presidential speeches using wl.comp. This time, we score them by passing the labmt sentiment lexicon to the lexicon method.

cl = cl.with_comparisons(
    wl.comp("Franklin D. Roosevelt", "Joe Biden")
        .score.lexicon(labmt, reference_score=labmt.center)
        .alias("roosevelt_biden_sentiment"),
)

We pass a reference_score when scoring the words—in this case it is the center of labMT's 1-to-9 sentiment scale, which is 5. Each token's score is compared to this reference score: if it is greater than 5, then it is "relatively positive"; if it is less than 5, then it is "relatively negative."

Plot the word shift

A comparison bar graph shows which tokens contribute to the difference between two corpora. A word shift graph also shows how they contribute. Here, there are four types of contributions:

  • \(\left(+ \uparrow \right)\) A relatively positive word is used more in the comparison corpus
  • \(\left(+ \downarrow \right)\) A relatively positive word is used less in the comparison corpus
  • \(\left(- \uparrow \right)\) A relatively negative word is used more in the comparison corpus
  • \(\left(- \downarrow \right)\) A relatively negative word is used less in the comparison corpus
chart = cl.plot.shift("roosevelt_biden_sentiment", show_totals=True)

0102l+↓−↑+↑−↓01lTotal−0.010−0.008−0.006−0.004−0.0020.0000.0020.0040.0060.0080.010Contribution15101520253035404550Rankourmillioniupchildrenmyinmendontjustvictoryalltogetherillamericaamericansbedemocracypresidenttruthamericannationalfamiliesfightingtaxhavewarmorecancercantwillgreatwelllikewepresentknowgodafghanistanofthankmelovegetyoudoviolencepeacejobsfamily

The total difference is positive, which means that the comparison corpus (Biden) has a higher average sentiment than the reference corpus (Roosevelt). The tokens are interpreted with respect to this difference. For example, Biden's speeches have a higher sentiment than Roosevelt's because they use relatively positive words more ("you", "america", "thank") and relatively negative words less ("war", "fighting"). The higher sentiment is partly offset by other words, such as Biden's speeches using some relatively positive words less ("great", "peace", "victory") and some relatively negative words more ("violence", "dont", "tax").

Save the figure

WordLevel plots are Altair charts, so they can be saved as PNG, SVG, PDF, or HTML files with the save method. Saving requires some extra dependencies, which are installed with the save extra.

chart.save("word_shift.png", ppi=300)

Where to go next

Learn more about word shift graphs and how to plot them, or get started with your own data by constructing a catalog and scoring it.