Skip to content

Word Shift Graphs

Weighted averages

Many measures for comparing corpora can be represented as a difference in weighted averages. Let \(p_\tau\) be the relative frequency of a token \(\tau\) and let \(\phi_\tau\) be the score assigned to that token. Then we can compute the difference:

\[ \Phi = \sum_\tau p_\tau^{(C)} \phi_\tau - \sum_\tau p_\tau^{(R)} \phi_\tau = \sum_\tau \phi_\tau \left(p_\tau^{(C)} - p_\tau^{(R)} \right) \]

where \(C\) and \(R\) indicate the comparison and reference corpora respectively.

Observe how the overall difference is a linear sum over the tokens. This means that we can measure how each token contributes individually to the difference. We denote the contribution as \(\delta_\tau\). Further, its sign is determined entirely by whether \(p_\tau^{(C)}\) or \(p_\tau^{(R)}\) is larger—if it's positive, then we know that the token was used more in the comparison corpus, and vice versa.

Reference scores

Sometimes, we can interpret each token's score qualitatively. For example, suppose that we have a sentiment lexicon that assigns a score \(\phi_\tau\) to each token on a scale from 1 to 9, where 1 is the most "negative" a token can be, and 9 is the most "positive" it can be. The middle of this scale is 5. Therefore, if a token has a score below 5, then it is "relatively negative," and if it has a score above 5, then it is "relatively positive."

We can encode the qualitative notion of a score being "relatively" positive or negative by introducing a reference score, \(\Phi^{(\text{ref})}\). It is mathematically equivalent to rewrite our difference in weighted averages \(\Phi\) as follows.

\[ \Phi = \sum_\tau \underbrace{\left(p_\tau^{(C)} - p_\tau^{(R)} \right)}_{\uparrow / \downarrow} \underbrace{\left(\phi_\tau - \Phi^{(\text{ref})} \right)}_{+ / -} \]

The contribution of each token \(\delta_\tau\) now depends on two things: whether it was used more or less in the comparison corpus than the reference corpus \(\left(\uparrow / \downarrow \right)\), and whether its score is relatively positive or negative \(\left( + / - \right)\). This means that there are four ways that a token can contribute, based on the signs of these components:

  • \(\left(+ \uparrow \right)\) A relatively positive token is used more.
  • \(\left(+ \downarrow \right)\) A relatively positive token is used less.
  • \(\left(- \uparrow \right)\) A relatively negative token is used more.
  • \(\left(- \downarrow \right)\) A relatively negative token is used less.

Basic word shifts

We rank the contributions by magnitude and partition them by these different qualitative interpretations. This is what is plotted in a word shift graph.

import wordlevel as wl

cl = (
    wl.Catalog(
        wl.Dataset("presidential_speeches"),
        corpora=["Franklin D. Roosevelt", "Joe Biden"],
    )
    .with_comparisons(
        wl.comp("Franklin D. Roosevelt", "Joe Biden")
            .score.lexicon(wl.lex.labMT("english"), reference_score="center")
            .alias("roosevelt_biden_sentiment"),
    )
)

chart = cl.plot.shift("roosevelt_biden_sentiment", show_totals=True)

0102l+↓+↑−↓−↑01lTotal−0.010−0.008−0.006−0.004−0.0020.0000.0020.0040.0060.0080.010Contribution15101520253035404550Rankcancertogetherithankbevictoryillchildrendemocracyjobswillyouwellpresentupouramericamymoregettaxdoafghanistanviolencefightingnationallovelikemepeaceweinmillionmenallhaveknowcantpresidentfamilyfamiliesdontoftruthgreatwaramericansamericangodjust

In the word shift above, we are comparing the presidential speeches of Biden (the comparison corpus) to Roosevelt (the reference corpus). Overall, Biden's speeches have a higher average sentiment than Roosevelt's speeches. Two types of contributions directly reinforce this:

  • \(\left(+ \uparrow \right)\) Relatively positive tokens are used more ("you", "america", "we", etc.).
  • \(\left(- \downarrow \right)\) Relatively negative tokens are used less ("war", "fighting", etc.).

Some tokens offset the higher sentiment, meaning that the difference would have been even more positive without them:

  • \(\left(+ \downarrow \right)\) Relatively positive tokens are used less ("great", "peace", "will", etc.).
  • \(\left(- \uparrow \right)\) Relatively negative tokens are used more ("violence", "tax", "cancer", etc.).

Corpus-specific scores

Scores may be associated with a specific corpus. For example, we may have sentiment lexicons that are constructed for the context of each corpus—such as lexicons that encode how positive or negative words were in different years. Note, when scores can depend on a specific corpus, they can be a function of a token's frequency itself. The Shannon entropy of a corpus, for example, measures the contribution of each token as \(p_\tau \log 1 / p_\tau\): \(p_\tau\) is the frequency, and \(\log 1 / p_\tau\) is the score.

We can represent corpus-specific scores by extending the word shift framework. The contributions \(\delta_\tau\) of each token \(\tau\) can be expressed as

\[ \delta_\tau = \underbrace{ \left( p_\tau^{(C)} - p_\tau^{(R)} \right) }_{\uparrow / \downarrow} \underbrace{ \left[ \frac{1}{2} \left( \phi_\tau^{(C)} + \phi_\tau^{(R)} \right) - \Phi^{(\text{ref})} \right] }_{+ / -} + \underbrace{ \frac{1}{2} \left( p_\tau^{(C)} + p_\tau^{(R)} \right) \left( \phi_\tau^{(C)} - \phi_\tau^{(R)} \right) }_{\bigtriangleup / \bigtriangledown} \]

Like before, we can see that a token's contribution depends on whether it was used more or less \(\left(\uparrow / \downarrow \right)\), and whether its average score is relatively positive or negative \(\left( + / -\right)\). However, it also now depends on the difference in the scores themselves \(\left( \bigtriangleup / \bigtriangledown \right)\).

  • \(\left(\bigtriangleup \right)\) A token's score is higher within the comparison corpus.
  • \(\left(\bigtriangledown \right)\) A token's score is higher within the reference corpus.

Detailed word shifts

Note how the new term in our word shift decomposition is additive. We represent this in the word shift graph as a stacked bar.

cl = cl.with_comparisons(
    wl.comp("Franklin D. Roosevelt", "Joe Biden")
        .score.lexicon(
            lexicon_reference=wl.lex.SocialSent("1940"),
            lexicon_comparison=wl.lex.SocialSent("2000"),
            reference_score="center",
        )
        .alias("roosevelt_biden_sentiment_by_decade"),
)

chart = cl.plot.shift("roosevelt_biden_sentiment_by_decade", show_totals=True)

010203l▽−↓+↑△+↓−↑01lTotal−0.008−0.006−0.004−0.0020.0000.0020.0040.0060.008Contribution15101520253035404550Rankeverymanyclearthese*surepeacepresidentlifeaircompleteq*fearammenfamiliesgodwholeactenemiestogetherdoingwell*jobgoingbigcountryagainstdemocracygotrightafghanistan*americalivingapplause*thosesaylivesviolencemr*tonightimportantgreatamericansmosthomenationpartthankhomesshall

This word shift graph tells a different story than before: when we use decade-specific sentiment lexicons (from SocialSent), Biden's speeches are more negative than Roosevelt's speeches. Most of this is driven by specific tokens being associated with more negative sentiment.

  • \(\left( \bigtriangledown \right)\) Tokens have more negative scores ("nation", "americans", "country", etc.). These contribute to the lower sentiment.
  • \(\left( \bigtriangleup \right)\) Tokens have more positive scores ("together", "thank", "president", etc.). These offset the negative sentiment difference; without them, the difference would be even more negative.

The two components of each contribution can counteract each other (one can be positive, the other can be negative). We display this by fading the bars by the counteracted component. For example, "nation" contributes negatively because its score is lower, but that is offset partly by the fact that it is a relatively negative token that was used less.

Further reading

Gallagher, R. J., Frank, M. R., Mitchell, L., Schwartz, A. J., Reagan, A. J., Danforth, C. M., & Dodds, P. S. (2021). Generalized word shift graphs: a method for visualizing and explaining pairwise comparisons between texts. EPJ Data Science, 10(1), 4.