Skip to content

Scoring with Lexicons

A lexicon assigns scores to tokens. For example, sentiment dictionaries are some of the most popular lexicons: tokens are given a score based on how "positive" or "negative" they are.

When tokens are given scores, we can score an entire corpus by taking the weighted average \(\sum \phi_\tau \, p_\tau\), where \(\phi_\tau\) is the score given to a token \(\tau\) by the lexicon, and \(p_\tau\) is how relatively often it appears in the corpus (the normalized frequency).

Example data

We will use two small corpora and put them into a Catalog. See the cookbook for more about working with catalogs.

import wordlevel as wl

before = {"good": 40, "bad": 20, "happy": 30, "sad": 10, "the": 100}
after = {"good": 20, "bad": 45, "happy": 15, "sad": 35, "the": 100}

cl = wl.Catalog({"before": before, "after": after})

Loading a lexicon

WordLevel provides several pre-made lexicons, constructed by academic researchers. They can be loaded using the wl.lex operator. Below, we load the labMT sentiment lexicon, which rates how positive ("happy") or negative ("sad") a word is on a scale from 1 to 9.

labmt = wl.lex.labMT("english")

Note

The first time that you load a lexicon it is downloaded from Hugging Face. The lexicon is cached locally, so subsequent calls do not download it repeatedly.

Inspecting a lexicon

The lexicon's mapping between tokens and scores is stored as a dataframe df.

print(labmt.df.head(n=5))
┌───────────┬───────┐
│ token     ┆ score │
│ ---       ┆ ---   │
│ str       ┆ f64   │
╞═══════════╪═══════╡
│ laughter  ┆ 8.5   │
│ happiness ┆ 8.44  │
│ love      ┆ 8.42  │
│ happy     ┆ 8.3   │
│ laughed   ┆ 8.26  │
└───────────┴───────┘

The lexicon also records the center of its scale. For labMT, this is the score at the middle of the scale indicating that a word is "neutral"—neither positive nor negative.

print(f"Center of the lexicon's scale: {labmt.center}")
Center of the lexicon's scale: 5.0

Scoring a comparison

We can compare our corpora by taking the difference in their weighted averages. That is, we can compute

\[ \sum \phi_\tau \, p^{(C)}_\tau - \phi_\tau \, p^{(R)}_\tau \]

where \(p_\tau^{(C)}\) and \(p_\tau^{(R)}\) are the normalized frequencies of a token \(\tau\) in the comparison corpus \(C\) and reference corpus \(R\) respectively.

To do this, we can pass our lexicon directly to the comparison's score.lexicon method. See the cookbook for more about scoring comparisons in general.

cl = cl.with_comparisons(wl.comp("before", "after").score.lexicon(labmt).alias("sentiment"))

Scoring with a custom lexicon

You can also use a custom lexicon to make comparisons. The score.lexicon method accepts a dictionary or dataframe that maps tokens to scores.

custom_lexicon = {"good": 4.5, "bad": 1.5, "happy": 5.0, "sad": 1.0}

cl = cl.with_comparisons(
    wl.comp("before", "after")
        .score.lexicon(custom_lexicon)
        .alias("custom_sentiment"),
)

Dropped tokens

Lexicons can only score the tokens that appear in them. All other tokens are dropped from the calculation. For example, "the" is the most common token in both of our corpora, but there is no score for it in our custom lexicon, so it is dropped from the comparison.

custom = cl.comparisons["custom_sentiment"]

print(custom.df.select("token", "comparison_score").sort("token"))
┌───────┬──────────────────┐
│ token ┆ comparison_score │
│ ---   ┆ ---              │
│ str   ┆ f64              │
╞═══════╪══════════════════╡
│ bad   ┆ 0.286957         │
│ good  ┆ -1.017391        │
│ happy ┆ -0.847826        │
│ sad   ┆ 0.204348         │
└───────┴──────────────────┘

Scoring with multiple lexicons

At times, we may have a different lexicon for each corpora. For example, tokens may be associated with different sentiments at different periods of time (e.g. "catfish" or "amazon" in 1920 vs. 2020) or social contexts (e.g. "trump" in left- and right-leaning news outlets).

We can still compute the difference in the weighted averages of the corpora, but now the scores depend on the reference and comparison corpora.

\[ \sum \phi^{(C)}_\tau \, p^{(C)}_\tau - \phi^{(R)}_\tau \, p^{(R)}_\tau \]

For our example data, we create two custom lexicons: one associated with our before corpus and one associated with our after corpus. We then pass these to the lexicon_reference and lexicon_comparison parameters when scoring the comparison.

scores_before = {"good": 8.0, "bad": 2.0, "happy": 9.0}
scores_after = {"good": 7.5, "bad": 1.5, "sad": 1.0}

cl = cl.with_comparisons(
    wl.comp("before", "after")
        .score.lexicon(
            lexicon_reference=scores_before,
            lexicon_comparison=scores_after,
            scored_in_one_lexicon="borrow",
        )
        .alias("drifting_sentiment"),
)

Borrowing scores

When we have a different lexicons per corpus, some tokens may be scored in both lexicons, but others may only appear in one or the other. For example, the words "doomscrolling" and "deepfake" did not exist until the 2010s, so they cannot have a sentiment score in a lexicon specific to the 1920s. Other words like "apple" and "cloud" are used more often today than in the past (because of their associations with technology), and so they may not appear in a lexicon of frequently used words in the 1920s.

In these cases, we can allow a token to "borrow" a score across lexicons. Effectively, this allows us to just compare the change in the relative frequency of the token, even if we cannot measure the change in its score across the corpora.

In our example lexicons, "happy" is only defined in the before lexicon, and "sad" is only defined in the after lexicon. We can see that those tokens "borrow" their scores across lexicons because we set scored_in_one_lexicon="borrow" when creating our comparison. If we instead set scored_in_one_lexicon="drop", those tokens would be dropped, and only tokens scored by both lexicons would contribute to the comparison.

drifting = cl.comparisons["drifting_sentiment"]

print(drifting.df.select("token", "score_ref", "score_cmp", "comparison_score").sort("token"))
┌───────┬───────────┬───────────┬──────────────────┐
│ token ┆ score_ref ┆ score_cmp ┆ comparison_score │
│ ---   ┆ ---       ┆ ---       ┆ ---              │
│ str   ┆ f64       ┆ f64       ┆ f64              │
╞═══════╪═══════════╪═══════════╪══════════════════╡
│ bad   ┆ 2.0       ┆ 1.5       ┆ 0.186957         │
│ good  ┆ 8.0       ┆ 7.5       ┆ -1.895652        │
│ happy ┆ 9.0       ┆ 9.0       ┆ -1.526087        │
│ sad   ┆ 1.0       ┆ 1.0       ┆ 0.204348         │
└───────┴───────────┴───────────┴──────────────────┘

Setting reference scores

Many lexicons—particularly sentiment lexicons—are built on continuous scales from a score that is most "negative" to a score that is most "positive." This means each token is relatively negative or positive: "sad" has a negative sentiment, while "happy" has a positive sentiment. If we distinguish between these tokens, it can change how we interpret our comparison. For example, a corpus can have a higher sentiment than another because it either used more positive words or it used fewer negative words.

To encode this, we can pass a reference score to the comparison. Each score is compared to this anchor, determining whether it is relatively positive or negative.

# Our custom lexicon scores tokens from 1 to 5, so the center of its scale is 3
cl = cl.with_comparisons(
    wl.comp("before", "after")
        .score.lexicon(custom_lexicon, reference_score=3.0)
        .alias("centered_sentiment"),
)

Warning

Since the reference score can change the per-token scores (though not the overall difference), it should only be used if you understand how to interpret word shift graphs.