Skip to content

Scoring Comparisons

Example data

We will use an example dataset of US presidential speeches from Franklin D. Roosevelt and Joe Biden. Each document is a speech, and each corpus contains all of the speeches by that particular president.

import wordlevel as wl

speeches = wl.Dataset("presidential_speeches")
cl = wl.Catalog(speeches, corpora=["Joe Biden", "Franklin D. Roosevelt"])

Scoring a comparison

We will compare Roosevelt's and Biden's speeches by measuring the difference in the normalized frequency of each token. Roosevelt's speeches will be our reference corpus, and Biden's speeches will be our comparison corpus. This means that we are calculating the difference \(p_\tau^{(\text{Biden})} - p_\tau^{(\text{Roosevelt})}\), where \(p_\tau\) is the normalized frequency of token \(\tau\).

We make a comparison by using the constructor wl.comp, and providing it to the Catalog's method with_comparisons.

cl = cl.with_comparisons(
    wl.comp("Franklin D. Roosevelt", "Joe Biden").score.proportion().alias("roosevelt_biden")
)

In the comparison above, wl.comp("Franklin D. Roosevelt", "Joe Biden") defines which corpora we want to compare. The first argument is the reference corpus (Roosevelt) and the second argument is the comparison corpus (Biden). We can use the names of any two corpora in our Catalog to create our comparison.

The method score.proportion() specifies that we want to compare the tokens' normalized frequencies (proportions). All scoring methods are called from the score namespace; see the API reference for other measures.

We call the alias method to give our comparison the name "roosevelt_biden". This is the name that we can use to refer to the comparison when working with the Catalog.

Note

wl.comp does not do anything by itself: it returns a lazily evaluated plan for how to make the comparison. Attaching it to the Catalog is how the comparison is actually executed.

Inspecting a comparison

The method with_comparisons stores the comparison in our Catalog. If you are familiar with Polars, this is similar to how df.with_columns() attaches expressions (columns) to a dataframe.

We can access our comparison using the alias that we gave it, "roosevelt_biden".

roosevelt_biden = cl.comparisons["roosevelt_biden"]

The comparison's dataframe df contains details about the token-by-token differences.

print(
    roosevelt_biden.df.select(
        "token",
        "Franklin D. Roosevelt__wl_normed",
        "Joe Biden__wl_normed",
        "comparison_score",
        "score_rank",
        "attribution"
    )
    .sort("score_rank")
    .head(n=10)
)
┌───────────┬────────────────────┬───────────────────┬──────────────────┬────────────┬─────────────┐
│ token     ┆ Franklin D. Roosev ┆ Joe               ┆ comparison_score ┆ score_rank ┆ attribution │
│ ---       ┆ elt__wl_normed     ┆ Biden__wl_normed  ┆ ---              ┆ ---        ┆ ---         │
│ str       ┆ ---                ┆ ---               ┆ f64              ┆ u32        ┆ enum        │
│           ┆ f64                ┆ f64               ┆                  ┆            ┆             │
╞═══════════╪════════════════════╪═══════════════════╪══════════════════╪════════════╪═════════════╡
│ of        ┆ 0.049903           ┆ 0.026764          ┆ -0.023139        ┆ 1          ┆ reference   │
│ the       ┆ 0.074072           ┆ 0.054592          ┆ -0.01948         ┆ 2          ┆ reference   │
│ you       ┆ 0.002945           ┆ 0.009628          ┆ 0.006683         ┆ 3          ┆ comparison  │
│ to        ┆ 0.032243           ┆ 0.036842          ┆ 0.004599         ┆ 4          ┆ comparison  │
│ which     ┆ 0.005207           ┆ 0.000748          ┆ -0.004459        ┆ 5          ┆ reference   │
│ i         ┆ 0.008973           ┆ 0.013128          ┆ 0.004154         ┆ 6          ┆ comparison  │
│ that      ┆ 0.017435           ┆ 0.013732          ┆ -0.003703        ┆ 7          ┆ reference   │
│ president ┆ 0.000566           ┆ 0.0042            ┆ 0.003634         ┆ 8          ┆ comparison  │
│ in        ┆ 0.022906           ┆ 0.019351          ┆ -0.003555        ┆ 9          ┆ reference   │
│ we        ┆ 0.012119           ┆ 0.015525          ┆ 0.003406         ┆ 10         ┆ comparison  │
└───────────┴────────────────────┴───────────────────┴──────────────────┴────────────┴─────────────┘

The column Joe Biden__wl_normed is \(p_\tau^{(\text{Biden})}\), the normalized frequency of the token in Biden's speeches. Similarly, the column Franklin D. Roosevelt__wl_normed is \(p_\tau^{(\text{Roosevelt})}\). The comparison_score is the difference \(p_\tau^{(\text{Biden})} - p_\tau^{(\text{Roosevelt})}\).

The attribution is determined by whether the score is positive or negative. When it is positive, the token was used more in Biden's speeches, and so it is "attributed" to his corpus (the comparison). When the score is negative, the token was used more in Roosevelt's speeches, and so it is attributed to the reference corpus.

Making multiple comparisons

A Catalog can contain multiple comparisons, differentiated by their aliases. We can construct them by simply calling the method with_comparisons again, which accepts any number of comparisons.

cl = cl.with_comparisons(
    wl.comp("Franklin D. Roosevelt", "Joe Biden")
        .score.rank_divergence(alpha=1/3)
        .alias("roosevelt_biden_rank_divergence"),
    wl.comp("Franklin D. Roosevelt", "Joe Biden")
        .score.shannon_entropy()
        .alias("roosevelt_biden_entropy")
)

Now, if we look at the catalog's comparisons, we see that the new ones are stored alongside the original comparison.

print(cl.comparisons.keys())
dict_keys(['roosevelt_biden', 'roosevelt_biden_rank_divergence', 'roosevelt_biden_entropy'])