Skip to content

Make a Comparison

This guide shows you how to get started comparing texts with WordLevel. To follow along, make sure that you first install the package.

Load data

A collection of documents is known as a corpus. We compare multiple corpora by loading them into a catalog. We will use an example dataset of US presidential speeches. Each document is a speech, and each corpus contains all of the speeches by a particular president. We will compare the speeches of Presidents Franklin D. Roosevelt and Joe Biden.

import wordlevel as wl

speeches = wl.Dataset("presidential_speeches")
cl = wl.Catalog(speeches, corpora=["Franklin D. Roosevelt", "Joe Biden"])

Inspect the Catalog

The Catalog stores the corpora together as a dataframe. There is one row per token, and one column per corpus, where the entry is the number of times that token appeared in that corpus. For example, the word "the" was used 9,559 times in speeches by Roosevelt.

Columns with the suffix __wl_normed contain the normalized frequencies of the tokens in each corpus, such that all of the frequency values within a corpus sum to 1.

print(cl.df.head(n=5))
┌───────┬───────────────────────┬───────────┬──────────────────────┬──────────────────────┐
│ token ┆ Franklin D. Roosevelt ┆ Joe Biden ┆ Franklin D.          ┆ Joe Biden__wl_normed │
│ ---   ┆ ---                   ┆ ---       ┆ Roosevelt__wl_normed ┆ ---                  │
│ str   ┆ u16                   ┆ u16       ┆ ---                  ┆ f64                  │
│       ┆                       ┆           ┆ f64                  ┆                      │
╞═══════╪═══════════════════════╪═══════════╪══════════════════════╪══════════════════════╡
│ the   ┆ 9559                  ┆ 5693      ┆ 0.074072             ┆ 0.054592             │
│ of    ┆ 6440                  ┆ 2791      ┆ 0.049903             ┆ 0.026764             │
│ and   ┆ 4675                  ┆ 3702      ┆ 0.036226             ┆ 0.0355               │
│ to    ┆ 4161                  ┆ 3842      ┆ 0.032243             ┆ 0.036842             │
│ in    ┆ 2956                  ┆ 2018      ┆ 0.022906             ┆ 0.019351             │
└───────┴───────────────────────┴───────────┴──────────────────────┴──────────────────────┘

Compare the corpora

WordLevel lets us quantify how two corpora differ token by token (at the "word level"). We do this by creating a comparison between them and using a measure to score their token-level differences.

We will use the rank-turbulence divergence as an example. Each token is ranked by frequency. The larger a token's difference in rank between the corpora, the more it contributes to the divergence.

We make comparisons by using wl.comp and providing the names of the corpora that we want to compare. We give this comparison the alias "roosevelt_biden", which we will use to retrieve it from our Catalog.

cl = cl.with_comparisons(
    wl.comp("Franklin D. Roosevelt", "Joe Biden")
        .score.rank_divergence(alpha=1/3)
        .alias("roosevelt_biden"),
)

Retrieve the comparison scores

The comparison is stored on the Catalog under the alias we gave it, "roosevelt_biden".

roosevelt_biden = cl.comparisons["roosevelt_biden"]

print(
    roosevelt_biden.df.select("token", "comparison_score", "score_rank", "attribution")
    .sort("score_rank")
    .head(n=10)
)
┌────────────┬──────────────────┬────────────┬─────────────┐
│ token      ┆ comparison_score ┆ score_rank ┆ attribution │
│ ---        ┆ ---              ┆ ---        ┆ ---         │
│ str        ┆ f64              ┆ u32        ┆ enum        │
╞════════════╪══════════════════╪════════════╪═════════════╡
│ thats      ┆ 0.000325         ┆ 1          ┆ comparison  │
│ im         ┆ 0.000298         ┆ 2          ┆ comparison  │
│ of         ┆ 0.000269         ┆ 3          ┆ reference   │
│ to         ┆ 0.000269         ┆ 4          ┆ comparison  │
│ which      ┆ 0.000267         ┆ 5          ┆ reference   │
│ applause   ┆ 0.000266         ┆ 6          ┆ comparison  │
│ weve       ┆ 0.000261         ┆ 7          ┆ comparison  │
│ ukraine    ┆ 0.00026          ┆ 8          ┆ comparison  │
│ you        ┆ 0.000258         ┆ 9          ┆ comparison  │
│ government ┆ 0.000256         ┆ 10         ┆ reference   │
└────────────┴──────────────────┴────────────┴─────────────┘

There is one row per token. Each token has a comparison_score, which comes directly from the scoring measure that we chose, the rank-turbulence divergence. It measures how much that token contributes to the difference between the corpora.

For the rank-turbulence divergence, the attribution indicates whether the token was used relatively more in the "reference" corpus or the "comparison" corpus. Because we defined our comparison as wl.comp("Franklin D. Roosevelt", "Joe Biden"), the reference corpus is Roosevelt's speeches, and the comparison corpus is Biden's speeches. Read more about comparisons for details on reference and comparison corpora.

Plot the scores

We can visualize the comparison scores as a bar graph. There is one bar per token, and its direction indicates whether it is attributed to the reference corpus (Roosevelt, left) or the comparison corpus (Biden, right).

chart = cl.plot.bar(
    "roosevelt_biden",
    show_totals=True,
    bar_config={
        "reference_symbol": "Roosevelt",
        "comparison_symbol": "Biden",
    },
)

01lBidenRoosevelt01lTotal0.00050.00040.00030.00020.000100.00010.00020.00030.00040.0005Contribution15101520253035404550Rankmenqshipstrumpnatooftheresjobspowersisraelwhichsuchuponpandemicindustryputinamericaindividualthankproductionyoujapanesethereforegovernmentafghanistangermanletsyoureukrainecantfarmthatsgoinggermanyplanesimaxispresidentdontapplausetalibanivelabortheyrebritishwevetoshallcovidvaccinated

This provides a quick but effective summary of how the two corpora differ. For example, we see that Roosevelt's speeches are characterized by words pertaining to World War II ("japanese", "germany", "british", "axis") and industrial mobilization ("industry", "ships", "labor", "production"), whereas Biden's speeches are characterized by the COVID-19 pandemic ("covid", "vaccinated", "pandemic") and the Russian invasion of Ukraine ("ukraine", "putin", "nato").

Save the figure

WordLevel plots are Altair charts, so they can be saved as PNG, SVG, PDF, or HTML files with the save method. Saving requires some extra dependencies, which are installed with the save extra.

chart.save("roosevelt_biden.png", ppi=300)

Where to go next

Continue to learn how to conduct sentiment analysis of Roosevelt's and Biden's speeches and make a word shift graph, or get started with your own data by constructing a catalog and scoring it.