Skip to content

Plotting Comparisons

Comparisons measure how much each token contributes to the difference between two corpora. We can visualize those differences by plotting them as a bar graph.

Example dataset

We will use an example dataset of US presidential speeches from Franklin D. Roosevelt and Joe Biden. Each document is a speech, and each corpus contains all of the speeches by that particular president. See the cookbook for more about working with catalogs.

We make two comparisons: one using the Jensen-Shannon divergence and one using the difference in Shannon entropy.

import wordlevel as wl

speeches = wl.Dataset("presidential_speeches")
cl = wl.Catalog(speeches, corpora=["Franklin D. Roosevelt", "Joe Biden"])

cl = cl.with_comparisons(
    wl.comp("Franklin D. Roosevelt", "Joe Biden")
        .score.jsd()
        .normalize(by="total_diff")
        .alias("roosevelt_biden_jsd"),
    wl.comp("Franklin D. Roosevelt", "Joe Biden")
        .score.shannon_entropy()
        .normalize(by="total_diff")
        .alias("roosevelt_biden_entropy"),
)

Bar plot

Once we make a comparison, we can plot it by passing its name to Catalog.plot.bar. The name is the alias that we gave it when adding the comparison to the Catalog.

chart = cl.plot.bar("roosevelt_biden_jsd")

1.5%1%0.5%0%0.5%1%1.5%Percent contribution15101520253035404550Rankgetthankgovernmentapplauseamericajobsdontcantivesuchjapanesegermanyjustshallletscovidqiupongoingshipstheyrewhatwereyouofafghanistannatoukrainemenindustryviolencelabortheimmywellthatswaraboutitstheresbyweveputinisraelpresidentyourewhichhe

Total contributions

There are two things that can give more context to the comparison bar graph. First, we can add a legend indicating what each bar means. Second, we can show the total contributions, and how they are attributed. The attribution depends on the type of comparison we made.

Both of these aspects are added to the plot if we set show_totals=True, which appends panels to the top of the bar graph summing the contributions of all tokens, including those that are not shown in the graph.

Attributed by corpus

Some comparison measures attribute each token to a specific corpus, based on whether its relatively more frequent in one or the other. For example, with the Jensen-Shannon divergence, a token used more in Roosevelt's speeches is attributed to the reference corpus, and a token used more in Biden's speeches is attributed to the comparison corpus. Each bar points in the direction of the attribution.

In the example below, "you," "ukraine," and "afghanistan" are attributed to Biden's speeches (the comparison corpus), and "of," "government," and "men" are attributed to Roosevelt's speeches (the reference corpus).

chart = cl.plot.bar("roosevelt_biden_jsd", max_rank=20, show_totals=True)

01lReferenceComparison01lTotal1.5%1%0.5%0%0.5%1%1.5%Percent contribution15101520Rankafghanistanthegoingivethankamericaimthatspresidentofapplausedontwevequkraineletsyoumengovernmentwhich

If a comparison attributes by corpus, then its bars can be configured with a BarConfig. For example, we can change their colors and the labels of the total contributions. Note, we can pass a dictionary instead of a BarConfig instance, and we do not have to specify every field—only the ones that we want to change.

chart = cl.plot.bar(
    "roosevelt_biden_jsd",
    max_rank=20,
    show_totals=True,
    bar_config={
        "reference": "#009E73",
        "comparison": "#7B3294",
        "reference_symbol": "Roosevelt",
        "comparison_symbol": "Biden",
    },
)

01lBidenRoosevelt01lTotal1.5%1%0.5%0%0.5%1%1.5%Percent contribution15101520Rankmenafghanistanwhichimdontgovernmentthethankukraineqpresidentiveapplauseletsgoingweveofyouamericathats

Attributed by sign

Some comparison measures cannot attribute each token to a specific corpus because there is a more complex interplay between how often a token appears and how it "scores" in a corpus. For example, when we compare two corpora by entropy, both the normalized frequency of a token \(p_\tau\) and its surprisal \(\log 1 / p_\tau\) in both corpora affect the final token-level contribution.

For these measures, we attribute a token by whether it reinforces or offsets the overall difference between the corpora. In the example below, we can see that Biden's speeches have a higher entropy than Roosevelt's speeches because the overall difference (the "Total") is positive. An individual token reinforces that difference when its contribution is also positive, pushing the entropy of Biden's speeches further above that of Roosevelt's speeches. A token offsets that difference when its contribution is negative, pulling back the difference such that it is not as large as it could be otherwise.

chart = cl.plot.bar("roosevelt_biden_entropy", max_rank=20, show_totals=True)

01lReinforcesOffsets01lTotal-1.2%-1%-0.8%-0.6%-0.4%-0.2%0%0.2%0.4%0.6%0.8%1%1.2%Percent contribution15101520Rankamericawhatiwhichthatsgoinggovernmentwetowereyoumenofpresidentgetitsthattheimby

If a comparison attributes by reinforcements and offsets, then its bars can be configured with a ScoreSignBarConfig. Like before, we can pass a dictionary instead of a ScoreSignBarConfig instance, and we do not have to specify every field—only the ones that we want to change.

chart = cl.plot.bar(
    "roosevelt_biden_entropy",
    max_rank=20,
    show_totals=True,
    bar_config={
        "reinforce": "#D55E00",
        "offset": "#CC79A7",
        "reinforce_symbol": "Raises",
        "offset_symbol": "Lowers",
    },
)

01lRaisesLowers01lTotal-1.2%-1%-0.8%-0.6%-0.4%-0.2%0%0.2%0.4%0.6%0.8%1%1.2%Percent contribution15101520Rankitsyouimthemenwhatipresidentthatstogetweofamericagovernmentgoingthatwhichbywere

Size and orientation

Number of tokens

As we have seen in some of the examples already, the max_rank parameter can be used to set how many tokens are shown in the bar graph. By default, the top 50 contributing tokens are displayed.

chart = cl.plot.bar("roosevelt_biden_jsd", max_rank=20)

1.5%1%0.5%0%0.5%1%1.5%Percent contribution15101520Rankgoingletsthethankapplauseiveimwhichdontqweveyouafghanistanmenukraineamericapresidentgovernmentthatsof

Orientation

By default, the bar graph is vertical. This makes it easier for a reader to interpret the tokens because they can be read horizontally. However, we can plot the bar graph horizontally by setting vertical=False. The rotation key of a LabelConfig can be used to rotate the bar labels so that they can be read more easily even in the horizontal layout.

chart = cl.plot.bar(
    "roosevelt_biden_jsd",
    max_rank=20,
    vertical=False,
    label_config={"rotation": -45},
)

15101520Rank1.5%1%0.5%0%0.5%1%1.5%Percent contributiongovernmentofafghanistanthatsimukraineivewevewhichpresidentgoingyouletsthankqthedontapplauseamericamen

Height and width

When the bar graph is vertical, the width parameter controls the width of the graph in pixels.

chart = cl.plot.bar("roosevelt_biden_jsd", max_rank=20, width=500)

1.6%1.4%1.2%1%0.8%0.6%0.4%0.2%0%0.2%0.4%0.6%0.8%1%1.2%1.4%1.6%Percent contribution15101520Rankthegoingthatsgovernmentapplausepresidentqimweveiveukraineyouamericaafghanistanofthankwhichletsdontmen

For a vertical bar graph, its height is automatically determined from two things: the number of tokens displayed, and how wide each bar is. The number of tokens is set by max_rank. The bar_width is specified as a configuration of the rank axis, which is the vertical axis. It indicates how many pixels to give each token vertically. Other aspects of the height can be controlled through some of the other RankAxisConfig parameters.

chart = cl.plot.bar(
    "roosevelt_biden_jsd",
    max_rank=5,
    rank_axis={"bar_width": 24},
)

1.5%1%0.5%0%0.5%1%1.5%Percent contribution15Rankyoupresidentthatswhichof

For a horizontal bar graph, when vertical=False, the width of the bars is still controlled by the RankAxisConfig because the rank axis is now the horizontal axis. However, the width parameter is ignored, and height becomes the tunable parameter that can be used to specify the height of the graph.

chart = cl.plot.bar(
    "roosevelt_biden_jsd",
    max_rank=20,
    vertical=False,
    height=300,
    rank_axis={"bar_width": 24},
)

15101520Rank1.5%1%0.5%0%0.5%1%1.5%Percent contributionimapplauseafghanistandontofweveamericaqyoupresidentgovernmentmenwhichukraineivethegoingletsthatsthank

Configuration

Axes

The score axis measures the magnitude of each token's contribution, and the rank axis indicates how tokens are ordered by those contributions. Their titles, font sizes, colors, and other aspects can be configured using a ScoreAxisConfig and RankAxisConfig respectively.

chart = cl.plot.bar(
    "roosevelt_biden_jsd",
    max_rank=20,
    rank_axis={
        "title": "Position",
        "title_color": "#7B3294",
        "label_tick_step": 3,
        "label_fontsize": 12,
    },
    score_axis={
        "title": "Percent contribution to the JSD",
        "title_fontsize": 14,
        "label_color": "#7B3294",
        "grid": True,
    },
)

1.5%1%0.5%0%0.5%1%1.5%Percent contribution to the JSD1369121518Positionapplauseimyoutheafghanistanthatsmengoinggovernmentpresidentivewhichqukrainewevethankdontofamericalets

Labels

Each bar is labeled with its corresponding token. How labels are displayed can be configured with a LabelConfig.

chart = cl.plot.bar(
    "roosevelt_biden_jsd",
    max_rank=20,
    label_config={"color": "#333333", "size": 13, "weight": "bold", "pad": 6},
)

1.5%1%0.5%0%0.5%1%1.5%Percent contribution15101520Rankgoingthatsyouthankukrainepresidentweveamericathedontapplauseletsgovernmentofimiveafghanistanwhichqmen

Totals

When the total contributions are shown (show_totals=True), the color and label of the grand total can be configured with a TotalShiftConfig. To configure the colors and labels of the totals per attribution, see above for using the bar_config.

chart = cl.plot.bar(
    "roosevelt_biden_jsd",
    max_rank=20,
    show_totals=True,
    total_shift_config={"total": "#D55E00", "total_symbol": "JSD"},
)

01lReferenceComparison01lJSD1.5%1%0.5%0%0.5%1%1.5%Percent contribution15101520Rankyouthatsiveapplauseletstheafghanistanamericaofdontthankwevewhichimmengovernmentpresidentgoingukraineq

Borders

The border and zero line can also be configured stylistically using a BorderConfig and ZeroLineConfig.

chart = cl.plot.bar(
    "roosevelt_biden_jsd",
    max_rank=20,
    border_config={"color": "#7B3294", "width": 1},
    zero_line_config={"color": "#009E73", "width": 2},
)

1.5%1%0.5%0%0.5%1%1.5%Percent contribution15101520Rankpresidentukrainewhichgovernmentimmenapplauseafghanistaniveweveyouofthankqthatsthedontletsgoingamerica

Saving the figure

WordLevel plots are Altair charts, so they can be saved as PNG, SVG, PDF, or HTML files with the save method. Saving requires some extra dependencies, which are installed with the save extra.

chart.save("bar_graph.png", ppi=300)