Normalizing Scores¶
We can put comparison scores on a common scale by normalizing them. This allows us to easily interpret the scores relative to one another and across comparisons, rather than in terms of their raw units.
In a comparison, each token \(\tau\) makes a contribution \(\delta_\tau\) to the difference. The overall difference between the two corpora is \(\sum \delta_\tau\).
Example data¶
We will use an example dataset of US presidential speeches from Franklin D. Roosevelt and Joe Biden. Each document is a speech, and each corpus contains all of the speeches by that particular president. See the cookbook for more about working with catalogs.
import wordlevel as wl
speeches = wl.Dataset("presidential_speeches")
cl = wl.Catalog(speeches, corpora=["Joe Biden", "Franklin D. Roosevelt"])
Normalizing by total difference¶
The first way that we can normalize scores is by the total difference, which is equivalent to dividing all of the scores by \(\sum | \delta_\tau |\). This allows us to interpret each token's score as a portion of the total difference between the two corpora.
cl = cl.with_comparisons(
wl.comp("Franklin D. Roosevelt", "Joe Biden")
.score.rank_divergence(alpha=1/3)
.normalize(by="total_diff")
.alias("roosevelt_biden_rank_divergence_normed"),
)
Normalizing by net difference¶
The second way that we can normalize is by the net difference, which is equivalent to dividing all of the scores by \(\left| \sum \delta_\tau \right|\). This allows us to interpret each token's score as its ratio to the net difference.
cl = cl.with_comparisons(
wl.comp("Franklin D. Roosevelt", "Joe Biden")
.score.shannon_entropy()
.normalize(by="net_diff")
.alias("roosevelt_biden_entropy_normed"),
)
Warning
By default, normalizing forces the comparison to be evaluated eagerly because WordLevel must check that the constant it divides by is nonzero. This behavior can be changed by setting the check_nonzero parameter to False.