Comparisons¶
What is a comparison?¶
A comparison quantifies the token-level differences between two corpora. It is constructed by executing a plan for how to compare the corpora, including how to measure their differences, which tokens to include, and more. If you are familiar with Polars, this is similar to how expressions are plans for how to transform and analyze columns in a dataframe.
The core of every comparison is a scoring method for quantifying the token-level differences. WordLevel implements a full suite of measures, including basic frequency differences, entropy-based measures like the Kullback-Leibler and Jensen-Shannon divergences, and other advanced measures.
Reference and comparison corpora¶
Each comparison has two corpora: a reference corpus and a comparison corpus. The reference corpus can be considered the baseline from which the comparison corpus deviates. It is important to be mindful of which corpus is which, since it can affect the interpretation of the results.
For example, say we have two corpora: the speeches of President Franklin D. Roosevelt and the speeches of President Joe Biden. Suppose Roosevelt's speeches are the reference corpus and Biden's are the comparison corpus, and we simply want to measure the difference in how often words are used. Since Roosevelt is our reference, this means we are calculating \(f^{(\text{Biden})}_\tau - f^{(\text{Roosevelt})}_\tau\), where \(f^{(C)}_\tau\) is the frequency of token \(\tau\) in corpus \(C\). If the difference is positive, then it means that the token was used more by Biden. Similarly, if it is negative, then it means it was used more by Roosevelt.
Some scoring methods, like the Jensen-Shannon divergence or rank-turbulence divergence, are symmetric by default, meaning that the token-level difference is the same regardless of which corpus is the reference and which is the comparison. However, it can still be important to track which corpus is which, particularly for interpreting a word shift graph.