Constructing Catalogs¶
Creating a corpus¶
A corpus is represented as a bag of words. One way to encode this is as a dictionary mapping each token to its frequency:
corpus = {
"the": 100, # The token "the" appears 100 times across all the documents in the corpus
"and": 88,
"sad": 24,
"happy": 35,
# And so on...
}
Similarly, we can represent a corpus as a set of rows and columns, where there is one row per token, one column indicating the token itself, and one column indicating its frequency.
| token | frequency |
|---|---|
| the | 100 |
| and | 88 |
| sad | 24 |
| happy | 35 |
In Python, we can represent these rows and columns as a dataframe. For example, using the package Polars, our corpus can be encoded as a dataframe.
import polars as pl
corpus = pl.DataFrame({
"token": ["the", "and", "sad", "happy"],
"frequency": [100, 88, 24, 35]
})
Tip
Currently, WordLevel does not have any built-in functionality for tokenizing documents and counting token frequencies. For tools that handle tokenization, consider using other packages like NLTK, spaCy, or scikit-learn.
Creating a Catalog¶
Say that we have two corpora, one about animals in general, and one about cats specifically. Their vocabularies will have some overlaps and some differences.
import polars as pl
import wordlevel as wl
animals = {"the": 92, "and": 74, "duck": 20, "dog": 14, "cat": 11, "lion": 4}
cats = {"the": 90, "and": 78, "cat": 55, "lion": 38, "tabby": 20, "tuxedo": 19}
If these are our corpora, we can create a Catalog in several ways, depending on which is the most
convenient.
From multiple dictionaries¶
We can create a Catalog by passing both corpora (represented as token-to-frequency dictionaries)
and giving them names. Note, corpus names must be listed in the same order as the dictionaries. If not provided, they will default to the generic names corpus1 and corpus2.
From a single dictionary¶
We can also pass a single dictionary which maps each corpus name to its token-to-frequency dictionary explicitly:
From a Polars dataframe¶
If we already have our corpora formatted as a dataframe, we can create a Catalog directly
from it. First, let's create a dataframe for our example data.
# To create a dataframe, we need ordered columns, filling in 0s if a token doesn't appear in a corpus
tokens = sorted(list(set(animals) | set(cats)))
animal_freqs = [animals[token] if token in animals else 0 for token in tokens]
cat_freqs = [cats[token] if token in cats else 0 for token in tokens]
df_corpora = pl.DataFrame({"token": tokens, "animals": animal_freqs, "cats": cat_freqs})
To create a Catalog, we need to provide the column indicating the names of the tokens, the
token_col, and the names of the columns indicating frequencies of tokens, the corpora. Note,
with a list of dictionaries, we use the corpora parameter to name the corpora—here, we use it to select them.
From a Narwhals-supported dataframe¶
WordLevel can create a Catalog from any Narwhals-supported dataframe, including pandas, DuckDB,
PySpark, and others.
import pandas as pd
df_pandas = pd.DataFrame({"token": tokens, "animals": animal_freqs, "cats": cat_freqs})
cl = wl.Catalog(df_pandas, token_col="token", corpora=["animals", "cats"])
Note
Polars is the internal engine for WordLevel. This means that even if you provide a non-Polars dataframe, any dataframe you receive back from WordLevel will be a Polars dataframe. See the Polars documentation to learn more about how to manipulate such dataframes.
If you pass a lazy non-Polars dataframe, it will be eagerly collected and materialized in memory when the Catalog is constructed.
Working with a Catalog¶
All of the above constructors create the same Catalog object. Regardless of which one we use, the
corpora are now represented as a Polars dataframe under the hood. We can inspect that dataframe
directly through the df attribute.
┌────────┬─────────┬──────┬────────────────────┬─────────────────┐
│ token ┆ animals ┆ cats ┆ animals__wl_normed ┆ cats__wl_normed │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ i64 ┆ i64 ┆ f64 ┆ f64 │
╞════════╪═════════╪══════╪════════════════════╪═════════════════╡
│ and ┆ 74 ┆ 78 ┆ 0.344186 ┆ 0.26 │
│ cat ┆ 11 ┆ 55 ┆ 0.051163 ┆ 0.183333 │
│ dog ┆ 14 ┆ 0 ┆ 0.065116 ┆ 0.0 │
│ duck ┆ 20 ┆ 0 ┆ 0.093023 ┆ 0.0 │
│ lion ┆ 4 ┆ 38 ┆ 0.018605 ┆ 0.126667 │
│ tabby ┆ 0 ┆ 20 ┆ 0.0 ┆ 0.066667 │
│ the ┆ 92 ┆ 90 ┆ 0.427907 ┆ 0.3 │
│ tuxedo ┆ 0 ┆ 19 ┆ 0.0 ┆ 0.063333 │
└────────┴─────────┴──────┴────────────────────┴─────────────────┘
The output shows that there is one row per token, and there is one column for each corpus: one for the animals corpus and one for the cats corpus.
There are two new columns with the suffix __wl_normed. These columns are automatically created
by the Catalog and indicate the normalized frequency of each token in each corpus: the
proportion between 0 and 1 derived by dividing a token's frequency by the total sum of
frequencies within a corpus.
If a token does not appear in a corpus, it has a frequency of zero in that corpus. For example,
"duck" is not in the cats corpus, but we see that in its row, it has a frequency of zero for that
corpus, even if we did not explicitly provide that information in the dictionaries. In other words,
the Catalog considers the union of vocabularies across all of the corpora, filling in zero
frequencies as needed.