Corpora & Catalogs¶
What is a corpus?¶
A corpus is a collection of documents, where each document contains an ordered sequence of tokens. A token is the smallest unit of text that we care about. Often, tokens are individual words separated by whitespace. For example, we can break the sentence "WordLevel is a Python package" into the tokens ("WordLevel", "is", "a", "Python", "package"). In other cases, "tokens" may refer to n-grams (n consecutive words) or subword pieces. We refer to the full set of unique tokens as the vocabulary.
When we have more than one corpus, we have corpora. For example, we may say that all of the social media posts made on June 1st are one corpus (where each post is a document), and all of the posts made on June 2nd are another corpus. The June 1st corpus and June 2nd corpus together are our corpora.
Warning
A corpus is meant to be a relatively large collection of text containing at least several thousand tokens. While what constitutes a "document" or "corpus" may differ depending on the context, a single short document with only tens or hundreds of tokens is not a corpus. WordLevel's methods are only scientifically suited for corpora that have large, stable frequency distributions across the vocabulary, and they are not recommended for comparing short documents individually.
What is a catalog?¶
A catalog is WordLevel's interface for working with multiple corpora. It is the workbench for transforming, manipulating, and comparing corpora in the same way that a dataframe is a workbench for manipulating columns.
To work with corpora and catalogs, see the cookbook.