Measure vocabulary and word frequency in any corpus scope
Word List counts every word in the Active selection, groups repeated forms together, and reports how often each form occurs. Use it to describe a selection, identify frequent or rare vocabulary, compare differently sized scopes, and open any listed word directly in Concordance.
One row represents a word form, not one occurrence
One occurrence of a word in the analyzed text. In the real example below, est occurs 15,646 times and contributes 15,646 word tokens.
One distinct form after Exempla applies the active letter-case and apostrophe rules. Those 15,646 occurrences of est occupy one Word List row.
A table with one row per unique word form, plus its exact Frequency and normalized Per 1,000 value.
Choose the Word List scope with Active selection
The scope is the part of the corpus included in the count. Word List uses the same Active selection as Concordance and Sequences. Confirm it before generating because every Frequency, Per 1,000 value, and Overview measure is calculated from that selection.
Counts every imported corpus text. Use this to establish the project’s complete vocabulary and frequency baseline.
Counts only the selected interviews or documents. The same speaker can contribute words from each selected text.
Counts words attributed to the selected speaker codes across every text in which those speakers appear.
Counts the combined members of one or more custom text groups, such as Ottawa interviews or a particular interview period.
Counts utterances attributed to the members of one or more custom speaker groups.
Click one item to select it. Use Shift-click for a continuous range. Use Command-click on macOS or Ctrl-click on Windows to combine separate items. Choose the × in Active selection to return to the entire corpus.
Generate a Word List step by step
Choose Word List in the main workspace navigation. Opening the page does not begin a potentially large count automatically.
Check the summary above the left panel. It must identify the corpus, texts, speakers, or groups you intend to analyze.
Decide whether apostrophe forms remain joined and whether uppercase and lowercase variants remain separate.
Exempla counts the selected language and reports percentage progress and elapsed time while keeping the interface responsive.
Read the scope in Overview before interpreting the table. Hover over an abbreviated scope label to see its complete contents.
Use the rows for individual word frequencies and Overview for measures describing the complete generated selection.
Appears when Active selection or a word-handling rule no longer matches the displayed result. It performs a fresh count using the current choices.
Changing a scope or rule marks the result for updating but does not interrupt your work or begin a new analysis without your command.
Understand exactly what Exempla counts as a word
Exempla first divides the selected transcript language into countable word tokens. It then normalizes those tokens according to the two visible Word handling controls. These rules determine which occurrences are grouped into the same row.
Off by default. When off, we’re contributes we and re. When on, it contributes the single word we’re. The choice is shared with Concordance, Sequences, and corpus word counts.
Off by default. When off, Apple, apple, and APPLE are combined in one normalized lowercase row. When on, each visibly different case form receives its own row and frequency.
Numbers, alphanumeric forms, and one-character words are counted. Examples include 2026, A1, I, and a.
Internal hyphens and periods can remain part of a word, as in mother-in-law, U.S.A, and 3.14. Surrounding punctuation is not part of the word.
[MARINA] We’re in Ottawa in 2026.With apostrophes split and letter case combined, the counted forms are we, re, in, ottawa, in, and 2026.
[MARINA] (laughs) We returned.The speaker code [MARINA] and the comment (laughs) are excluded. Word List has no Include comments option. Only we and returned are counted.
- Square-bracketed speaker codes identify attribution but never become Word List entries.
- Correctly paired parenthetical comments are excluded from Word Lists and word totals.
- Accented letters are preserved. Forms such as près and pres remain different words.
- Common straight and curly apostrophes are normalized so typography alone does not create separate word forms.
- Changing either visible Word handling control requires Update Word List because it can change both the rows and every derived measure.
Read the Word List table
This screenshot shows real Word List results from a French corpus, sorted by Frequency. Click it to see the controls and rows at a larger size. In the app, choose any column heading to sort the complete table.
The unique word form after Exempla applies the current letter-case and apostrophe rules.
The exact number of times that word occurs in the generated scope. This is a raw count.
The frequency expected for every 1,000 analyzed word tokens at the observed rate. This is a normalized frequency, not an additional count.
Frequency ÷ total words in the generated scope × 1,000. The screenshot shows est with a Frequency of 15,646 and a normalized rate of 29.28 per 1,000 words.
Interpret every Overview measure using its denominator
Overview describes the complete generated scope before any text filter or Minimum frequency threshold is applied. Filtering rows does not recalculate these measures.
Text files36
Speakers111
Unique words15,696
Hapax legomena6,852
Average word length3.78
14,842words / text
4,814words / speaker
Illustration based on the real Word List Overview screenshot.
Corpus and Vocabulary
All counted word tokens in the generated scope. Repeated occurrences are included every time.
The number of corpus texts represented in the scope. For a speaker scope, this is the number of texts containing counted utterances from the selected speakers.
The number of attributed speaker codes represented in the counted scope. Text before any speaker code can contribute words without creating a speaker.
The number of distinct Word List forms after the active normalization rules are applied.
Unique word forms occurring exactly once in the generated scope. The count may change when the scope changes.
The mean number of characters per counted token. Retained internal apostrophes, hyphens, and periods count as characters.
Relative to the entire corpus
Words in the generated selection ÷ words in the entire corpus, using the same apostrophe and letter-case rules. An entire-corpus list reports 100%.
Unique forms present in the selection ÷ unique forms present anywhere in the entire corpus. It measures vocabulary coverage, not the selection’s share of token occurrences.
Within the generated selection
Hapax legomena ÷ unique words × 100. It answers what percentage of the selection’s distinct forms occur only once.
Unique words ÷ all words × 100. This is Type–Token Ratio expressed as a percentage.
Total words ÷ represented text files. It is an arithmetic average and does not mean every text has that length.
Total words in the generated scope ÷ represented speakers. It is an arithmetic average, not each speaker’s individual contribution, and can hide large differences between speakers.
Filter and sort the completed Word List
Enter one or more characters to keep words containing that text anywhere in the form. Matching is not case-sensitive, even when Separate letter case produced distinct rows.
Enter the smallest raw Frequency to display. Minimum 5 keeps only forms occurring at least five times. It does not use the Per 1,000 value.
Use the × in Filter results and return Minimum to 1. The complete generated rows reappear immediately without a new Word List generation.
- Filtering affects the detailed table and the Word List navigator in the left panel.
- Filtering does not change the generated scope, token counts, or Overview measures.
- Export includes the rows currently displayed after Filter results and Minimum are applied.
- Filter and Minimum can remain active for the next generation, so check them if a new table seems unexpectedly short.
Sort from any table heading
The first click on a newly selected Word heading sorts A–Z. Click it again for Z–A.
The first click sorts from the highest raw count to the lowest. Click again for the lowest count first.
The first click sorts from the highest normalized rate to the lowest. In a single Word List, this order matches Frequency because every row has the same denominator.
Words with equal numeric values use alphabetical order as a stable tiebreaker.
Double-click a word to inspect every occurrence
Double-click any row in the detailed Word List. Exempla opens Concordance and runs an Exact word search for that form using the current Active selection. The Concordance result supplies speaker, left and right context, source text, and transcript line.
Choose the word form whose uses you want to inspect.
Exempla searches using the Word List letter-case and apostrophe policy.
Compare contexts, open the transcript, collect evidence, or export the results.
Compare two Word Lists using normalized frequency
Choose Compare scopes… to compare two independent selections. Scope A and Scope B can each be the entire corpus, one text, one speaker, one text group, or one speaker group. The two choices must be different. They do not replace Active selection.
The real example compares Speaker 4 Augustin (A) with Speaker 037 Louise (B). The table and Overview below are from the same comparison; click the screenshot to read them at full size.
The raw count and normalized rate within Scope A.
The raw count and normalized rate within Scope B.
A per 1,000 minus B per 1,000. Positive values favour A; negative values favour B; zero means equal normalized rates.
For c, A has 55 occurrences (46.03 per 1,000) and B has 124 (26.09 per 1,000), yet A is higher on the normalized scale: A − B is +19.93. Exempla subtracts the full-precision rates before rounding the result, so subtracting the displayed rates may differ by 0.01.
- The table contains the combined vocabulary of both scopes. A word absent from one scope receives zero for that scope.
- Clicking A − B orders by the size of the difference regardless of sign. This brings the largest contrasts in either direction together.
- In comparison mode, Minimum uses the combined raw frequencies from A and B.
- The A and B key remains visible above the table so the direction of every difference stays clear.
- Overview changes to side-by-side A and B measures for words, unique words, average length, corpus shares, hapax share, and lexical diversity.
Export the Word List or comparison
Choose Export after generation. Export is unavailable while counting or when the table has no displayed rows.
The CSV contains Word, Frequency, and Per 1,000 for every currently displayed row.
The CSV contains Word, A frequency, A per 1,000, B frequency, B per 1,000, and A − B. The headers include the complete A and B scope labels.
If Filter results or Minimum hides rows, those rows are not included. Reset both controls before exporting the complete generated table.
Exempla proposes a date-first name containing the project name and either Word List or Word List Comparison. If the project name would make the filename unsafe, it is omitted instead of being cut.
The file uses UTF-8 with an encoding signature so accented characters open reliably in spreadsheet software on macOS and Windows.
An export of at least 5,000 rows uses a responsive progress window with Cancel. Cancelling does not leave an incomplete CSV in place.
When a Word List does not look as expected
Confirm that the project contains at least one imported corpus text.
Active selection or a Word handling rule changed after the visible table was generated. Choose Update Word List to recount the current choices.
Confirm that Active selection contains spoken text and that speaker codes and parentheses are correctly paired. Codes and comments are intentionally excluded.
Clear Filter results and return Minimum to 1. The generated list exists, but the current view controls are hiding every row.
Turn on Separate letter case, then choose Update Word List. Filtering alone cannot separate forms that were combined during generation.
Check Join at apostrophes. With it off, we’re appears under we and re, not as one joined row.
Verify that both tools use the same Active selection, letter-case rule, apostrophe rule, and comment policy. Update the Word List if its button indicates an outdated result.
The two columns use different units. Frequency is a count; Per 1,000 is a rate based on the scope’s total words.
Read the A and B key. A − B is positive when the rate is higher in A and negative when it is higher in B.
Export follows the displayed table. Reset Filter results to empty and Minimum to 1 before exporting the complete generated list.
Was this chapter helpful?
Your response will help improve the guide.