Word List

Measure vocabulary and word frequency in any corpus scope

Word List counts every word in the Active selection, groups repeated forms together, and reports how often each form occurs. Use it to describe a selection, identify frequent or rare vocabulary, compare differently sized scopes, and open any listed word directly in Concordance.

22 min
The essential distinction

One row represents a word form, not one occurrence

Word token

One occurrence of a word in the analyzed text. In the real example below, est occurs 15,646 times and contributes 15,646 word tokens.

Unique word form

One distinct form after Exempla applies the active letter-case and apostrophe rules. Those 15,646 occurrences of est occupy one Word List row.

Word List

A table with one row per unique word form, plus its exact Frequency and normalized Per 1,000 value.

Choose the Word List scope with Active selection

The scope is the part of the corpus included in the count. Word List uses the same Active selection as Concordance and Sequences. Confirm it before generating because every Frequency, Per 1,000 value, and Overview measure is calculated from that selection.

Active selection1 text group· Ottawa interviews
Entire corpus

Counts every imported corpus text. Use this to establish the project’s complete vocabulary and frequency baseline.

One or more text files

Counts only the selected interviews or documents. The same speaker can contribute words from each selected text.

One or more speakers

Counts words attributed to the selected speaker codes across every text in which those speakers appear.

Text groups

Counts the combined members of one or more custom text groups, such as Ottawa interviews or a particular interview period.

Speaker groups

Counts utterances attributed to the members of one or more custom speaker groups.

Click one item to select it. Use Shift-click for a continuous range. Use Command-click on macOS or Ctrl-click on Windows to combine separate items. Choose the × in Active selection to return to the entire corpus.

Generate a Word List step by step

1
Open Word List

Choose Word List in the main workspace navigation. Opening the page does not begin a potentially large count automatically.

2
Confirm Active selection

Check the summary above the left panel. It must identify the corpus, texts, speakers, or groups you intend to analyze.

3
Set the word-handling rules

Decide whether apostrophe forms remain joined and whether uppercase and lowercase variants remain separate.

4
Choose Generate Word List

Exempla counts the selected language and reports percentage progress and elapsed time while keeping the interface responsive.

5
Verify the completed scope

Read the scope in Overview before interpreting the table. Hover over an abbreviated scope label to see its complete contents.

6
Inspect the table and Overview

Use the rows for individual word frequencies and Overview for measures describing the complete generated selection.

Update Word List

Appears when Active selection or a word-handling rule no longer matches the displayed result. It performs a fresh count using the current choices.

No automatic recount

Changing a scope or rule marks the result for updating but does not interrupt your work or begin a new analysis without your command.

Understand exactly what Exempla counts as a word

Exempla first divides the selected transcript language into countable word tokens. It then normalizes those tokens according to the two visible Word handling controls. These rules determine which occurrences are grouped into the same row.

Join at apostrophes

Off by default. When off, we’re contributes we and re. When on, it contributes the single word we’re. The choice is shared with Concordance, Sequences, and corpus word counts.

Separate letter case

Off by default. When off, Apple, apple, and APPLE are combined in one normalized lowercase row. When on, each visibly different case form receives its own row and frequency.

Numbers and mixed forms

Numbers, alphanumeric forms, and one-character words are counted. Examples include 2026, A1, I, and a.

Internal punctuation

Internal hyphens and periods can remain part of a word, as in mother-in-law, U.S.A, and 3.14. Surrounding punctuation is not part of the word.

Counted spoken text
[MARINA] We’re in Ottawa in 2026.

With apostrophes split and letter case combined, the counted forms are we, re, in, ottawa, in, and 2026.

Excluded annotations
[MARINA] (laughs) We returned.

The speaker code [MARINA] and the comment (laughs) are excluded. Word List has no Include comments option. Only we and returned are counted.

  • Square-bracketed speaker codes identify attribution but never become Word List entries.
  • Correctly paired parenthetical comments are excluded from Word Lists and word totals.
  • Accented letters are preserved. Forms such as près and pres remain different words.
  • Common straight and curly apostrophes are normalized so typography alone does not create separate word forms.
  • Changing either visible Word handling control requires Update Word List because it can change both the rows and every derived measure.

Read the Word List table

This screenshot shows real Word List results from a French corpus, sorted by Frequency. Click it to see the controls and rows at a larger size. In the app, choose any column heading to sort the complete table.

Exempla Word List

Real Exempla Word List results in Aurora light mode
Real Exempla Word List results in Aurora dark mode
Real Exempla Word List results in high-contrast light mode
Real Exempla Word List results in high-contrast dark mode
1Word

The unique word form after Exempla applies the current letter-case and apostrophe rules.

2Frequency

The exact number of times that word occurs in the generated scope. This is a raw count.

3Per 1,000

The frequency expected for every 1,000 analyzed word tokens at the observed rate. This is a normalized frequency, not an additional count.

How Per 1,000 is calculated

Frequency ÷ total words in the generated scope × 1,000. The screenshot shows est with a Frequency of 15,646 and a normalized rate of 29.28 per 1,000 words.

Interpret every Overview measure using its denominator

Overview describes the complete generated scope before any text filter or Minimum frequency threshold is applied. Filtering rows does not recalculate these measures.

OverviewEntire corpus
Corpus534,323Words

Text files36

Speakers111

Vocabulary

Unique words15,696

Hapax legomena6,852

Average word length3.78

Relative to corpus
Share of all words100.0%
Share of all unique words100.0%
Within selection
Hapax share43.7%
Lexical diversity2.9%

14,842words / text

4,814words / speaker

Illustration based on the real Word List Overview screenshot.

Corpus and Vocabulary

Words

All counted word tokens in the generated scope. Repeated occurrences are included every time.

Text files

The number of corpus texts represented in the scope. For a speaker scope, this is the number of texts containing counted utterances from the selected speakers.

Speakers

The number of attributed speaker codes represented in the counted scope. Text before any speaker code can contribute words without creating a speaker.

Unique words

The number of distinct Word List forms after the active normalization rules are applied.

Hapax legomena

Unique word forms occurring exactly once in the generated scope. The count may change when the scope changes.

Average word length

The mean number of characters per counted token. Retained internal apostrophes, hyphens, and periods count as characters.

Relative to the entire corpus

Share of all words

Words in the generated selection ÷ words in the entire corpus, using the same apostrophe and letter-case rules. An entire-corpus list reports 100%.

Share of all unique words

Unique forms present in the selection ÷ unique forms present anywhere in the entire corpus. It measures vocabulary coverage, not the selection’s share of token occurrences.

Within the generated selection

Hapax share

Hapax legomena ÷ unique words × 100. It answers what percentage of the selection’s distinct forms occur only once.

Lexical diversity

Unique words ÷ all words × 100. This is Type–Token Ratio expressed as a percentage.

Words per text

Total words ÷ represented text files. It is an arithmetic average and does not mean every text has that length.

Words per speaker

Total words in the generated scope ÷ represented speakers. It is an arithmetic average, not each speaker’s individual contribution, and can hide large differences between speakers.

Filter and sort the completed Word List

Filter results

Enter one or more characters to keep words containing that text anywhere in the form. Matching is not case-sensitive, even when Separate letter case produced distinct rows.

Minimum

Enter the smallest raw Frequency to display. Minimum 5 keeps only forms occurring at least five times. It does not use the Per 1,000 value.

Clear the filters

Use the × in Filter results and return Minimum to 1. The complete generated rows reappear immediately without a new Word List generation.

  • Filtering affects the detailed table and the Word List navigator in the left panel.
  • Filtering does not change the generated scope, token counts, or Overview measures.
  • Export includes the rows currently displayed after Filter results and Minimum are applied.
  • Filter and Minimum can remain active for the next generation, so check them if a new table seems unexpectedly short.

Sort from any table heading

Word

The first click on a newly selected Word heading sorts A–Z. Click it again for Z–A.

Frequency

The first click sorts from the highest raw count to the lowest. Click again for the lowest count first.

Per 1,000

The first click sorts from the highest normalized rate to the lowest. In a single Word List, this order matches Frequency because every row has the same denominator.

Equal values

Words with equal numeric values use alphabetical order as a stable tiebreaker.

Double-click a word to inspect every occurrence

Double-click any row in the detailed Word List. Exempla opens Concordance and runs an Exact word search for that form using the current Active selection. The Concordance result supplies speaker, left and right context, source text, and transcript line.

Double-click a Word List row

Choose the word form whose uses you want to inspect.

Exact Concordance begins

Exempla searches using the Word List letter-case and apostrophe policy.

Read each occurrence

Compare contexts, open the transcript, collect evidence, or export the results.

Compare two Word Lists using normalized frequency

Choose Compare scopes… to compare two independent selections. Scope A and Scope B can each be the entire corpus, one text, one speaker, one text group, or one speaker group. The two choices must be different. They do not replace Active selection.

The real example compares Speaker 4 Augustin (A) with Speaker 037 Louise (B). The table and Overview below are from the same comparison; click the screenshot to read them at full size.

Exempla Word List scope comparison

Real Exempla Word List scope comparison results in Aurora light mode
Real Exempla Word List scope comparison results in Aurora dark mode
Real Exempla Word List scope comparison results in high-contrast light mode
Real Exempla Word List scope comparison results in high-contrast dark mode
AA frequency and A per 1,000

The raw count and normalized rate within Scope A.

BB frequency and B per 1,000

The raw count and normalized rate within Scope B.

±A − B

A per 1,000 minus B per 1,000. Positive values favour A; negative values favour B; zero means equal normalized rates.

For c, A has 55 occurrences (46.03 per 1,000) and B has 124 (26.09 per 1,000), yet A is higher on the normalized scale: A − B is +19.93. Exempla subtracts the full-precision rates before rounding the result, so subtracting the displayed rates may differ by 0.01.

  • The table contains the combined vocabulary of both scopes. A word absent from one scope receives zero for that scope.
  • Clicking A − B orders by the size of the difference regardless of sign. This brings the largest contrasts in either direction together.
  • In comparison mode, Minimum uses the combined raw frequencies from A and B.
  • The A and B key remains visible above the table so the direction of every difference stays clear.
  • Overview changes to side-by-side A and B measures for words, unique words, average length, corpus shares, hapax share, and lexical diversity.

Export the Word List or comparison

Choose Export after generation. Export is unavailable while counting or when the table has no displayed rows.

Single Word List

The CSV contains Word, Frequency, and Per 1,000 for every currently displayed row.

Scope comparison

The CSV contains Word, A frequency, A per 1,000, B frequency, B per 1,000, and A − B. The headers include the complete A and B scope labels.

Filters are respected

If Filter results or Minimum hides rows, those rows are not included. Reset both controls before exporting the complete generated table.

Suggested filename

Exempla proposes a date-first name containing the project name and either Word List or Word List Comparison. If the project name would make the filename unsafe, it is omitted instead of being cut.

Text encoding

The file uses UTF-8 with an encoding signature so accented characters open reliably in spreadsheet software on macOS and Windows.

Large export

An export of at least 5,000 rows uses a responsive progress window with Cancel. Cancelling does not leave an incomplete CSV in place.

Check scope, rules, and view controls in order

When a Word List does not look as expected

Generate Word List is unavailable

Confirm that the project contains at least one imported corpus text.

The button says Update Word List

Active selection or a Word handling rule changed after the visible table was generated. Choose Update Word List to recount the current choices.

No words were found

Confirm that Active selection contains spoken text and that speaker codes and parentheses are correctly paired. Codes and comments are intentionally excluded.

The table says no results match the filters

Clear Filter results and return Minimum to 1. The generated list exists, but the current view controls are hiding every row.

Uppercase and lowercase forms are unexpectedly combined

Turn on Separate letter case, then choose Update Word List. Filtering alone cannot separate forms that were combined during generation.

An apostrophe form is missing

Check Join at apostrophes. With it off, we’re appears under we and re, not as one joined row.

A frequency differs from Concordance

Verify that both tools use the same Active selection, letter-case rule, apostrophe rule, and comment policy. Update the Word List if its button indicates an outdated result.

Per 1,000 looks smaller than Frequency

The two columns use different units. Frequency is a count; Per 1,000 is a rate based on the scope’s total words.

The comparison direction seems reversed

Read the A and B key. A − B is positive when the rate is higher in A and negative when it is higher in B.

Export contains fewer rows than expected

Export follows the displayed table. Reset Filter results to empty and Minimum to 1 before exporting the complete generated list.