Describe the composition and distribution of your corpus
Corpus Statistics summarizes the entire project by text file and speaker. It shows how large the corpus is, how its words are distributed, which sources contribute most, and how each text and speaker relates to the rest of the collection.
Use Statistics to understand what your corpus contains
Measure the total number of words, unique word forms, text files, and speakers in the project.
See the proportion of all corpus words contributed by every text file and speaker.
See which speakers appear in a selected text and which texts contain a selected speaker.
Sort complete tables to identify large, small, broad, or narrowly represented sources.
Use Corpus Statistics to check corpus shape and balance
You need project-level checks: corpus size, distribution of contributions by text and speaker, and whether certain sources dominate the dataset.
When you need exact contextual evidence, jump directly to Concordance and inspect matches.
Use Word List for normalized token frequencies; use Statistics to ask whether vocabulary patterns are broadly representative.
Before interpreting a specific collocation or topical trend, confirm your corpus is not skewed by a very large file or low speaker coverage.
Statistics always describes the entire corpus
Corpus Statistics does not use the shared Active selection. Selected text files, selected speakers, and custom groups do not narrow this page. Every measure refers to all text files currently imported into the project.
Exempla counts the word tokens in the prepared project copies of every imported text.
Codes in square brackets identify turns and attribution, but the codes themselves are excluded from word totals.
Text inside matched parentheses is masked before counting, including comments that continue across stored lines.
Forms differing only by letter case contribute to the same normalized word form.
Choose how apostrophes behave
Join at apostrophes is the shared project setting. When selected, text on both sides of an apostrophe is treated as one word. When it is not selected, the parts are counted separately. Changing this setting makes existing statistics out of date and requires a refresh.
Generate the statistics step by step
The page initially says Not generated. Opening it does not begin counting automatically.
Choose the tokenization rule you want applied throughout the calculation.
Exempla counts in the background. A thin progress indicator and a live elapsed-time display show that work is continuing.
The overview, rankings, main tables, and detail panels appear together when the complete result is ready.
After generation, the button becomes Refresh. Use it to calculate the same current corpus again.
This status appears after a relevant corpus change, such as adding or deleting a text, editing or replacing transcript content, renaming a text or speaker code, or changing Join at apostrophes. The main button becomes Refresh Statistics so you can replace the older calculation.
Read the corpus-wide summary first
Illustration
Every counted word occurrence in the complete corpus. Repeated words count each time they occur.
The number of corpus texts currently included in the project.
The number of distinct non-empty speaker codes recognized across those texts.
The number of distinct normalized word forms after comments and speaker codes are excluded.
Unique words ÷ Total words × 100. This is the type-token ratio for the complete corpus.
Identify the five largest text files and speakers
The two cards below the Overview rank text files and speakers by word count. Each row shows its name, share of all corpus words, and exact word total. The small bar provides a visual comparison of corpus share.
1Sociolinguistic interview 1120.52%54,117
2Sociolinguistic interview 1200.50%51,763
3Sociolinguistic interview 2090.49%51,437
4Sociolinguistic interview 1560.49%50,981
5Sociolinguistic interview 1780.48%49,630
1Sarah0.27%28,226
2Camille0.22%22,851
3Alexis0.22%22,477
4Morgan0.21%21,930
5Nadia0.20%20,415
Illustration · each bar represents a share of the entire corpus.
Read one row for every text file or speaker
The real example selects Work and Opportunity in Saint-Ouen: 31,261 words, 6.53% of corpus words, with 4 of 111 speakers. Click the screenshot to read the table and speaker breakdown at full size.
Columns are Text file, Words, Corpus share, Speakers, and Unique words.
Columns are Speaker, Words, Corpus share, Texts, and Unique words.
Use the segmented control above the table. Switching views does not regenerate the statistics.
Both tables initially place the highest word count first. Equal values use the source name as a stable reference.
What the shared columns mean
The number of counted word occurrences attributed to that text file or speaker.
The row’s Words ÷ Total words × 100. The denominator is always the entire corpus.
For a text, the number of distinct speakers it contains. For a speaker, the number of different texts in which that code appears.
The number of distinct normalized word forms found within that individual row’s source.
Select a row to inspect its internal distribution
The right panel changes with the selected main-table row. It provides both coverage and a complete breakdown.
The summary reports its words, corpus share, number of speakers, and speaker coverage. The table then shows each speaker’s words and percentage within that text.
The summary reports the speaker’s words, corpus share, number of texts, and text coverage. The table then shows how that speaker’s words are distributed across texts.
For a text, each speaker’s share uses that text’s word total. For a speaker, each text’s share uses that speaker’s word total.
Double-click a detail row to open the corresponding speaker or text in the opposite main table.
Filter and sort the complete generated tables
In Texts, enter part of a text-file name. Matching is case-insensitive and fixed text, not a regular expression.
In Speakers, enter part of a speaker code. The same field automatically changes purpose when you switch tabs.
Choose a column heading to sort the complete current table. Choose it again to reverse the direction.
Filtering, sorting, selecting rows, and switching tabs only change the view of the completed calculation.
Use raw totals, proportions, and vocabulary measures for different questions
Use the raw total when you need the absolute amount contributed by a source.
Use it to compare a source with the complete corpus, especially when totals differ greatly.
Use it to see whether one text contains many corpus speakers or one speaker appears across many texts.
It is sensitive to source length. Longer samples normally have more opportunities to contain new word forms.
The speaker totals can be lower than the corpus total when words cannot be attributed to a recognized speaker code. Corpus share for a speaker still uses the entire corpus word total, so all speaker shares are not guaranteed to add to 100%.
Export one table as CSV or the complete result as Excel
Saves the complete Texts table in its current filter and sort order. The current Speakers tab does not affect this choice.
Saves the complete Speakers table in its current filter and sort order. The choice is explicit, so “current table” is never ambiguous.
Creates Summary, Text Files, and Speakers worksheets. Summary includes Overview and both top-five rankings. The table worksheets include coverage and all right-panel breakdowns.
A progress window reports the current stage and provides Cancel. Cancelling leaves no incomplete workbook.
The workbook also records the project name, generation date and time, entire-corpus scope, apostrophe rule, and Exempla version. Exempla proposes a date-first filename containing the project name and Corpus Statistics.
When Corpus Statistics looks unexpected
Confirm that the project contains at least one imported corpus text, then choose Generate Statistics.
A relevant corpus detail or the apostrophe rule changed after generation. Choose Refresh Statistics before relying on the displayed totals, labels, and tables.
Clear the filter field. A filter can hide every row without removing the generated data.
Review square-bracket speaker codes and speaker-code decisions made during import. Statistics counts distinct recognized codes.
Remember that speaker codes and parenthetical comments are excluded and apostrophe handling can change token counts.
Some corpus words may not be attributed to a recognized speaker. Every speaker share also uses total corpus words as its denominator.
Confirm Join at apostrophes and remember that letter case is combined and comments are excluded.
CSV export follows the corresponding filtered table. Clear its filter before exporting all rows.
Was this chapter helpful?
Your response will help improve the guide.