Corpus Statistics

Describe the composition and distribution of your corpus

Corpus Statistics summarizes the entire project by text file and speaker. It shows how large the corpus is, how its words are distributed, which sources contribute most, and how each text and speaker relates to the rest of the collection.

14 min
The analytical picture

Use Statistics to understand what your corpus contains

Size

Measure the total number of words, unique word forms, text files, and speakers in the project.

Composition

See the proportion of all corpus words contributed by every text file and speaker.

Coverage

See which speakers appear in a selected text and which texts contain a selected speaker.

Comparison

Sort complete tables to identify large, small, broad, or narrowly represented sources.

Use Corpus Statistics to check corpus shape and balance

Choose this page when

You need project-level checks: corpus size, distribution of contributions by text and speaker, and whether certain sources dominate the dataset.

Prefer Concordance for local questions

When you need exact contextual evidence, jump directly to Concordance and inspect matches.

Pair with Word List for vocabulary questions

Use Word List for normalized token frequencies; use Statistics to ask whether vocabulary patterns are broadly representative.

Use before claims

Before interpreting a specific collocation or topical trend, confirm your corpus is not skewed by a very large file or low speaker coverage.

Statistics always describes the entire corpus

Corpus Statistics does not use the shared Active selection. Selected text files, selected speakers, and custom groups do not narrow this page. Every measure refers to all text files currently imported into the project.

Spoken words count

Exempla counts the word tokens in the prepared project copies of every imported text.

Speaker codes do not count

Codes in square brackets identify turns and attribution, but the codes themselves are excluded from word totals.

Comments do not count

Text inside matched parentheses is masked before counting, including comments that continue across stored lines.

Case is combined

Forms differing only by letter case contribute to the same normalized word form.

Choose how apostrophes behave

Join at apostrophes is the shared project setting. When selected, text on both sides of an apostrophe is treated as one word. When it is not selected, the parts are counted separately. Changing this setting makes existing statistics out of date and requires a refresh.

Generate the statistics step by step

1
Open Corpus Statistics

The page initially says Not generated. Opening it does not begin counting automatically.

2
Confirm Join at apostrophes

Choose the tokenization rule you want applied throughout the calculation.

3
Choose Generate Statistics

Exempla counts in the background. A thin progress indicator and a live elapsed-time display show that work is continuing.

4
Wait for Statistics up to date

The overview, rankings, main tables, and detail panels appear together when the complete result is ready.

Refresh

After generation, the button becomes Refresh. Use it to calculate the same current corpus again.

Update required

This status appears after a relevant corpus change, such as adding or deleting a text, editing or replacing transcript content, renaming a text or speaker code, or changing Join at apostrophes. The main button becomes Refresh Statistics so you can replace the older calculation.

Read the corpus-wide summary first

Total words10,409,395
Text files708
Speakers893
Unique words18,260
Lexical diversity0.18%

Illustration

1Total words

Every counted word occurrence in the complete corpus. Repeated words count each time they occur.

2Text files

The number of corpus texts currently included in the project.

3Speakers

The number of distinct non-empty speaker codes recognized across those texts.

4Unique words

The number of distinct normalized word forms after comments and speaker codes are excluded.

5Lexical diversity

Unique words ÷ Total words × 100. This is the type-token ratio for the complete corpus.

Identify the five largest text files and speakers

The two cards below the Overview rank text files and speakers by word count. Each row shows its name, share of all corpus words, and exact word total. The small bar provides a visual comparison of corpus share.

Top text filesRanked by word count

1Sociolinguistic interview 1120.52%54,117

2Sociolinguistic interview 1200.50%51,763

3Sociolinguistic interview 2090.49%51,437

4Sociolinguistic interview 1560.49%50,981

5Sociolinguistic interview 1780.48%49,630

Top speakersRanked by word count

1Sarah0.27%28,226

2Camille0.22%22,851

3Alexis0.22%22,477

4Morgan0.21%21,930

5Nadia0.20%20,415

Illustration · each bar represents a share of the entire corpus.

Read one row for every text file or speaker

The real example selects Work and Opportunity in Saint-Ouen: 31,261 words, 6.53% of corpus words, with 4 of 111 speakers. Click the screenshot to read the table and speaker breakdown at full size.

Exempla Corpus Statistics

Real Exempla Corpus Statistics text-file table and speaker breakdown in Aurora light mode
Real Exempla Corpus Statistics text-file table and speaker breakdown in Aurora dark mode
Real Exempla Corpus Statistics text-file table and speaker breakdown in high-contrast light mode
Real Exempla Corpus Statistics text-file table and speaker breakdown in high-contrast dark mode
Text Files table

Columns are Text file, Words, Corpus share, Speakers, and Unique words.

Speakers table

Columns are Speaker, Words, Corpus share, Texts, and Unique words.

Texts or Speakers

Use the segmented control above the table. Switching views does not regenerate the statistics.

Default order

Both tables initially place the highest word count first. Equal values use the source name as a stable reference.

What the shared columns mean

Words

The number of counted word occurrences attributed to that text file or speaker.

Corpus share

The row’s Words ÷ Total words × 100. The denominator is always the entire corpus.

Speakers or Texts

For a text, the number of distinct speakers it contains. For a speaker, the number of different texts in which that code appears.

Unique words

The number of distinct normalized word forms found within that individual row’s source.

Select a row to inspect its internal distribution

The right panel changes with the selected main-table row. It provides both coverage and a complete breakdown.

Selected text file

The summary reports its words, corpus share, number of speakers, and speaker coverage. The table then shows each speaker’s words and percentage within that text.

Selected speaker

The summary reports the speaker’s words, corpus share, number of texts, and text coverage. The table then shows how that speaker’s words are distributed across texts.

Share in the detail table

For a text, each speaker’s share uses that text’s word total. For a speaker, each text’s share uses that speaker’s word total.

Move to the counterpart

Double-click a detail row to open the corresponding speaker or text in the opposite main table.

Filter and sort the complete generated tables

Filter text files

In Texts, enter part of a text-file name. Matching is case-insensitive and fixed text, not a regular expression.

Filter speakers

In Speakers, enter part of a speaker code. The same field automatically changes purpose when you switch tabs.

Sort any heading

Choose a column heading to sort the complete current table. Choose it again to reverse the direction.

No recount required

Filtering, sorting, selecting rows, and switching tabs only change the view of the completed calculation.

Use raw totals, proportions, and vocabulary measures for different questions

Words answers “how much?”

Use the raw total when you need the absolute amount contributed by a source.

Corpus share answers “what proportion?”

Use it to compare a source with the complete corpus, especially when totals differ greatly.

Coverage answers “how widely present?”

Use it to see whether one text contains many corpus speakers or one speaker appears across many texts.

Unique words answers “how many forms?”

It is sensitive to source length. Longer samples normally have more opportunities to contain new word forms.

The speaker totals can be lower than the corpus total when words cannot be attributed to a recognized speaker code. Corpus share for a speaker still uses the entire corpus word total, so all speaker shares are not guaranteed to add to 100%.

Export one table as CSV or the complete result as Excel

Export Text Files table as CSV

Saves the complete Texts table in its current filter and sort order. The current Speakers tab does not affect this choice.

Export Speakers table as CSV

Saves the complete Speakers table in its current filter and sort order. The choice is explicit, so “current table” is never ambiguous.

Export complete statistics as Excel workbook

Creates Summary, Text Files, and Speakers worksheets. Summary includes Overview and both top-five rankings. The table worksheets include coverage and all right-panel breakdowns.

Responsive workbook export

A progress window reports the current stage and provides Cancel. Cancelling leaves no incomplete workbook.

The workbook also records the project name, generation date and time, entire-corpus scope, apostrophe rule, and Exempla version. Exempla proposes a date-first filename containing the project name and Corpus Statistics.

Check the status and counting rules

When Corpus Statistics looks unexpected

Nothing is generated

Confirm that the project contains at least one imported corpus text, then choose Generate Statistics.

The status says Update required

A relevant corpus detail or the apostrophe rule changed after generation. Choose Refresh Statistics before relying on the displayed totals, labels, and tables.

The table appears empty

Clear the filter field. A filter can hide every row without removing the generated data.

A speaker count looks wrong

Review square-bracket speaker codes and speaker-code decisions made during import. Statistics counts distinct recognized codes.

A text total differs from the visible transcript

Remember that speaker codes and parenthetical comments are excluded and apostrophe handling can change token counts.

Speaker shares do not total 100%

Some corpus words may not be attributed to a recognized speaker. Every speaker share also uses total corpus words as its denominator.

Unique words seems high or low

Confirm Join at apostrophes and remember that letter case is combined and comments are excluded.

The CSV contains fewer rows

CSV export follows the corresponding filtered table. Clear its filter before exporting all rows.