One spelling, more than one lemma. Every word in the corpus has been read twice by machine — once by the morphological pipeline, once by the DeepSeek analysis — and these are the spellings where at least one of those readings names more than one headword. Sometimes the word really is ambiguous. Often the two readings simply disagree, and the disagreement is a lemmatisation error rather than a property of the language.
What the page shows
A wordform is the exact spelling as it stands in a poem; a lemma is the headword it is filed under. Each card is one spelling, and each row beneath it is one lemma that a reading assigned it to, with the number of occurrences that reading gave to that lemma.
The two readings
The morphological reading is the project's own analyser chain — EstNLTK with a dialectal fallback chain for Estonian, Omorfi and Voikko with Stanza and a dialect dictionary for Finnish. The DeepSeek reading is the AI analysis run over the poems. They work independently and see the same tokens.
Neither is a human gold standard. Where they split a spelling between lemmas, the cause may be real morphological ambiguity — a form that is genuinely both a noun case and a verb form — or it may be a lemmatisation error. The page does not decide that for you; it lays out the evidence. Reading the verses is usually the quickest way to settle it.
How to read a card
The heading is the spelling, followed by how often it occurs in the corpus. Follow it to see those verses in the concordance.
On each row, the green figure is how many occurrences the morphological reading filed under that lemma; the purple figure is how many the DeepSeek reading did.
A row with only a purple figure is a lemma only the AI proposes; only green means only the analyser chain does.
A card is flagged when the two readings put the spelling under different lemmas first.
Sort options
Most contested — by the occurrences the leading lemma does not account for, so a spelling split nearly evenly between two lemmas comes first.
Readings disagree — keeps only the spellings whose two readings lead with different lemmas, commonest first.
Most frequent — by how often the spelling occurs.
Most candidate lemmas — by how many lemmas the two readings name between them.
Why it matters
Runosong language is archaic, dialectal, and full of syncopated and truncated forms. Which lemma a word belongs to decides every frequency count and every semantic analysis built on top of it, so the places where two readings of the same line disagree are the places where those counts are least trustworthy.
morphological readingDeepSeek readingBoth numbers count occurrences of that spelling, not of the lemma.