What this shows
Every NOUN or ADJ lemma that a morphological analyser (EstNLTK, libvoikko, uralicNLP) decomposes into two or three real parts is surfaced here, together with collection places and glosses. The previous naive substring splitter has been replaced so common false positives (verb forms, diminutives, cross-language leakage) no longer leak through.
How to read it
- Sidebar is capped at 500 visible rows - refine the search to see more. Sorted by frequency by default; use Sort to switch to place spread or alphabetical.
- Toggle the Search mode between "compound" (matches the whole lemma or gloss) and "part" (matches either half). In part mode, the matched half is highlighted in the split visualisation.
- Use Confidence to filter by detection tier. Default is Confirmed (all analyser-verified compounds, no heuristic). Switch to Dual for the strictest slice, or + Heuristic to include the dialectal fallback.
- Each row shows the compound + its split + a tier badge. Click to open the full detail view on the right: a large poster of the compound with both parts glossed, a map of collection places, and a Related compounds section showing siblings that share the left or right part.
- Click either half in the detail view to open that part in the Word Explorer.
Confidence tiers
- dual confirmed - both EstNLTK and libvoikko produced the same parts tuple. Highest precision (~95%).
- estnltk - EstNLTK morph_analysis.root contains a compound boundary. Covers Estonian lemmas.
- voikko - libvoikko WORDBASES has two or more non-suffix segments. Covers Finnish lemmas.
- uralicnlp - uralicNLP Cmp# marker split the lemma (fallback for Finnish dialectal forms voikko does not know).
- heuristic - last-resort naive splitter for lemmas no analyser recognises. Lower precision, hidden by default.
Caveats
- Some entries still slip through (e.g. loan words voikko happens to decompose). A small hand-curated blocklist removes the worst offenders.
- Component glosses are pulled from the DeepSeek gloss index; a small hand-curated table fills in the most common roots.
- Place data is aggregated across all wordforms of the lemma.
What this shows
Wordforms where DeepSeek AI morphological analysis mentioned "compound" in its comment. This provides broader coverage than tool detection but includes some derivations and edge cases.
How to read it
- Frequency = corpus attestation count. Decomposition shows AI-parsed compound parts (when available in the comment). Tool overlap indicates independent confirmation by morphological analysers.
- Search across wordforms, comments, translations, and parsed parts. Use filters to narrow by language, overlap status, or decomposition availability.
- Sidebar is capped at 500 rows. Use search and filters to navigate the full set of entries.
Wordforms vs lemmas
This tab centers on wordforms because lemma assignments may be imprecise for dialectal forms. The lemma shown in the detail pane is for context - not a definitive claim.
Caveats
- DeepSeek annotation may classify some derivations (e.g., -line, -ton suffixes) as compounds. Use the "Tool-confirmed" filter to focus on high-confidence entries.
- Parsed decomposition is extracted from free-text AI comments via pattern matching - about 75K of 183K entries have parseable parts.