RunoVerse

Compound Words

Browse compounds from morphological tools and AI annotation — two complementary views. The Tool-detected tab shows 51,606 lemmas verified by EstNLTK, libvoikko, and uralicNLP. The DeepSeek tab shows 183,321 wordforms where AI annotation identified compound morphology.

Morpheme productivity

Which morphemes build the most compounds, split by position. Modifiers are first elements (the qualifier), heads are final elements (what the compound fundamentally is). Counts are distinct compound lemmas containing the morpheme in that position. Click any row to see those compounds.

Top modifier morphemes

first element · qualifier

Top head morphemes

final element · base
Clicking a morpheme switches to “Search by part” and filters the list.
Press / to search, ? for help, J/K to navigate
Loading compounds...
Loading compound index...

Compound Words - Help

What this shows

Every NOUN or ADJ lemma that a morphological analyser (EstNLTK, libvoikko, uralicNLP) decomposes into two or three real parts is surfaced here, together with collection places and glosses. The previous naive substring splitter has been replaced so common false positives (verb forms, diminutives, cross-language leakage) no longer leak through.

How to read it

  • Sidebar is capped at 500 visible rows - refine the search to see more. Sorted by frequency by default; use Sort to switch to place spread or alphabetical.
  • Toggle the Search mode between "compound" (matches the whole lemma or gloss) and "part" (matches either half). In part mode, the matched half is highlighted in the split visualisation.
  • Use Confidence to filter by detection tier. Default is Confirmed (all analyser-verified compounds, no heuristic). Switch to Dual for the strictest slice, or + Heuristic to include the dialectal fallback.
  • Each row shows the compound + its split + a tier badge. Click to open the full detail view on the right: a large poster of the compound with both parts glossed, a map of collection places, and a Related compounds section showing siblings that share the left or right part.
  • Click either half in the detail view to open that part in the Word Explorer.

Confidence tiers

  • dual confirmed - both EstNLTK and libvoikko produced the same parts tuple. Highest precision (~95%).
  • estnltk - EstNLTK morph_analysis.root contains a compound boundary. Covers Estonian lemmas.
  • voikko - libvoikko WORDBASES has two or more non-suffix segments. Covers Finnish lemmas.
  • uralicnlp - uralicNLP Cmp# marker split the lemma (fallback for Finnish dialectal forms voikko does not know).
  • heuristic - last-resort naive splitter for lemmas no analyser recognises. Lower precision, hidden by default.

Caveats

  • Some entries still slip through (e.g. loan words voikko happens to decompose). A small hand-curated blocklist removes the worst offenders.
  • Component glosses are pulled from the DeepSeek gloss index; a small hand-curated table fills in the most common roots.
  • Place data is aggregated across all wordforms of the lemma.

Keyboard shortcuts

/Focus search box ?Toggle this help overlay SShare URL (copy to clipboard) XExport filtered list as CSV RSelect a random compound JNext compound in list KPrevious compound in list EscClose panels / deselect