Explore word co-occurrence patterns across more than 4.3 million verse lines — discover formulaic pairs and poetic associations.
What this shows
Co-occurrence analysis of words in Finnic runosong verses. Two words are "collocates" if they frequently appear in the same verse line.
Association strength is measured by PMI (Pointwise Mutual Information) — high PMI means the pair appears together far more often than chance predicts.
How to read the results
PMI score — association strength (higher = stronger). Naturally downweights common function words.
Co-occurrence count — verses containing both words.
Mutual badge — both words list each other as top collocates (bidirectional association).
Sorting
By PMI — strongest statistical associations (rare but specific pairs)
By count — most frequent RunoVerse verse-level co-occurrences
FILTER Metrics
FILTER (SKVR Filter Project, University of Helsinki / Tampere University) is a computational platform for analysing Finnic oral poetry. RunoVerse imports its word co-occurrence statistics as a second set of association metrics alongside PMI. Coverage: 52,695 words with corpus frequency ≥ 10.
Use the metric selector to compare rankings:
Log-likelihood (LogL) — statistical significance; robust across frequency ranges. Primary FILTER metric.
Mutual Information (MI) — highlights rare but exclusive pairs.
Frequency — raw co-occurrence count (may differ from RunoVerse verse-line counts).
Ranking disagreements are highlighted: a word ranked top-5 by one metric but not another reveals pairs that are significant but infrequent, or vice versa.
Top Pairs
The 200 most notable word pairs across the entire corpus, ranked by combined PMI and frequency. These often represent core runosong formulas — parallelism pairs and refrains.
Pair types
Each word pair is classified by what drives the co-occurrence — sound pattern or meaning:
Sound-symbolic — words that echo each other phonologically: similar endings with different onsets (e.g. kannel–kummel). Detected via edit distance and shared suffixes.
Alliterative — words sharing the same initial sound (e.g. veli–veno). Alliteration is a core structural device in runosong.
Content — pairs that co-occur because of meaning or narrative context, not because of shared sounds (e.g. maja–kus, laev–meri). This is a residual category: any pair that is neither sound-symbolic nor alliterative.
The classification is based on word form only — it detects phonological patterns automatically but does not perform semantic analysis. High-PMI content pairs typically reflect thematic formulas (place + action, object + attribute) that oral poets used as building blocks.
Part of speech
Each word carries a grammatical badge: N (noun), V (verb), Adj (adjective), Adv (adverb), or a rarer class abbreviation. Tags come from the lexicon (UD tagset, ~98.5% coverage).
The POS filter narrows the collocate list to one class — useful for separating, say, nouns from verbs. "Other POS" covers proper nouns, numerals, pronouns, particles, etc. Words with no tag appear only under "All POS". The network graph is not affected by this filter.
Lemma view — spellings are pooled under one dictionary word
(neiu / neiut / neidu → neiu). The pooling is automatic and not sense-aware, so unrelated
words can merge and one word can split: nominative tuli "fire" is filed under the verb
tulema "came" while its other forms stay under tuli. Regional spellings
(koju → kodu) and identical Estonian/Finnish spellings merge too — use
Surface forms for dialect- or language-specific collocations. To check any
number here, click verses: the concordance lists the spellings currently
attributed to this word. For a few very frequent function words the count was computed over
a wider, older set of spellings, so the verse list can be slightly narrower than the number.
The concordance cannot search words shorter than three letters, so for those (ei)
the link returns nothing.
Loading collocate data...
Top Formulaic Pairs
The strongest pairs are genuine sound-symbolic refrain and echo formulas — paired nonsense-rhyme syllables like entten / tentten or pau / piu that recur together across the corpus. They score highly because the two words almost always appear side by side, not because of a counting error. Use the pair-type filter to surface alliterative or content-word pairs instead.