RunoVerse

Word Collocates

Explore word co-occurrence patterns across more than 4.3 million verse lines — discover formulaic pairs and poetic associations.

What this shows

Co-occurrence analysis of words in Finnic runosong verses. Two words are "collocates" if they frequently appear in the same verse line.

Association strength is measured by PMI (Pointwise Mutual Information) — high PMI means the pair appears together far more often than chance predicts.

How to read the results

Sorting

FILTER Metrics

FILTER (SKVR Filter Project, University of Helsinki / Tampere University) is a computational platform for analysing Finnic oral poetry. RunoVerse imports its word co-occurrence statistics as a second set of association metrics alongside PMI. Coverage: 52,695 words with corpus frequency ≥ 10.

Use the metric selector to compare rankings:

Ranking disagreements are highlighted: a word ranked top-5 by one metric but not another reveals pairs that are significant but infrequent, or vice versa.

Top Pairs

The 200 most notable word pairs across the entire corpus, ranked by combined PMI and frequency. These often represent core runosong formulas — parallelism pairs and refrains.

Pair types

Each word pair is classified by what drives the co-occurrence — sound pattern or meaning:

The classification is based on word form only — it detects phonological patterns automatically but does not perform semantic analysis. High-PMI content pairs typically reflect thematic formulas (place + action, object + attribute) that oral poets used as building blocks.

Part of speech

Each word carries a grammatical badge: N (noun), V (verb), Adj (adjective), Adv (adverb), or a rarer class abbreviation. Tags come from the lexicon (UD tagset, ~98.5% coverage).

The POS filter narrows the collocate list to one class — useful for separating, say, nouns from verbs. "Other POS" covers proper nouns, numerals, pronouns, particles, etc. Words with no tag appear only under "All POS". The network graph is not affected by this filter.

Lemma view — approximate dictionary-lemma aggregation, not sense-disambiguated. Inflected forms of one word merge (neiu / neiut / neidu → neiu), which is the main effect. But resolution is per spelling, so it can also split one concept across lemmas and merge unrelated ones: nominative tuli "fire" is filed under the verb tulema "came", while oblique fire forms stay under tuli — and pooling inflections enlarges such homonym grab-bags rather than fixing them. Regional spellings (koju → kodu) and some shared-spelling Estonian/Finnish forms also merge. Use Surface forms to study dialect- or language-specific collocations. (Cognates spelled differently, e.g. ladva / latva, stay separate.)

Loading collocate data...