RunoVerse
Recent clusters:

Verse Similarity Analysis

Cross-algorithm comparison of verse-level similarity across 4.3 million Finnic runosong verses, with formulaic cluster analysis and geographic spread.

What this page shows

A statistical analysis of verse-level similarity across the entire Finnic runosong corpus (4.3 million verse occurrences, 289,702 poems). The analysis compares how five different algorithms find similar verses and identifies formulaic patterns.

Algorithm Comparison Dashboard

Cross-Algorithm Discordance

For verses that appear in multiple algorithms, how much do their neighbor lists overlap? High discordance means the algorithms find fundamentally different similar verses. Low overlap suggests the algorithms capture complementary aspects of similarity.

Cluster Distribution Charts

Formulaic Verse Clusters

Geographic Spread Charts

Scatter plot shows how cluster size relates to geographic distribution. Histogram shows the frequency of clusters by number of distinct places. Widely distributed formulas represent the most universal elements of Finnic oral poetry.

Keyboard Shortcuts

/Focus cluster search box JSelect next cluster row KSelect previous cluster row EnterOpen selected cluster overlay EscClose overlay / clear search / blur input SShare: copy current URL to clipboard XExport clusters to CSV RRandom cluster from current filtered set ?Toggle this help panel TToggle verse translations 1Scroll to Algorithm Comparison 2Scroll to Discordance section 3Scroll to Clusters section 4Scroll to Geographic section
Algorithm Comparison Dashboard
Side-by-side statistics for each similarity algorithm. Match types: s = same language, x = cross-lingual, w = within-poem.
Cross-Lingual Match Percentage
Average Similarity Score
Verse Coverage (Verses with Matches)
Total Similarity Pairs
Cross-Algorithm Discordance
Pairwise comparison of how much the algorithms agree on their top neighbors for the same verses.
Language Distribution of Formulaic Patterns
How the 200 formulaic clusters divide across Estonian-dominant, Finnish-dominant, and bilingual categories.
Cluster Size Distribution (200 largest clusters)
Distribution of the 200 formulaic clusters by verse count, colored by dominant language composition. These are the very top of the distribution — a typical cluster in the full clustering holds only a handful of verses (median 7; see below).
The Full Clustering at a Glance
For scale: the complete verse clustering behind this site's similarity pages (May 2026 build) groups 2,902,367 of the corpus's 4,316,744 verse lines (67.2%) into 682,343 clusters of at least 5 similar verses. A typical cluster holds just a handful — median 7 verses, mean 9.1, largest 159.
versesclusters
5
145,822 (21.4%)
6–10
359,786 (52.7%)
11–20
149,107 (21.9%)
21–50
27,140 (4.0%)
51–100
463 (<0.1%)
101–159
25 (<0.01%)
The 200 showcase clusters on this page come from a separate formula-mining step that merges overlapping similarity neighbourhoods into much larger families, so their sizes are not directly comparable with this distribution.
Geographic Spread vs Frequency
Relationship between cluster size and geographic distribution across distinct collection places, for the 200 largest clusters.
Formulaic Verse Clusters (Top 200)
The 200 largest groups of nearly identical verses found across the corpus. These represent formulaic expressions shared across poems and regions.
Geographic Spread of Formulaic Clusters
How widely the 200 largest formulaic verse clusters are distributed across distinct collection places.
Cluster Size vs. Distinct Places
Distribution of Geographic Spread (# Places)