About RunoVerse
For questions about RunoVerse, contact veskis.kaarel@gmail.com
What this page covers
This About page is the reference guide for the entire RunoVerse platform. It documents the four source corpora (SKVR, JR, KR, ERAB), corpus-wide statistics (439K lemmas, 15.3M tokens, 292K poems), and the methodology behind the lemmatization and AI annotation pipelines.
Sections on this page
Source corpora — descriptions and poem counts for each collection, with a visual bar chart.
Lexicon statistics — counts for lemmas, word forms, tokens, cognate pairs, translation/etymology confidence, and gloss coverage.
Data sources — explains the three source categories (Corpus Only, Both Sources, DeepSeek Only) and what the color-coded word forms and agreement badges mean.
Similarity — documents all 5 poem-level and 5 verse-level similarity algorithms with their basis and coverage.
Dictionary annotations — lists all 9 lexicographic sources (EMS, EKSS, IMS, ERLA, VMS, Seto, SMS, KKS, VKS).
Lemmatization — describes the Estonian (EstNLTK) and Finnish (Omorfi/Voikko/Stanza) processing pipelines, plus the DeepSeek AI annotation layer.
Explore cards
The card grid below links to the main pages of RunoVerse (70+ pages in all). Each card shows a short description of the page. For longer explanations, see the five Feature Guides (Lexicon, Similarity, Poetics, Cross-Lingual, Corpus).
Navigating the site
How to Cite
If you use RunoVerse in academic work, please cite:
BibTeX entry
@misc{veskis2026runoverse,
author = {Veskis, Kaarel},
title = {{RunoVerse}: A Digital Explorer for Finnic Runosong Corpora},
year = {2026},
publisher = {Estonian Folklore Archives of the Estonian Literary Museum},
url = {https://runoverse.org/},
note = {Version 1.0. Accessed: }
}
BibLaTeX users may prefer @online (core BibLaTeX type with native url/urldate fields) or @software (requires the biblatex-software package). @misc is used here for maximum compatibility with all BibTeX processors.
To cite an individual poem, use the Cite button next to the poem title in the Reader — it produces a citation naming the source corpus or archive first, with RunoVerse as the access platform. Poem deep links of the form reader.html?poem=<id> are stable and safe to use in references.
What is this?
RunoVerse is an exploration toolkit for the Finnic runosong (folk poetry) corpora. It brings together the poems, their lemmatized word data, dictionary annotations and translations from four collections, so the shared Finnic poetic tradition can be read, searched, compared and mapped across language boundaries.
Please note that RunoVerse is under active development. The lemmatization of historical dialectal texts is inherently approximate, and the AI-generated translations, etymological analyses, and similarity metrics should be considered experimental. The statistics and counts shown may change as the data is refined. This tool is intended as an exploratory aid, not as a definitive reference.
Source corpora
| Collection | Language | Description |
|---|---|---|
| SKVR | Finnish, Karelian, Ingrian, Votic | Suomen Kansan Vanhat Runot – the Finnish Literature Society's published collection of Finnic runosong. Finnish Literature Society (SKS). 89,247 poems. |
| JR | Finnish, Karelian, Ingrian, Votic, Ludic, Veps | Julkaisemattomat Runot – unpublished folk poetry from SKS folklore archives. 85,585 poems. |
| ERAB | Estonian (incl. Seto); some Ingrian Finnish songs | Eesti Regilaulude Andmebaas – Database of Estonian Runosongs. Estonian Folklore Archives of the Estonian Literary Museum. 108,969 poems. |
| KR | Finnish, Karelian, Estonian | Kirjalliset Runot – a corpus of literary poems based on public domain sources, including Kalevala (1849), Wanha Kalevala (1835), Kanteletar, Kalevipoeg, and other published works. Compiled by Sakari Korpikallio (Finnish Literature Society, SKS) within the FILTER and REFOP projects (Research Council of Finland). CC-BY 4.0. 10,544 poems. |
The KR corpus differs in kind from the other three. Its name, Kirjalliset Runot ("Written Poems"), reflects its origin: Kalevala-metre verse preserved in written and printed sources rather than recorded from oral singers in the field. As Kati Kallio (Finnish Literature Society) documents, verse preserved this way survives largely as short pieces embedded in other documents – such as proverbs, charms, dedications, and congratulatory or condolence verses – alongside the large 19th-century literary editions such as the Kalevala, Kanteletar, and Kalevipoeg that also form part of this corpus. See Kallio's study of literary Kalevala-metre poetry for background: folklore.ee/folklore/vol67/kallio.pdf.
The three archival corpora, too, are predominantly runosong but not homogeneous in genre. Alongside the runosong core they include charms and incantations (a substantial share of the classified Finnish poems), transitional-form and end-rhymed newer songs, dance and children's songs, and other folklore genres recorded into the same collections. RunoVerse keeps this material together and loosely refers to the whole as runosong; the Genre Explorer shows the genre breakdown for the Estonian corpus.
Neither corpus group is a single language. SKVR and JR are predominantly Finnish, but they hold substantial Karelian (including Olonets Karelian) and Ingrian (Izhorian and Ingrian Finnish) material, smaller amounts of Votic, Ludic and Veps songs, and a few Sami and Estonian texts; by the dialect area of the collection place, roughly a third of the SKVR poems and a quarter of the JR poems were recorded in Ingria and Karelia. ERAB spans the North and South Estonian traditions, including Seto songs, and holds a few hundred songs collected in Ingria and in Finnish-speaking settlements. KR is mostly Finnish and Karelian, with the Estonian Kalevipoeg in addition. The corpus labels FI and ET used across the site therefore name the two corpus groups, not two languages; the dialect layer of the Places page shows where each language variety was recorded.
Lexicon statistics
| Measure | Count | Notes |
|---|---|---|
| Unique lemmas | 439,746 | Distinct base forms across all corpora (incl. 183,137 DeepSeek-only) |
| Unique wordforms | 1,480,455 | Distinct word forms occurring across all poem texts |
| Wordform–lemma mappings | 2,083,995 | Total mappings from inflected forms to lemmas (one wordform can map to multiple lemmas) |
| Total tokens | 15,264,640 | Total word occurrences in source texts |
| Poems | 292,089 | Unique poems with full verse texts available in the poem context viewer (294,345 total in source corpora; some excluded due to missing verse text data) |
| Finnish-only lemmas | 206,518 | Lemmas from Finnish corpus only (SKVR/JR collections) |
| Estonian-only lemmas | 100,835 | Lemmas from Estonian corpus only (ERAB) |
| Shared (Finnic) lemmas | 1,240 | Lemmas found in both Finnish and Estonian sources |
| Cognate pairs (ET↔FI) | 4,913 | Automatically discovered Estonian-Finnish cognate pairs based on translation overlap, etymological roots, and orthographic similarity (186 shared forms, 2,013 near-exact, 2,714 translation-bridged; a further 72 exact-form pairs are demoted and not shown) |
| Translation confidence | 192,236 | Lemmas with DeepSeek translation consistency score (8,446 strong, 11,069 good, 41,650 moderate, 131,071 low; 113,502 no data) |
| Etymology confidence | 211,919 | Lemmas with DeepSeek etymology consistency score (11,748 strong, 13,183 good, 46,078 moderate, 140,910 low; 93,819 no data) |
| Gloss coverage | 91.3% | Word forms with English translation (1,344,094 of 1,472,442), including 11,423 Claude Opus supplementary glosses |
| Corpus attestations | 15,264,640 | SKVR: 4,522,811 · JR: 3,398,967 · KR: 1,389,200 · ERAB: 7,341,908 |
Statistics reflect the current state of the lemmatized data and may change as lemmatization is refined.
Data sources and source filter
Each lemma in the lexicon has been tagged with one of three source categories, reflecting how it was identified. The source filter dropdown in the main view lets you filter by these categories:
| Source | Lemmas | Meaning |
|---|---|---|
| Corpus Only | 5,217 | Lemma was identified by the corpus lemmatization pipeline. None of the word forms listed under this lemma were matched to DeepSeek annotations during the merge. However, the lemma string itself may still appear as a word form in DeepSeek data, which means some “Corpus Only” entries can still have AI-generated translations visible via the A–Z browse. |
| Both Sources | 251,392 | Lemma comes from the corpus pipeline, and at least one of its word forms also appears in the DeepSeek annotations (possibly under a different lemma). These entries typically include AI-generated translations and may have cross-references (dsLemma) to alternative lemmatizations. Word forms in “Both Sources” entries are color-coded: green when both systems agree on the lemma, amber when DS assigns a different lemma, and gray when the word form is not in DS data. |
| DeepSeek Only | 183,137 | Lemma exists only in the DeepSeek annotations. The underlying word forms often appear in the corpus under different lemmas (96% of cases), but this particular lemmatization is unique to the AI analysis. |
The source categories reflect word-form-level overlap between the two lemmatization systems, not whether an entry has translations. Because the corpus pipeline and DeepSeek sometimes lemmatize the same word forms differently, a word form can belong to a “Corpus Only” lemma while also appearing independently in the DeepSeek data under a different lemma. The “Both Sources” category captures entries where the same word forms were recognized by both systems.
The agreement badge in the DeepSeek tab shows a ratio like “30/35 agree +15 n/a”, meaning 30 out of 35 DS-covered word forms have the same lemma in both systems, and 15 word forms are not present in the DS data. Hover over the badge for a full breakdown.
DeepSeek AI annotations
A subset of the corpus was independently annotated using DeepSeek, a large language model, to provide additional linguistic analysis. The AI annotations include:
- English translations of Estonian and Finnish word forms
- Etymological notes and cognate identification
- Morphological descriptions (case, number, tense, etc.)
- Part-of-speech tagging
| Measure | Count | Notes |
|---|---|---|
| DeepSeek tokens | 5,962,070 | AI-annotated token occurrences (ET: 2,867,388 + FI: 3,094,682) |
| DeepSeek-only lemmas | 183,137 | Lemmas unique to the AI analysis |
| English translations | 241,141 | Unique English terms extracted from AI annotations, browsable via A–Z (1,252,781 total mappings) |
| Cross-references | 91,754 | Entries linking to alternative lemmatizations between corpus and DeepSeek |
AI-generated annotations are provided as supplementary material and have not been manually verified. They should be used with appropriate caution, particularly for etymological claims and translations of rare dialectal forms.
Similarity and embedding data
The lexicon includes two word-level similarity systems to help explore relationships between word forms:
- Word form similarity – Edit-distance and phonological similarity between inflected forms across the corpus. Covers 1,166,348 word forms with ranked nearest neighbors and lemma-agreement indicators.
- BERT embeddings – Contextual nearest neighbors from a BERT model fine-tuned on Estonian runosong texts. Provides 190,975 query lemmas with their 10 nearest semantic neighbors, capturing meaning-based rather than form-based similarity.
Poem similarity
Five poem-level similarity algorithms identify related poems across a corpus of roughly 290,000 poems. Results are available in the Poem Reader (Related Poems panel) and the standalone Similarity Explorer with side-by-side comparison, network graphs, and geographic/temporal analytics.
| Algorithm | Basis | Description |
|---|---|---|
| TF-IDF Lemma | Lemma-level | Cosine similarity on TF-IDF vectors of lemmatized poem texts. Captures thematic similarity through shared vocabulary, weighted by corpus-level term importance. Top 50 neighbors per poem. |
| Wordform Overlap (Jaccard) | Exact wordforms | Jaccard index (|A∩B| / |A∪B|) over raw wordform sets. Identifies poems sharing exact surface forms, useful for detecting recurring lines and direct textual parallels. |
| Thematic (Translation-pivot) | Cross-lingual | Boolean-IDF cosine similarity over English translations derived from DeepSeek annotations, with lemma-level fallback for improved coverage. Enables cross-lingual comparison between Estonian and Finnish poems via a shared semantic space. Top 50 neighbors per poem. |
| Alignment | Character n-gram | Verse sequence alignment using character bigram cosine similarity and Wagner-Fischer dynamic programming, from the FILTER project (Janicki, Kallio & Sarv 2023). Covers 256,970 poems across SKVR, JR, KR, and ERAB. Captures structural similarity — poems that follow the same verse order score high. Shows aligned verse pair excerpts for top matches. Top 50 neighbors per poem. |
| Verse-level RRF | Verse-level fusion | Fuses Jaccard, TF-IDF, Translation, CharBigram, and Sentence similarity at the verse level using Average-Best-Per-Verse aggregation, then combines all five via Reciprocal Rank Fusion (k=60) into a single poem-level ranking. Shows T/J/Tr/C/Se algorithm contribution badges. |
The Similarity Explorer shows cross-algorithm agreement badges (BOTH) when poems appear in multiple algorithms' results, and ET↔FI badges for cross-corpus matches in the Translation-pivot, Alignment, and Verse-level RRF algorithms.
Verse similarity
Five algorithms (Jaccard, TF-IDF, Translation-pivot, CharBigram, Sentence) operate at the individual verse level across 4.32 million verse occurrences. Each verse is compared against all verses in other poems, with up to 20 nearest neighbors stored per algorithm.
| Metric | Value |
|---|---|
| Total verses indexed | 4,316,744 |
| Poems with verse data | 289,702 |
| Unique verse types (search index) | 2,906,535 |
| Verse clusters | 200 |
Verse similarity is available in the Poem Reader (click the expand arrow on any verse line) and the Verse Similarity Explorer, which also provides full-text verse search and a browser for the top 200 verse clusters – recurring verse lines ranked by cluster size across all four corpora.
Explore RunoVerse
Feature guides
RunoVerse contains over 30 interconnected tools for exploring the Finnic runosong tradition. These guides describe each tool in detail — what it shows, what data powers it, and how to use it.
Key features
- Search by lemma, word form, or English translation with diacritics-insensitive matching
- Filter by language, part of speech, and data source (corpus, DeepSeek, or both)
- Dictionary annotations from 9 Estonian and Finnish lexicographic sources
- DeepSeek AI translations, etymology, and morphological descriptions for 165K poems
- Five poem similarity algorithms with network graphs, geographic maps, and temporal analytics
- Verse-level similarity across 4.32M verses with inline expansion and full-text search
- Cross-lingual exploration: 4,913 cognate pairs, 1,240 shared lemmas, 49K etymology families
- Poetic analysis: alliteration patterns, semantic parallelism, recurring phrases, and meter
- Poem reader with word-level glosses, POS tags, and per-verse similarity
- Geographic and temporal corpus analysis across 803 collection places and four centuries
- Bookmarkable deep links, keyboard navigation, and CSV export
Dictionary annotations
Word entries are enriched with definitions from Estonian and Finnish lexicographic sources:
- EMS – Eesti murrete sõnaraamat (Dictionary of Estonian Dialects). Institute of the Estonian Language.
- EKSS – Eesti keele seletav sõnaraamat (Explanatory Dictionary of Estonian). Institute of the Estonian Language.
- IMS – Ida-Eesti murdesõnastik (Eastern Estonian Dialect Dictionary). Institute of the Estonian Language.
- ERLA – Harva ja vähem kasutatavate sõnade sõnastik (Glossary of Rare Folk-Song Words). Estonian Folklore Archives of the Estonian Literary Museum.
- VMS – Vähemtuntud murdesõnade seletusi (Glossary of Lesser-Known Dialect Words). Estonian Folklore Archives of the Estonian Literary Museum.
- Seto – Seto sõnastik (Seto Dictionary). Inge Käsi, Institute of the Estonian Language, 2016.
- SMS – Suomen murteiden sanakirja (Dictionary of Finnish Dialects). Kotimaisten kielten keskus (Kotus). CC BY 4.0.
- KKS – Karjalan kielen sanakirja (Dictionary of the Karelian Language). Kotimaisten kielten keskus (Kotus). CC BY 4.0.
- VKS – Vanhan kirjasuomen sanakirja (Dictionary of Old Literary Finnish). Kotimaisten kielten keskus (Kotus). CC BY 4.0.
Lemmatization
Estonian texts were lemmatized using EstNLTK morphological analysis combined with the University of Tartu dialect corpus (Lindström, Lippus & Tuisk, 2019), runosong word lists annotated by Kristiina Ross (Ross & Lohk, 2017), multiple lexical resources (EMS, EKSS, VES, ERLA, and others), expert manual annotations (37% of the corpus), and iterative automated correction cycles.
Finnish texts were lemmatized using a combinatory approach with a multi-tier fallback chain including Omorfi, Voikko, and Stanza, supplemented by the dictionaries Suomen murteiden sanakirja (SMS), Karjalan kielen sanakirja (KKS), and Vanhan kirjasuomen sanakirja (VKS).
In addition, approximately 165,000 poems were independently annotated using DeepSeek-R1, a large language model, run on the LUMI supercomputer. The annotation was carried out by Lidia Pivovarova (University of Helsinki), following the prompt design methodology developed collaboratively by computational linguists and folklorists (Pivovarova et al., 2025). The AI analysis produced lemmatizations, English translations, morphological descriptions, and etymological roots for each word token. These annotations yielded 107,110 additional lemmas not present in the corpus pipeline, and provided cross-references between the two lemmatization systems for entries where both recognized the same word forms.
References and acknowledgements
RunoVerse builds on the work of the FILTER project (Formulaic Intertextuality, Thematic Networks and Poetic Variation across Regional Cultures of Finnic Oral Poetry), whose corpus infrastructure and verse-alignment methods underlie several similarity layers here. Runoregi, FILTER's own interface for exploring poem and verse similarity in Finnic oral poetry, is the natural companion tool for readers who want a second, independent view of many of the same texts.
Corpora
- SKVR – Finnish Literature Society (Suomalaisen Kirjallisuuden Seura, SKS). Suomen Kansan Vanhat Runot. Digital corpus. skvr.fi. CC BY 4.0.
- JR – Finnish Literature Society (SKS). Julkaisemattomat Runot. aineistot.finlit.fi/skvr/runoluettelo-jr
- ERAB – Oras, J.; Saarlo, L.; Sarv, M.; Labi, K.; Uus, M.; Šmitaite, R. (comps.). Eesti Regilaulude Andmebaas. Estonian Folklore Archives of the Estonian Literary Museum. 2003–present. folklore.ee/regilaul/andmebaas
- KR – Korpikallio, S. (comp.). Kirjalliset Runot. A corpus of literary poems in Kalevala metre (1658–1866), compiled at the Finnish Literature Society (SKS) within the FILTER and REFOP projects (Research Council of Finland, nos. 333138, 340647). CC-BY 4.0. github.com/sks190/KR
Place coordinates
The Villages, communes and other places layer on the Places map draws 6,279 places finer than the collection place – villages, communes, manors, settlements – covering 124,001 of the 130,757 records that name one (94.8%). Its coordinates come from the registers below. A marker shows a place as a register records it, not a surveyed centre: register coordinates are name-cluster reference points, good to about the diameter of a village.
- Maa- ja Ruumiamet / Rahvusarhiiv – Ajalooline haldusjaotus: vallad 1922, vallad 1938; mõisad kuni 1917 (mõisakeskused); kihelkonnad kuni 1917. Data as of 2025-05-09. Also mõisamaad kuni 1917 (manor estate lands), data as of 2026-09-01: these supply the approximate extent drawn for communes that predate both commune registers – a manor’s estate land is related to, but not identical with, the commune attached to it. geoportaal.maaamet.ee
- Maa- ja Ruumiamet – Asustusüksused, data as of 2024-12-01. geoportaal.maaamet.ee
- Maanmittauslaitos – Nimistö 05/2026. CC BY 4.0. maanmittauslaitos.fi
- Maanmittauslaitos – Kiinteistörekisterikartta (cadastral index map), INSPIRE WFS, pulled 2026-09-08. CC BY 4.0. Village rings on Finnish pages of the villages map are inferred from these parcels, grouped by the register-village element of the parcel identifier.
- WarSampo Karelian places – Semantic Computing Research Group (SeCo), Aalto University / University of Helsinki. CC BY 4.0, retrieved August 2026. ldf.fi/dataset/warsa
- Karjalan kartat – place-name index of the ceded-Karelia map sheets, compiled by Jyrki Tiittanen. © Maanmittauslaitos / Puolustusvoimat; used by written permission of the National Land Survey, August 2026. karjalankartat.fi
- Nimiarkisto – the Names Archive of the Institute for the Languages of Finland (Kotimaisten kielten keskus). Database-form material, CC BY 4.0, read August 2026; each village placed from it cites its own archive record in its popup. The scanned collection maps and name slips are outside that licence and none is reproduced here. nimiarkisto.fi
- Inkeri.ru / Inkeri.fi – Ingrian parish and village registers of the Inkerin kulttuuriseura (Ingrian Culture Society), pairing the Finnish village names with their Russian counterparts. Read August 2026; each village placed from them cites the register page in its popup. inkeri.ru
- Etomesto.com – old-map service; twenty-one points for Ingrian villages that survive only on historical map sheets – six as tracts (урочище) in its place index, fifteen read off its georeferenced sheets (Finnish prewar topographic maps, the 1890 one-verst map of St Petersburg guberniya, the RKKA km map, the 1930 Leningrad-surroundings map) – all read in August 2026, each citing its index page or map view in the popup. etomesto.com
- Wikidata – Wikidata contributors. CC0 1.0, retrieved August 2026. wikidata.org
- GeoNames – CC BY 4.0, retrieved August 2026. geonames.org
- OpenStreetMap – © OpenStreetMap contributors; data available under the Open Database License (ODbL 1.0), retrieved August 2026. openstreetmap.org/copyright
- Wikipedia – Russian, Finnish and Swedish Wikipedia articles, CC BY-SA 4.0, retrieved August 2026. wikipedia.org
- Regionum.ru – register of Leningrad-oblast settlements (two points), retrieved August 2026. regionum.ru
- Alta.ru railway directory – station coordinates (one point), retrieved August 2026. alta.ru/railway
- KNAB – Kohanimeandmebaas, Eesti Keele Instituut, compiled by Peeter Päll. Database last updated 2022-07-28. Individual names were looked up here through August 2026; a data extract received 2026-08-31 covers the Estonian names and the Lutsi, Leivu and Crimean language islands, and placed 171 further points. Each such point names its KNAB record number. arhiiv.eki.ee/knab
- Yandex Maps – two Ingrian tract points (Räkälä, Oussimäki), read off the service in August 2026. yandex.com/maps
- Olonets province settlement list, 1905 – Blagoveshchenskiy, I. I. (comp.). Список населённых мест Олонецкой губернии по сведениям за 1905 год [List of populated places of Olonets Governorate, data for 1905]. Olonets Governorate Statistical Committee, Petrozavodsk, 1907. An official publication, in the public domain; digitised and located village by village by the олонецкая-губерния.рф project, retrieved 2026-08-19. олонецкая-губерния.рф
- Luovutetun Etelä-Karjalan pitäjät – Lappeenrannan maakuntakirjasto, Carelica collection. Village rosters for five Karelian Isthmus places, retrieved 2026-08-18; no coordinate comes from it. luovutetunetelakarjalanpitajat.fi
- Juminkeko – Juminkeko Foundation, Viena rune-village pages. The location of one destroyed rune village (Tšena, Vuokkiniemi), retrieved 2026-08-19; the coordinate itself is GeoNames’. juminkeko.fi
About 850 of these places are pinned one by one rather than by the automatic matcher. Most are read from the registers above; a few dozen rest on a single cited source – a parish history, a 1905 volost roster, a map service – and each of those names its own source in that place’s popup.
Lemmatization tools
- EstNLTK – Laur, S.; Orasmaa, S.; Särg, D.; Tammo, P. (2020). EstNLTK 1.6: Remastered Estonian NLP Pipeline. Proceedings of LREC 2020, pp. 7154–7162. github.com/estnltk/estnltk
- Vabamorf – Kaalep, H. J.; Vaino, T. (2001). Complete morphological analysis in the linguist’s toolbox. Congressus Nonus Internationalis Fenno-Ugristarum, 5, pp. 9–16.
- Omorfi – Pirinen, T. A. (2015). Omorfi – Free and open source morphological lexical database for Finnish. Proceedings of NODALIDA 2015, pp. 313–315. github.com/flammie/omorfi
- Voikko – Pitkänen, H. Voikko – Free linguistic software for Finnish. voikko.puimula.org
- Stanza – Qi, P.; Zhang, Y.; Zhang, Y.; Bolton, J.; Manning, C. D. (2020). Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. Proceedings of ACL 2020: System Demonstrations. stanfordnlp.github.io/stanza
- DeepSeek-R1 – DeepSeek-AI (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv preprint arXiv:2501.12948. deepseek.com
- Pivovarova et al. – Pivovarova, L.; Kallio, K.; Kanner, A.; Lindström, J.; Mäkelä, E.; Saarlo, L.; Veskis, K.; Väina, M. (2025). Benchmarking Large Language Models for Lemmatization and Translation of Finnic Runosongs. Proceedings of the 10th International Workshop on Computational Linguistics for Uralic Languages (IWCLUL 2025), pp. 87–105. Association for Computational Linguistics. aclanthology.org/2025.iwclul-1.12
- Lindström, Lippus & Tuisk – Lindström, L.; Lippus, P.; Tuisk, T. (2019). The online database of the University of Tartu archives of Estonian dialects and kindred languages and the Corpus of Estonian dialects. In Björklof, S. and Jantunen, S. (eds.), Multilingual Finnic: Language Contact and Change. Helsinki: Finno-Ugrian Society, pp. 327–350. doi.org/10.33341/uh.85040
- Ross & Lohk – Ross, K.; Lohk, A. (2017). Words, forms and phrases in Estonian folksongs and hymns. Folklore: Electronic Journal of Folklore, 67, pp. 49–64. doi.org/10.7592/FEJF2017.67.ross_lohk
- Janicki, Kallio & Sarv – Janicki, M.; Kallio, K.; Sarv, M. (2023). Exploring Finnic written oral folk poetry through string similarity. Digital Scholarship in the Humanities, 38(4), pp. 1510–1525. doi.org/10.1093/llc/fqac034
Lexical resources
- SMS – Kotimaisten kielten keskus (Kotus). Suomen murteiden sanakirja (Dictionary of Finnish Dialects). kaino.kotus.fi/sms. CC BY 4.0.
- KKS – Kotimaisten kielten keskus (Kotus). Karjalan kielen sanakirja (Dictionary of the Karelian Language). kaino.kotus.fi/kks. CC BY 4.0.
- VKS – Kotimaisten kielten keskus (Kotus). Vanhan kirjasuomen sanakirja (Dictionary of Old Literary Finnish). CC BY 4.0.
- EMS – Institute of the Estonian Language. Eesti murrete sõnaraamat. eki.ee/dict/ems
- EKSS – Institute of the Estonian Language. Eesti keele seletav sõnaraamat. eki.ee/dict/ekss
- VES – Võro Institute. Võro-eesti synaraamat (comp. Jüvä Sullõv). folklore.ee/Synaraamat
- ERLA – Estonian Folklore Archives of the Estonian Literary Museum. Harva ja vähem kasutatavate sõnade sõnastik. folklore.ee/laulud/erla
Rights and reuse
The poem and legend texts remain under the terms of their source collections (see the corpus references above: SKVR and KR are published under CC BY 4.0; the JR and ERAB materials belong to the Finnish Literature Society and the Estonian Folklore Archives). Authorship of the annotation layers is described above — the AI-produced lemmatizations, English translations, and morphological and etymological analyses come from the DeepSeek-R1 annotation carried out by Lidia Pivovarova (University of Helsinki), following the methodology of Pivovarova et al. (2025), and should be treated as experimental. The RunoVerse compilation itself — the site, its indexes, and its visualizations — is © Kaarel Veskis. Use in research and teaching with attribution is welcome; for other uses, please get in touch. Bulk reuse or redistribution of the underlying data files (annotation layers, indexes, translations) is not covered by this and needs the author’s prior agreement.
Contact
For questions about RunoVerse, contact veskis.kaarel@gmail.com