Two collection places, side by side. For each place you'll see a small map, the usual corpus metrics (poems, verses, unique lemmas, collectors), a distinctive vocabulary list, and a semantic-domain radar chart. The design answers: how does the sung vocabulary of one parish differ from another's?
When both places have enough text and are similar in size (both ≥ 500 tokens, size ratio < 50×), the list shows words statistically overrepresented in one place relative to the other using the log-likelihood ratio (G², p < 0.05). The number next to each word is the relative frequency ratio — e.g. “13.4×” means the word's per-token rate is 13.4 times higher here than in the other place; “unique” means it does not appear in the other place at all. Common function words (pronouns, conjunctions, negation) are filtered out.
When one place is very small or the corpora are too asymmetric for a reliable pair comparison, the page falls back to pre-computed TF-IDF distinctive words (each place vs all same-language places in the corpus). The header changes from “Distinctive Words (vs other place)” to just “Distinctive Words” to signal this.
The bar chart shows how much of each place's annotated vocabulary falls into the major semantic domains (body, kinship, nature, emotions, actions, artefacts, …). A place with a big kinship bar but a small nature bar sings mostly wedding/family themes; the reverse might be hunting or work songs.