Dharmalib

Glossary

Corpus

The collection of texts and linguistic annotations available on this site. Corpus statistics describe this collection, not Sanskrit as a whole.

Devanāgarī

The script used for Sanskrit in the reader. The same passage can also be shown in IAST transliteration.

IAST

International Alphabet of Sanskrit Transliteration, a scholarly Latin-script convention using diacritics such as ā, ṛ, ṣ, and ṇ.

Token

A word-level unit in a parsed passage. A token can be a surface form, an inflected word, or in some cases an analysed compound.

Form

The exact written or transliterated token selected in a passage.

Lemma

The normalized lexical form under which related inflected forms are grouped. For a verb, this may differ substantially from the particular form in the verse.

Dhātu / root

A verbal root used to organise derivational and verbal relationships. A root is not always the most useful explanation for every noun or compound, but it can reveal a productive lexical family.

Gaṇa

A traditional verbal class. The Dhātu explorer displays the supplied gaṇa as root metadata; it is distinct from pada.

Pada

A root’s supplied inflectional designation: Parasmaipada, Ātmanepada, or Ubhayapada. It may be absent when the source data does not provide it.

Guṇa and vṛddhi

Traditional vowel grades. The Dhātu explorer uses initial guṇa- and vṛddhi-grade patterns only to organise linked words for browsing; those groups are not complete derivational analyses.

Upasarga

A verbal prefix. Where root-source data supplies upasargas, the Dhātu explorer lists them on the root record.

Compound

A Sanskrit expression combining two or more lexical members into one unit. The site shows a breakdown only when the current analysis supplies one.

Concordance

A collection of contexts in which a lemma or form occurs. A sample concordance is an aid to comparison, not an exhaustive edition.

Occurrence

One attested instance in the indexed corpus. “Texts appeared in” counts works, while “total occurrences” counts individual instances.

Concept

A semantic grouping used to connect mapped lemmas for exploration. Concepts on this site use a WordNet-derived classification; they are not a Sanskrit-native philosophical taxonomy or a final interpretation of a passage.

Supersense

One of the broad semantic classes used for corpus mappings, such as Act, Person, State, or Time. Supersenses are the most useful level of the Concepts section for comparing mapped Sanskrit vocabulary.

Synset

A more specific node in the imported WordNet hierarchy. The current corpus maps lemmas chiefly to broad supersenses, so synsets are best read as a reference hierarchy rather than as exhaustive Sanskrit semantic groups.

Reference

The location of a passage in the text’s own system: for example, chapter and verse, or maṇḍala, sūkta, and ṛc.

CoNLL-U

A plain-text format for syntactically and morphologically annotated language data. It is part of the site’s data pipeline; readers need not know the format to use the corpus.

A project by designLab at Bodha.