Glossary
Corpus
The collection of texts and linguistic annotations available on this site. Corpus statistics describe this collection, not Sanskrit as a whole.
Devanāgarī
The script used for Sanskrit in the reader. The same passage can also be shown in IAST transliteration.
IAST
International Alphabet of Sanskrit Transliteration, a scholarly Latin-script convention using diacritics such as ā, ṛ, ṣ, and ṇ.
Token
A word-level unit in a parsed passage. A token can be a surface form, an inflected word, or in some cases an analysed compound.
Form
The exact written or transliterated token selected in a passage.
Lemma
The normalized lexical form under which related inflected forms are grouped. For a verb, this may differ substantially from the particular form in the verse.
Dhātu / root
A verbal root used to organise derivational and verbal relationships. A root is not always the most useful explanation for every noun or compound, but it can reveal a productive lexical family.
Gaṇa
A traditional verbal class. The Dhātu explorer displays the supplied gaṇa as root metadata; it is distinct from pada.
Pada
A root’s supplied inflectional designation: Parasmaipada, Ātmanepada, or Ubhayapada. It may be absent when the source data does not provide it.
Guṇa and vṛddhi
Traditional vowel grades. The Dhātu explorer uses initial guṇa- and vṛddhi-grade patterns only to organise linked words for browsing; those groups are not complete derivational analyses.
Upasarga
A verbal prefix. Where root-source data supplies upasargas, the Dhātu explorer lists them on the root record.
Compound
A Sanskrit expression combining two or more lexical members into one unit. The site shows a breakdown only when the current analysis supplies one.
Concordance
A collection of contexts in which a lemma or form occurs. A sample concordance is an aid to comparison, not an exhaustive edition.
Occurrence
One attested instance in the indexed corpus. “Texts appeared in” counts works, while “total occurrences” counts individual instances.
Concept
A semantic grouping used to connect mapped lemmas for exploration. Concepts on this site use a WordNet-derived classification; they are not a Sanskrit-native philosophical taxonomy or a final interpretation of a passage.
Supersense
One of the broad semantic classes used for corpus mappings, such as Act, Person, State, or Time. Supersenses are the most useful level of the Concepts section for comparing mapped Sanskrit vocabulary.
Synset
A more specific node in the imported WordNet hierarchy. The current corpus maps lemmas chiefly to broad supersenses, so synsets are best read as a reference hierarchy rather than as exhaustive Sanskrit semantic groups.
Reference
The location of a passage in the text’s own system: for example, chapter and verse, or maṇḍala, sūkta, and ṛc.
CoNLL-U
A plain-text format for syntactically and morphologically annotated language data. It is part of the site’s data pipeline; readers need not know the format to use the corpus.