We build language technology for Sanskrit and other Indic languages: corpora, tokenizers and language models trained from scratch. Every training run we do, finished or live, is published at the Sansar Lab.
Sansar is a family of small causal language models trained only on Sanskrit text, with a Sanskrit tokenizer of its own. Each model is trained from scratch; none is a fine-tune of an English model.
| Model | Parameters | Bits per byte (held-out, ex-Gītā) ↓ |
|---|---|---|
| sansar-700m | 704M | 0.5307 |
| sansar-350m | 318M | 0.5547 |
| sansar-125m | 97M | 0.6039 |
| sansar-60m | 63M | 0.6434 |
| sansar-20m | 27M | 0.7177 |
Scores are measured on frozen held-out Sanskrit sets: DCS gold sentences, prose, Vedic and out-of-domain texts. Held-out text is masked out of training. Every model card gives the per-set numbers, the training data and a version history.
Scored the same way, general base models need more bits per byte: Gemma 3 4B 0.6965, Qwen3-4B 0.7071, Llama 3.2 3B 0.7122, Sarvam-1 0.7465. sansar-700m beats all four on every held-out set with 704M parameters (Krutrim-2 12B, scored earlier in a pass that is not strictly comparable, is not in this list). A live demo is at saansar.com/demo.
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("MuseMesh/sansar-700m", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("MuseMesh/sansar-700m", trust_remote_code=True)
ids = tok("संस्कृतं नाम दैवी वाक्", return_tensors="pt")
print(tok.decode(model.generate(**ids, max_new_tokens=60, do_sample=True, top_k=100)[0]))
Small models trained from scratch with the same recipe (Muon, rotary positions, QK-norm, squared ReLU) and a shared 32k tokenizer, released with their evaluation sets and the exact list of training documents. Collection.
| Model | Parameters | Trained on | Result |
|---|---|---|---|
| mume-english-125m | 134M | 1.33B FineWeb tokens | 1.106 bits per byte on frozen FineWeb val; on the validation stream 1.132 vs 1.176 for a plain-AdamW GPT-2 124M at the same tokens (3.7% lower) |
| mume-math-125m | 134M | 1.33B OpenWebMath tokens | MATH test (clean) 62.6% fewer bits than the English model; branch sft is a GSM8K/MATH fine-tune |
The corpus records keep their source licences. Model weights and the tokenizer are released under CC BY-NC 4.0 for non-commercial research use, and the modelling code under Apache-2.0. For other uses, contact us.
Links: muse-mesh.com · GitHub · kushal@muse-mesh.com