A 110M-parameter Malayalam language model trained from Malayalam text, with a tokenizer built for the script rather than adapted to it. Weights, tokenizer and training recipe are published in full.
32,000 tokens of byte-level BPE trained on Malayalam alone, at 1.36 chars/token — so more real Malayalam fits inside a 512-token window.
1.9M documents of native Malayalam. Deduplicated, filtered by script, with every 200th document held out for validation.
A decoder-only transformer in the GPT-2 shape — 12 layers, 12 heads, 768 dimensions. Deliberately conventional.
1.39B tokens to a held-out loss of 1.189, then instruction tuning at a tenth of the learning rate.
Bits-per-character, not perplexity — it is the only measure that survives a change of tokenizer. Published in full, failures included.
| Release | Language | Parameters | Context | Date | Weights |
|---|---|---|---|---|---|
| Gargi-M1-Instruct | Malayalam | 110M | 512 | Aug 2026 | Hugging Face |
| Gargi-M1-Base | Malayalam | 110M | 512 | Aug 2026 | Hugging Face |
| Gargi-T1 | Tamil | — | — | 2027 | Planned |
| Gargi-K1 | Kannada | — | — | 2027 | Collecting |