Models

Gargi-M1

A 110M-parameter Malayalam language model trained from Malayalam text, with a tokenizer built for the script rather than adapted to it. Weights, tokenizer and training recipe are published in full.

Parameters
110M
Context
512
Training tokens
1.39B
Vocabulary
32,000
License
Apache 2.0
Status
Research preview

Fundamentals

01

Tokenizer

32,000 tokens of byte-level BPE trained on Malayalam alone, at 1.36 chars/token — so more real Malayalam fits inside a 512-token window.

02

Corpus

1.9M documents of native Malayalam. Deduplicated, filtered by script, with every 200th document held out for validation.

03

Architecture

A decoder-only transformer in the GPT-2 shape — 12 layers, 12 heads, 768 dimensions. Deliberately conventional.

04

Training

1.39B tokens to a held-out loss of 1.189, then instruction tuning at a tenth of the learning rate.

05

Evaluation

Bits-per-character, not perplexity — it is the only measure that survives a change of tokenizer. Published in full, failures included.

Releases

ReleaseLanguageParametersContextDateWeights
Gargi-M1-InstructMalayalam110M512Aug 2026Hugging Face
Gargi-M1-BaseMalayalam110M512Aug 2026Hugging Face
Gargi-T1Tamil2027Planned
Gargi-K1Kannada2027Collecting
A billion people should not have to think in English to be understood by a machine.
Research access