About

Language is shared knowledge.

Gargi trains small, highly capable models for Indian languages and releases them openly. Everything we build is meant to be used, copied and improved by anyone — because the knowledge a language carries belongs to the people who speak it, not to whoever happens to hold the weights.

Our vision

Every person reaching computers — and everything computers know — in their own language, through models that belong to everyone.

A model is human knowledge, compressed — trained on what millions of people wrote, spoke and handed down. It was never ours to fence off. Language and the knowledge it carries are held in common, so the models built from them must be too.

What is ours to build — and to earn from — is what sits on top: the tooling, the systems, the specialized ways that knowledge gets put to work. So the order is deliberate: first, models that handle each language superbly; then, the systems built on them. The knowledge stays free. The value lives in what you make with it.

01

Small models, genuinely capable

Small models that handle Indian languages well enough to depend on — trained from the language itself, not translated into it.

02

Open source, permanently

Weights, tokenizers, corpora and evaluation code are released as each is finished, and they stay released.

03

The whole knowledge layer

Not just models, but the infrastructure beneath them — so information becomes reachable in any language people read, write or speak.

04

Free at the core

The models and the language layer stay free; the tooling and systems built around them carry the business.

Built the way open source is built

Gargi develops in the open, in public repositories, with the same review and contribution model that produced most of the software the world runs on. If you work on language, machine learning, data or the languages themselves, there is work here for you.

ResearchersPretraining, tokenization, evaluation design for Indic scripts.
EngineersTraining infrastructure, inference on small hardware, tooling.
LinguistsCorpus curation, annotation standards, dialect coverage.
SpeakersReading model output and telling us where it is wrong.
A billion people should not have to think in English to be understood by a machine.
Work with us