Gargi trains small, highly capable models for Indian languages and releases them openly. Everything we build is meant to be used, copied and improved by anyone — because the knowledge a language carries belongs to the people who speak it, not to whoever happens to hold the weights.
Every person reaching computers — and everything computers know — in their own language, through models that belong to everyone.
A model is human knowledge, compressed — trained on what millions of people wrote, spoke and handed down. It was never ours to fence off. Language and the knowledge it carries are held in common, so the models built from them must be too.
What is ours to build — and to earn from — is what sits on top: the tooling, the systems, the specialized ways that knowledge gets put to work. So the order is deliberate: first, models that handle each language superbly; then, the systems built on them. The knowledge stays free. The value lives in what you make with it.
Small models that handle Indian languages well enough to depend on — trained from the language itself, not translated into it.
Weights, tokenizers, corpora and evaluation code are released as each is finished, and they stay released.
Not just models, but the infrastructure beneath them — so information becomes reachable in any language people read, write or speak.
The models and the language layer stay free; the tooling and systems built around them carry the business.
Gargi develops in the open, in public repositories, with the same review and contribution model that produced most of the software the world runs on. If you work on language, machine learning, data or the languages themselves, there is work here for you.