Indian language AI is often framed as a translation problem. Soket AI Labs is taking a broader view: it is building models that can understand, reason, write code and solve complex problems across Indian and other under-represented languages.
Founded in 2019 by Abhishek Upperwal, Soket AI is developing a sovereign, open-source AI stack through Project EKA, a 120 Bn-parameter model focused on Indian and 20+ Global South languages. The company’s central bet is that language coverage will matter only if the underlying models can also perform well in mathematics, coding and reasoning.
Why Indian language AI needs more than translation
For Soket AI, Indian languages are the starting point rather than the final product. The startup is working on a multilingual model that will cover more than 20 languages across the Global South, where digital representation and high-quality datasets remain limited.
The approach is to build for India’s diverse language environment first and then apply those methods to other under-represented language ecosystems. The source does not specify which languages will be included in the model.
This distinction matters because language access alone does not make an AI system useful for enterprise or government applications. Soket AI is also investing in mathematics, coding and structured logic, with the expectation that stronger reasoning models could support sectors such as banking, finance, cybersecurity, defence, law and scientific research.
In February, the startup showed an internal prototype that allowed developers to generate and understand code using Hindi prompts. The demonstration was intended to show that advanced coding tools do not have to remain limited to English speakers.
Project EKA is Soket AI’s larger test
Project EKA is the company’s main model-building effort. Soket AI says it is developing a 120 Bn-parameter, open-source model that will extend beyond Indian languages to serve under-represented languages across the Global South. The model is expected to launch sometime in 2027, according to the company.
The company has not disclosed the exact details of its new architectural approach. It says it spent nearly a year testing different approaches to balance capability and efficiency rather than simply following existing architectures.
Before scaling to Project EKA, Soket AI is building a 24 Bn-parameter model as a proof point for its architecture and data strategy. This could arrive by the end of 2026. The company plans to test the smaller model with a small group of enterprise partners before scaling to the larger model by 2027.
Those partners are expected to help validate the architecture, datasets and performance in real-world environments. Soket AI’s stated objective is not necessarily to match global leaders such as Anthropic on day one.
Instead, it wants to address specific enterprise requirements at significantly lower costs and improve the models through deployment feedback. Both models will be open-sourced, according to the source.
The difficult part is data, not just compute
Building a frontier model requires more than access to compute. Soket AI identifies data, research depth and specialised talent as equally important constraints.
The startup found that public data sources were not consistently strong enough for advanced reasoning tasks. As Upperwal said, “We hit a wall where many of the content in the domain was not sufficient. Data quality was not that great,”
In response, Soket AI designed its own pipelines to collect, clean and structure data for Indian and other low-resource languages. It has also open-sourced portions of its Indic-language data work through Indic Corpus v1, described in the source as a large-scale Indic pre-training dataset.
The company is working with organisations such as ICRISAT, or the International Crops Research Institute for the Semi-Arid Tropics, to build specialised datasets. One focus is agricultural advisory data, which could eventually be used to train systems that generate farming recommendations. The source does not provide details on the dataset’s size or current deployment status.
Why orchestration will shape commercial use
Soket AI is also investing in what Upperwal calls harness engineering. This refers to the layers, workflows and tools that allow AI systems to run reliably in production and connect model capabilities to enterprise applications.
The company’s view is that a foundation model alone is not enough for real-world deployment. The model layer and orchestration stack need to work closely together so that research can translate into commercial-ready systems.
This focus has expanded the company’s hiring needs. Much of the research effort was previously driven by the founding team, but Soket AI has begun hiring specialised researchers and engineers as model development, infrastructure engineering and dataset creation progress in parallel.
The build-out is supported by IndiaAI Mission backing of ₹177 Cr, including ₹162 Cr in compute resources and ₹15 Cr in non-compute support. Soket AI is also looking to raise money and has received some commitments, but the source does not disclose the amount, investors or valuation.
How Soket AI plans to monetise the research
Revenue is not the company’s immediate priority. “We want to do some amazing work on the research front first, and then I think the priority of business on revenue should kickstart,” the founder said.
The company also intends to release model weights and other assets developed during its research process. Those weights would allow developers to run or fine-tune the models for their own applications.
What Soket AI still has to prove
Soket AI’s challenge is not simply to make a model work across Indian languages. It has to show that the model can reason effectively, operate at a cost enterprises can support and generalise beyond India’s language environment.
The 24 Bn-parameter model is the next practical test. Enterprise partners will provide feedback on the architecture, datasets and performance before the company attempts to scale to Project EKA in 2027.
Whether Soket AI can match the capabilities of the world’s leading AI labs remains unresolved. But its approach highlights a broader issue in Indian language AI: language coverage, reasoning capability, data quality and production infrastructure have to be developed together.
The outcome will depend less on the size of a model alone than on whether those pieces work in combination.