← The notebook
Volume № 01 · § 08 · filed 02 Aug 2026

India doesn't need to catch up

Everyone keeps asking whether India can build a model that beats the frontier. I think that's the wrong question, and the right one has an answer India can afford.

24 minpositionindiadefencecomputepolicy
Yearly rise in the cost of keeping up3.5×Fig. 1
Of the AI mission's money that has gone out3.9%§ 03
Of the mission that pays for checking models0.2%§ 07

I've spent a few months reading everything India has published about AI strategy. Mission documents, budget lines, committee reports, speeches. One sentence keeps showing up, wearing slightly different clothes each time: we will build a model to rival the best in the world.

Every time I read it I want to ask the same thing back. Do you know what that costs next year?

Because I looked it up, and the answer turns out to be the whole strategy. Training a frontier model gets about 3.5 times more expensive every year. That's measured. Meanwhile the cost of hitting any fixed level of capability falls about four times a year, because the training tricks keep getting better and the chips keep getting cheaper per unit of work.

Two rates. One climbing fast, one falling almost as fast. Put them on the same chart and the strategy picks itself.

§ 01

The two curves

cost of one training run, per year
$1.8B$900M$0
2023202420252026202720282029
Keeping up with the frontier$1.8BStaying twelve months behind it$23M
Fig. 1An argument built from measured rates, so treat the climbing line as a projection. Both lines leave the same point in 2023 and then follow the two rates. The climb is drawn to Epoch’s own bent projection, which passes $1B in 2027, rather than the raw 3.5×, which would put 2026 at $3.7B and overshoot that projection about fourfold.

Here's the receipt for the falling line, and it's a couple of years old now because that's how long it takes to get clean numbers. GPT-4 cost between $41 and $78 million to train in early 2023. Twenty-one months later DeepSeek-V3 landed in the same class for a final run of $5.58 million. Same capability. About a fourteenth of the price. Nobody had to invent anything magical to get there.

Divide the two rates and you get 1.17, which means holding a constant gap behind the frontier gets about 15% cheaper every year. Cheaper. In nominal rupees, while the thing you're chasing gets 3.5 times harder to reach.

Catching up is a treadmill that speeds up. Staying twelve months behind is a walk.

So what does a fixed distance behind actually buy you, today, in August 2026? More than most people in this conversation seem to think.

Kimi K3 put its weights on the internet on 27 July: 2.8 trillion parameters, mixture of experts, and it landed around third on the Artificial Analysis index. Not third among open models. Third. DeepSeek's V4 Flash line is working the cheap-and-fast end of the same tier, which is arguably the more useful end if you're a government paying per token. Epoch measures the open pack at roughly four months behind the closed frontier. Practitioners put it at three to five since K3 shipped. Coding is basically closed at this point, and what's left of the gap lives in the hardest reasoning benchmarks.

3 to 5 months24 months
0 months30 months behind
The dark span is where open weights sit today: Epoch measures about four months on its capability index, and Nathan Lambert put it at three to five after Kimi K3. Past the ember line, proximity stops being something you can buy and this whole note falls over.
Fig. 2Where the open pack sits against the closed frontier, and the line past which everything in this note stops working. The gap has been stable to slightly widening since 2023, which is the single most decision-relevant fact here.

Now look at which Chinese models are open and which aren't. Qwen 3.6 Plus: closed. MiniMax's top tier: closed. Xiaomi's MiMo-V2-Pro, the single highest-volume model on OpenRouter last quarter: closed. China open-sources the tier below its own frontier and keeps the frontier itself. That's a deliberate strategy and it works, which is exactly why Singapore and Malaysia have both built sovereign AI programmes on Qwen without anyone quite deciding to.

One caveat on my own chart, because I'd rather say it than have you find it. The climbing line is a projection and it's already bending. If you take 3.5× literally from GPT-4 to now you land on $3.7 billion for the largest current run, which overshoots Epoch's own number by about four times. Capital and power are starting to bite. That bend is good news for India, since it drags the frontier down toward the affordable tier faster than the raw trend says.

And here's the part I'd argue with people about. Waiting is cheap in money and ruinous in everything else. Start in 2029 and you'll pay less for a 2026-class model. You'll also have no team, no tooling, and no record of the things that didn't work. Models get cheaper every year. The eight engineers who know why the last run fell over do not.

§ 02

Nobody's talking about the actual bill

Everyone treats the training run as the thing that costs money. For a government it's the cheap half. Training happens once. Serving happens every day, at the scale of the population, for as long as the country exists, and the newer models think harder per question, so each answer costs more than last year's did.

Try it on one real workload. One frontier-model reasoning pass over a year of UPI transactions comes to roughly ₹15,100 crore. One workload. One year. About one and a half times the entire five-year national AI mission. I ran that twice because I didn't believe it the first time.

And UPI isn't some weird edge case. India's district courts sit on just under five crore pending cases and take in about 2.19 crore new filings a year. Bhashini already handles 15 million inferences a day across 36 text and 23 voice languages. This is just what running a country looks like once the software gets smart.

Now the bit that genuinely annoys me. Around ₹4,590 crore of that UPI number buys no intelligence whatsoever. It gets paid because the text is in Indian scripts.

tokens per unit of text, English = 1
1.4–2.1×
4–8×
7.8×
EnglishIndic tokenizerTypical multilingualOdia on LLaMA-4
Fig. 3How many tokens the same passage costs, relative to English. The marked column is the fix. The measurement on the right is on LLaMA-4's tokenizer, which is a generation old now, and the newer frontier tokenizers have not fixed this.

A tokenizer is the piece of software that chops your text into billable chunks. The ones shipped with the big models were built for English. Feed them Devanagari or Odia and they chop badly, so the same sentence costs four to eight times more to run. Indian teams have already built tokenizers that get this down to about 1.4 to 2.1 times, which is close enough to parity to stop mattering.

It's the least glamorous line item in Indian AI and I'm fairly convinced it's the highest return one. At population scale you're looking at thousands of crore a year. Nobody gets a photo op for fixing a tokenizer, which is probably why it hasn't been fixed.

Owning the model and renting the serving is like owning the car and renting the road.
§ 03

38,231 GPUs is not 38,231 GPUs

India has provisioned 38,231 GPUs. That is a lot of GPUs. They sit across fourteen or fifteen different providers, which is where it stops being a lot of GPUs.

What India has: 38,231 GPUs, spread across 14 to 15 providers
What one flagship training run needs: 8,000 of them, wired together, exclusively, for months
one square = roughly 250 GPUs
Fig. 4Provider-level shares aren't published, so the clumps are drawn evenly. The fleet size and the provider count are the reported figures. The gaps between the clumps are the whole point of the drawing.

Big training runs need thousands of chips on one fabric, talking to each other at full speed, reserved for one job for months. Split the fleet across vendors and the share of each chip you actually use drops from around 35% to somewhere between 15 and 20%, which roughly doubles your bill. Go below about 8,000 chips on a single fabric and you don't get a slow run, you get no run at all, at any budget. This is the one failure mode that more money does not fix.

The pricing has its own version of the problem. The headline is ₹65 per GPU hour, which sounds like an achievement until you look underneath it.

₹65, the quoted average
₹4 an hour₹1,234 an hour
The spread across listed SKUs on the national compute programme.
Fig. 5The published per-hour rates on the national compute programme, and the average everyone quotes. Averaging across SKUs this far apart produces a number, and the number tells you nothing about what your run will cost.

Nobody outside the programme knows how much of that capacity is being used, because utilisation isn't published. From what is public you can derive an allocation ratio somewhere under 3%, and I want to flag that the two figures may not be measuring commensurate things. Which is precisely the argument for publishing the real one. Meanwhile the trade press keeps reporting that the mission can't find enough takers.

Something surprised me while I was doing this. Training is not a power problem for India at all. 8,192 H100s running for 26 days draws about 11.9 MW, against a record peak demand of 270.82 GW. That's 0.0044%. Even a 40,000-GPU national cluster at roughly 58 MW disappears against 532.7 GW of installed capacity. Every time someone tells you India can't afford the electricity for a frontier model, they're wrong by four orders of magnitude.

Inference is where power bites, and it bites hard. AI-capable data centre capacity was about 275 MW in 2025 and needs to reach something like 6,500 MW by 2030: a 24-fold build in five years. Only about 14% of India's current data centre stock can host this hardware at all. Power has already overtaken land and capital as the binding constraint on new sites. Water stress in Bengaluru, Chennai and Gurugram is an under-priced risk nobody wants to talk about, and Maharashtra plus Tamil Nadu hold 65% of installed IT load, which is a concentration problem of its own.

The good news is just as real. Mumbai builds data centre capacity at roughly $6.6 per watt against Tokyo's $15.2 and Singapore's $14.5, which makes India the second most cost-effective large market on earth for this. It's a durable advantage and almost nobody in the sovereignty conversation mentions it, because substations and cooling towers don't trend.

§ 04

The market you're buying from isn't really a market

Go one layer above the GPUs and the supply chain gets narrow fast. TSMC took 70.4% of the top ten foundries' revenue last quarter. CoWoS advanced packaging is sold out with most of it pre-allocated. SK hynix holds about 62% of high bandwidth memory, Micron 21%, Samsung 17%. ASML is the only company on earth that makes an EUV machine. None of those numbers are moving toward diversification.

And every credible alternative to Nvidia belongs to an American hyperscaler: Google's TPU, Amazon's Trainium, Meta's MTIA, Microsoft's Maia. Buy one of those and you've diversified your vendor risk while doing precisely nothing to your American risk. For a foreign government those are the same risk in different packaging.

  1. Oct 2022
    Controls by the unitAmerica starts restricting AI chips by count and capability.
  2. Oct 2023
    Controls by densityThe rule switches to performance per chip, so every faster generation gets caught automatically.
  3. Jan 2025
    Countries get tiersIndia lands in Tier 2, with a 1,700 chip licence-free allowance per company and a 7% single-country cap.
  4. Aug 2025
    Access gets a priceA 15% revenue share on H20 sales to China. Call it what it is: a toll.
  5. Dec 2025
    The toll goes up25% on H200.
  6. 12 May 2026
    The rescission that wasn'tThe GAO finds the withdrawal of the tiering rule procedurally defective. Tier 2, the 1,700 chip allowance and the 7% cap are all still sitting in the Code of Federal Regulations, reactivable without new rulemaking and without a vote.
  7. 12 Jun 2026
    It reaches the APIA frontier lab suspends two models for foreign nationals. Partial restoration on 26 June, for about a hundred vetted partners.
  8. FY 2026
    $30 billion of sovereign AIWhat Nvidia booked under that heading, roughly tripled year on year.
chip controlspricingmodel accessthe market
Fig. 6Four years of the squeeze, sorted by date and coloured by what kind of squeeze it was. The May 2026 entry is the one I'd pin to a wall.

Two details from that list deserve more attention than they get. First, caps written in FLOPs tighten on their own: the same allowance that permits about 50,000 H100-equivalents permits only about 43,500 B200s. Nobody has to write a new rule to squeeze you, better chips just have to exist. Second, that $30 billion of sovereign AI revenue means a serious chunk of the global conversation about sovereignty is a sales funnel. Worth remembering when you read anyone's deck on it. Mine included.

A strategy that can't tell sovereignty from procurement will buy an enormous amount of hardware and very little independence.
§ 05

Defence is where I stop being polite about this

Everything above is an argument about money. This part isn't, and it's why I think the use-don't-build camp is answering a question nobody asked.

Right now most government AI is advisory. It drafts, summarises, translates, triages. Somewhere between 2027 and 2030 it goes operational: agents that file, route, dispatch and decide inside workflows with legal effect. America's CDAO has already put four frontier labs on contracts with $200 million ceilings, explicitly for agentic workflows in national security missions. France is rolling Mistral out across its military for operational decision support. This is happening now.

The moment a model is inside the loop rather than beside it, three things become true that weren't true before.

01

An outage stops being an inconvenience

If the system routes something and the system is gone, the thing doesn't get routed. You now have an operational failure with a foreign company's status page attached to it.

02

A version change becomes a re-accreditation event

Commercial deprecation clocks run at six months' notice for generally available models, three for specialised ones, about two weeks for previews. Indian record-retention obligations run in years and decades. No amount of goodwill reconciles those two numbers, and if you paper over it you'll eventually be unable to reproduce a decision you have to defend in court.

03

A refusal becomes a mission failure that belongs to somebody else

Frontier vendors publish usage policies that prohibit weapons work and grant exceptions to what one of them calls carefully selected government entities, assessed against that company's own judgment of whether you have independent and democratic oversight. Read that again. A private firm in another country decides whether your armed forces qualify for an exemption from its rules.

If it can refuse you, you're renting it.

This isn't a fringe view, either. Five Eyes joint guidance already tells its own governments to keep model weights in a protected storage vault inside a highly restricted zone. The UK Ministry of Defence's JSP 936 already requires suppliers to show data provenance and architecture traceable to requirements. Allied rules, for allied systems, and every one of them quietly assumes you hold the weights yourself.

India already has the instrument for this. DRDO published its Evaluating Trustworthy AI framework in 2024 and it simply isn't binding on procurement. Making it binding costs a signature. I've read around this for months and I still can't work out what the holdup is.

One more thing on defence, and it's the least exciting paragraph in this whole piece. The stated primary constraint on Indian military AI has nothing to do with models. The services run separate legacy systems that don't talk to each other, which blocks joint command and control. No model fixes that. An interoperable defence data layer is the prerequisite, it's deeply unglamorous, and it's where the first tranche of money should go. Buying a model before fixing the plumbing gets you a tap for a house with no pipes.

And on Chinese open weights, since it comes up every single time. You don't need to claim anyone has slipped a backdoor into a shipped checkpoint. There's no public evidence for that and asserting it loses you the argument in the first thirty seconds. The verified case stands on its own: inserting a behaviour is cheap and doesn't get harder at scale, safety fine-tuning doesn't reliably remove it, inspecting weights doesn't reliably detect it, one politically conditioned code-security defect has already been measured in a shipped model, and Chinese law makes a developer's assurances unverifiable. For a weapons system, being unable to check is disqualifying all by itself. Much narrower claim than the one people usually make. Unlike the loud version, it holds.

§ 06

Everyone else already ran this experiment

The one advantage of being late is that other countries have run your experiment for you, in public, with receipts. So look at what the chips actually bought them.

accelerators used, or held
EuroLLM-22Bone grant-funded team
Shipped. Europe's best fully open model
400
Mistral Large 3about 40 people
Shipped. 675B, Apache 2.0
3,000
JUPITER, Jülicheurope's fastest machine
No flagship model
24,000
India's national fleet14 to 15 providers
No flagship run yet
38,231
010,00020,00030,00040,000
Fig. 7Fleet size against whether a flagship model came out of it, sorted by fleet. The sort is what makes the pattern impossible to miss.

The two that shipped had the fewest chips. That's not a fluke of my sample either. Europe announced €200 billion for AI, and eighteen months into a €52 million flagship consortium with fifteen-plus partners it has produced a data catalogue, some pretraining code and a judging framework. No model. Its AI gigafactories go to tender this summer and won't produce a token before 2028. Europe's best fully open model, meanwhile, was trained by one grant-funded team on 400 GPUs, on a different machine entirely.

Mistral makes the same point from the winning side. Large 3 is a 675B mixture-of-experts model under Apache 2.0, trained on roughly 3,000 H200s, and the team says the binding constraint was how many good people they had. The moat that actually protects them is a framework agreement with the French armed forces. Guaranteed demand from their own government, signed, while American rivals were taking public heat for defence work.

The rest of the record is blunter. Aleph Alpha, Germany's national champion, concluded it couldn't compete on models and joined forces with a Canadian company. Silo AI ended up inside AMD. Korea, which runs the most disciplined version of this programme anywhere, cut two of five funded teams at the first six-month gate, and cut one of them specifically for building on a frozen Alibaba vision component. The UAE bought its way in and came out more dependent on America than it went in, which is the cautionary tale I'd tape to the wall of any Indian procurement office.

So the seat is empty. And India walks up to it holding things nobody else has: 38,000 GPUs already bought, a from-scratch 105B Apache-2.0 model already out of the door, 22 scheduled languages nobody else will ever serve properly, and the largest annotation workforce on the planet. Europe runs on American cloud. The Gulf runs on American export licences that can be pulled. Neither of them can credibly host a model that's genuinely non-aligned. India can, and right now that position is going spare.

§ 07

The cheapest thing on the list is the one nobody funded

An independent body that tests models and publishes what it finds costs under $5 million a year to run. India funds this at ₹20.46 crore: 0.2% of the mission, and less than a fifth of what the mission spends on its own administration.

It shows in the obvious places. India hosted the largest AI summit ever held, in February 2026. Ten months later the international body that came out of that moment lists Australia, Canada, the EU, France, Japan, Kenya, South Korea, Singapore, the UK and the US as members. India isn't in it. India convened the room and then didn't join the institution the room produced.

At home the problem is tangled rather than empty. The benchmarks used to judge Indian language models were built by AI4Bharat. AI4Bharat's founders went on to found Sarvam, whose models now run air-gapped inside UIDAI and serve 80 million SBI Life customers, and whose benchmark results are self-reported on an evaluation suite the company designed, translated and judged itself.

I want to be careful here, because this reads as an accusation and I don't mean it as one. Nobody acted badly. There's simply no scoreboard, so everyone marks their own homework, and good faith cannot fix a governance defect. It's a vacancy, and vacancies get filled by whoever shows up.

Whatever India builds without an independent scoreboard will be contestable. And it will get contested, by people who aren't friendly.

Korea shows the alternative and there's nothing exotic about it: 40% benchmarks, 35% an expert panel, 25% user feedback, administered by two agencies that don't build models. That's the whole trick. The scorer isn't the builder.

The cost of not having one is already on the books. ₹1,058 crore of public money produced a 17-billion-parameter model, released under a licence that forbids commercial use, with a context window of 4,096 tokens. No engineer failed here. Somebody signed the wrong licence, and a one-page procurement standard would have caught it before the money left the building.

§ 08

What I'd actually do, and how you'd know I was wrong

Five things, in this order, because the first three unblock the rest. None of them is a moonshot and the first two are nearly free.

01

Stand up an evaluation body that doesn't build models

Under $5 million a year. It tests everything going into a government system, publishes what it finds, and is barred from grading anything its own people have a stake in. First three jobs: re-test the Indian models already in production, audit Chinese open weights on India's own territorial questions using the actual weights instead of a phone app, and build an Indian language benchmark that nobody selling a model had a hand in.

02

Write down what a sovereign model means, then make it a condition of purchase

One document. Declare the base model you fine-tuned from and its licence. Ship weights in a format that can't execute code when loaded. Publicly funded models go out under a permissive licence. Anything with legal effect stays reproducible for as long as the record has to be kept. Today nothing would stop a ministry shipping a fine-tune of a foreign model without anyone finding out, and Korea cut a national champion for exactly that.

03

Put the chips in one place

One fabric, at least 8,000 accelerators, wired together, with guaranteed exclusive time measured in months. Until that exists no flagship run is possible, however many GPUs the press release counts. Everything else on this list can wait a quarter. This one can't, because procurement cycles are long and the chips you order today arrive next year.

04

Fix the plumbing before asking for more money

₹400.94 crore released against ₹10,371.92 crore sanctioned, at roughly 40% of the way through the mission. Spending it evenly would need about ₹2,075 crore a year and the real rate is a fifth of that. Multi-year authority, a fund that doesn't lapse every March, one named person accountable for getting money out of the door. That does more in eighteen months than doubling the allocation.

05

Turn the coalitions India already founded into actual programmes

The Trusted AI Commons has 22 member countries, a mandate covering benchmarks and shared tooling, and no budget line. The India-Japan statement of 2 July 2026 commits both governments in writing to co-develop open multilingual models, and it's the only agreement of its kind anywhere. It has no implementing body. Both are signed. Funding them is the cheapest strategic win on the table and I genuinely cannot work out what's stopping it.

All of it, including inference capacity and power, comes to somewhere between ₹30,000 and ₹47,000 crore over five years. Around one percent of a single year's defence budget, annually, for the thing that underwrites every other capability. The model programme, the part that sounds most ambitious, is the smallest line on the list after diplomacy.

So how would you know if any of it worked? One test. If every foreign model became unavailable to India tomorrow, nothing the state does should stop. Degrade, sure. Stop, no. Everything else on the list is an input to that one sentence.

And here's what would prove me wrong, which I'd rather say myself. The whole argument rests on open weights tracking the closed frontier, which they've done since 2023 and were still doing when Kimi K3 shipped in July. If that gap widens past two years and stays there, proximity stops being something you can buy, and India’s real options shrink to negotiating the best access terms it can get. Go back and look at Fig. 2. That’s the number I’d put on a dashboard and check every quarter. As far as I can tell almost nobody in Indian policy is watching it at all.

Forget winning. The target is being impossible to cut off.

watchglass (2026). India doesn't need to catch up. The notebook, volume № 01, § 08. watchglass.app.