An H100 hour in India costs anywhere from ₹70 to ₹1,088 depending entirely on who you buy it from. That fifteen-fold spread is the most quoted fact about GPU as a Service in India, and it is the least important one. The number that decides whether the business works is utilisation — and roughly 83 per cent of enterprises running their own GPUs report it at half or less.
Table of Contents
- What an H100 Hour Actually Costs in India
- Why the Spread Is Not the Story
- GPU as a Service Lives or Dies on Utilisation
- The Fractional SKU Gap
- How Indian Providers Actually Bill
- Three Billing Models, One Reconciliation Problem
- The Indian Invoicing Constraints That Shape the Product
- What Good Looks Like
- Where Hybr® Fits
- Frequently Asked Questions
- References

What an H100 Hour Actually Costs in India
Rates for GPU as a Service in India span roughly fifteen times on identical silicon. The IndiaAI portal’s subsidised twelve-month reserved rate works out near ₹70 per GPU-hour. Google Cloud’s Mumbai on-demand rate is about ₹1,088. Every figure below comes from a live published price card read on 7 September 2026. Rate cards in this market move quickly — treat the shape of the spread as the finding, not any individual number.
The subsidised lane is a policy instrument rather than a market price. The IndiaAI Mission empanelled providers through competitive bidding, discovering an L1 rate card of 53 configurations — an H100 SXM at ₹153 on demand and ₹134.10 at commitment — and then applies a subsidy of up to 40 per cent, paid by IndiaAI directly to the provider. The headline “₹65 per GPU-hour” that appears in government backgrounders is best read as an entry price for the cheapest subsidised configurations, not a blended H100 rate; the arithmetic only reconciles against the earlier tranche-one base rate.
Access to that lane is also narrow. It is open to DPIIT-registered startups and MSMEs, academic researchers meeting citation thresholds, students, government entities and approved public projects — with allocations valid for thirty days to commence and twenty-three days to consume. It is not a commercial supply route for an enterprise or a service provider.
Why the Spread Is Not the Story
Indian providers undercut hyperscaler on-demand pricing by roughly three to one. They do not undercut hyperscaler committed pricing at all — and committed is what a serious buyer signs.
AWS Mumbai’s three-year all-upfront reserved rate for a p5 instance lands near ₹293 per GPU-hour. Google’s Mumbai three-year committed-use rate is around ₹349. Against that, Yotta’s on-demand rate is ₹356 and its one-month committed rate is ₹263; E2E’s on-demand rate is ₹274. The Indian lane wins comfortably against a customer paying list, and wins narrowly or not at all against a customer willing to commit for three years.
Which means price alone is not a durable position in this market. What is durable is everything a hyperscaler three-year reservation does not give you: rupee billing without foreign-exchange exposure at a rate that has moved to ₹94.49 to the dollar, a residency posture you can evidence per tenant, a channel that can resell you, and terms that flex in months rather than years. Those are commercial and platform capabilities, not price-list entries.
GPU as a Service Lives or Dies on Utilisation
A GPU costs the same whether it runs or not. At a ₹274 sell price, an owned GPU returns about ₹137 per owned hour at 50 per cent utilisation — below the roughly ₹160 needed to cover depreciation and interest alone, before colocation, power or a single operations salary.
The evidence on where operators actually sit is uncomfortable. A survey of 573 technical leaders fielded in June 2026 found that about 83 per cent of enterprises running their own GPUs report utilisation of 50 per cent or less, and only 44 per cent rigorously track what their AI compute costs against what it returns. Datadog’s field CTO for the region put it plainly for the Indian market: the constraint is not capacity, it is utilisation, and the first failure point is almost always cost attribution.
Depreciation policy compounds it. Published neocloud schedules for identical hardware range from four to six years — a fifty per cent swing in annual depreciation on the same GPU. Two operators with the same silicon, the same price card and the same utilisation can report materially different margins on depreciation policy alone, before any difference in how they run the estate.
There is a working counter-example in the Indian listed market. E2E Networks reported Q1 FY27 revenue of ₹1,568 million on roughly 5,100 deployed GPUs, with a 75.2 per cent EBITDA margin. Dividing that revenue by the GPU-hours available in the quarter gives an implied blended realisation of roughly ₹141 per GPU-hour across the whole fleet. That is our own arithmetic on published figures, not a company disclosure, and it sits well below the listed H100 on-demand rate — which is what a mixed fleet at real-world utilisation looks like, and why a top-SKU rate card tells you little about a business.
The Fractional SKU Gap
Checking the published rate cards of Yotta, E2E Networks, Jarvislabs, Neysa, Cyfuture, Tata Communications and AceCloud on 7 September 2026, the smallest listed unit at each was a whole GPU; we found no published sub-GPU SKU. Private or unlisted arrangements may well exist. What the public catalogues show is a billing gap rather than a hardware one.
The hardware supports partitioning. NVIDIA’s Multi-Instance GPU gives seven hardware-isolated instances on A100, H100, H200, GB200 and B200, and four on A30 and the RTX PRO 6000 Blackwell. NVIDIA’s own documentation makes the multi-tenant argument explicitly: for cloud service providers, MIG ensures one client cannot impact the work or scheduling of other clients, with separate paths through crossbar ports, L2 cache banks, memory controllers and DRAM address buses.
There is one significant catch that shapes the Indian market specifically. L40S and L4 do not support MIG at all — they do not appear in NVIDIA’s supported-GPU table — and L40S is the volume SKU at Yotta, Cyfuture, E2E, Neysa and Tata. Providers selling L40S multi-tenant are using time-slicing or vGPU, not hardware partitioning, and the isolation guarantees are correspondingly weaker.
| GPU | Maximum MIG instances | Notes |
|---|---|---|
| A100 40GB / 80GB | 7 | Ampere, the original MIG platform |
| H100 80GB / 94GB, H200 141GB | 7 | Profiles down to 1g.10gb-class slices |
| B200 180GB, GB200 186GB | 7 | Blackwell; 1g.23gb through 7g.180gb |
| A30 24GB | 4 | |
| RTX PRO 6000 Blackwell 96GB | 4 | 1g.24gb / 2g.48gb / 4g.96gb |
| L40S, L4 | Not supported | Time-slicing or vGPU only — and these are India’s volume SKUs |
The commercial consequence is direct. An inference workload that needs a quarter of an H100 has to rent a whole one, which prices out the exact mid-market and startup demand that Indian providers are chasing — and leaves three-quarters of a card idle inside a paid-for hour. Government data on its own compute programme shows the mismatch: training accounts for around 80 per cent of GPU-hours consumed, while more than 60 per cent of startups requesting access want GPUs for inference. Inference is the fractional workload. Nobody is selling it fractionally.
How Indian Providers Actually Bill
Billing granularity, currency and commitment depth vary more between Indian providers than the headline rates do — and those choices, not the rate card, are what make a service resellable.
| Provider | Currency | Granularity | Deepest published commit discount |
|---|---|---|---|
| Yotta Shakti Cloud | INR | Hourly | −26% at one month |
| E2E Networks | USD | Per-minute | −14% at monthly |
| Jarvislabs | USD | Per-minute, pause/resume | −9% at twelve months |
| Neysa | USD | Hourly | −51% at thirty-six months |
| Cyfuture AI | INR | Per-second, plus per-token serverless | −33% to −51% by SKU |
| Tata Communications | INR | Hourly | On request |
| AceCloud | INR | Monthly blocks only | −10% at twelve months |
| Sarvam AI | INR | Per million tokens, per audio hour, per page | Prepaid credits |
Two things stand out. First, the commitment ladder ranges from nine per cent to fifty-one per cent, which is a wide spread for the same underlying asset — a fifty-one per cent discount for thirty-six months implies a very different view of capital recovery than nine per cent for twelve. Second, the market is split on currency. Several Indian providers publish in rupees and several in dollars, so an Indian buyer’s exposure to a rate that has moved to ₹94.49 to the dollar depends on which provider they pick rather than on where the GPU sits. Currency of listing is a product decision, and for a domestic buyer it is a material one.
Sarvam is the clearest counter-example and worth studying: pure consumption billing, denominated in rupees, per million tokens for text, per hour of audio, per ten thousand characters for speech, per page for document work. That is what a token-metered AI service looks like when it is priced for an Indian buyer. We covered the general mechanics in AI billing and chargeback for enterprise.
Three Billing Models, One Reconciliation Problem
Capacity is sold three ways — prepaid capacity commitments, on-demand per-minute, and prepaid tokens or credits — and each recognises revenue differently. Getting that wrong is not a billing inconvenience; it is an audit finding.
- Prepaid capacity commitments. Multi-year, cash up front, revenue recognised as capacity is delivered. This is the model behind CoreWeave’s $99.4 billion revenue backlog and $8.2 billion of deferred revenue at Q1 2026 — numbers that only work if you can prove what was delivered, period by period.
- On-demand, per-minute or per-second. Revenue as usage occurs. Simple to recognise, brutal on utilisation, and the model most exposed to idle capacity.
- Prepaid tokens or credits. Revenue recognised as tokens are consumed, which requires a token meter that finance trusts as much as engineering does.
All three need the same three numbers reconciled: what was committed under the contract, what was delivered as provisioned capacity per period, and what was drawn down against the commitment. Two of those three come from a meter. Manual approaches — joining usage exports to a price book in a spreadsheet — slow invoicing and fail during audits and financing rounds, which is precisely when an infrastructure business is least able to afford it. The mechanics of doing this properly are in usage-based and metered billing.

The Indian Invoicing Constraints That Shape the Product
Indian tax and invoicing rules are not an afterthought bolted onto GPU as a Service. They constrain the product design.
- GST at 18 per cent on cloud and hosting services under SAC 998315, unchanged through the September 2025 rate reform.
- E-invoicing is mandatory above ₹5 crore of aggregate annual turnover, and once you cross it the obligation is permanent even if turnover later falls. Above ₹10 crore, invoices must be reported to the Invoice Registration Portal within 30 days of the invoice date.
- Place of supply follows the recipient’s location and determines IGST versus CGST plus SGST — which matters when the GPU is in one state and the customer in another.
- Reverse charge applies when an Indian business buys GPU capacity from a foreign provider: 18 per cent IGST self-assessed, creditable as input tax credit.
- Export of services zero-rating requires payment in convertible foreign exchange, with a Letter of Undertaking and FIRC documentation. Until the currency lands and is documented, the supply has not qualified as an export.
- The equalisation levies are gone — the 2 per cent from August 2024 and the 6 per cent from April 2025 — replaced by the Significant Economic Presence test.
The practical consequence: an invoice for Indian GPU capacity has to carry the right SAC code, the right place of supply, the right tax treatment for a domestic versus export customer, and reach the IRP inside the reporting window — every month, per tenant, automatically. That is a rating-and-invoicing capability, and it is why pricing and rate management and invoicing belong in the same system as the meter.
What Good Looks Like
Five capabilities separate a GPU estate that earns margin from one that rents raw hours.
- Fractional and time-sliced SKUs in the catalogue. MIG profiles on the hardware that supports them, time-slicing where it does not, sold as priced units rather than left as an engineering capability nobody can buy.
- Per-tenant metering of the units you sell — GPU-hours, tokens, storage, egress — continuously, so utilisation is a number you watch rather than one you discover.
- Commit-and-drawdown tracking so committed, delivered and consumed reconcile without a spreadsheet.
- Rupee rate cards with reseller margin, because the Indian channel is how mid-market demand is reached. See reseller management.
- Compliant invoicing per tenant, and a residency posture that does not force a separate cluster per regulated customer — the fastest way to destroy utilisation is to run five half-idle estates, which is the trap set out in AI data residency in India.
Watch the two-minute version. The same argument in video, including what per-tenant metering and reseller pricing look like in a working console. See how the service layer works →
Where Hybr® Fits
Hybr® is the commerce and metering layer for GPU as a Service in India, not another GPU cloud. It registers the Kubernetes and GPU clusters you already run rather than provisioning them, connects to the LLM gateway you already operate rather than hosting models, meters GPU-hours and tokens per tenant, rates them against that customer’s rupee price list and reseller tier, and puts them on one invoice alongside Microsoft CSP seats and managed services.
The useful first test is small: take the SKU you sell most, and check whether you can currently state its realised revenue per owned GPU-hour last month. If that number requires a spreadsheet and a week, the problem is not your price card. Start at Kubernetes and GPU-as-a-Service, or see the market context in India’s AI data centre build-out.
Frequently Asked Questions
What does GPU as a Service cost in India in 2026?
H100-class rates run from about ₹70 per GPU-hour on the IndiaAI Mission’s subsidised twelve-month reserved lane to about ₹1,088 on Google Cloud Mumbai on-demand. Indian commercial providers cluster between ₹200 and ₹415 depending on commitment depth, against AWS Mumbai on-demand at roughly ₹780. All figures from published rate cards, September 2026.
Is GPU as a Service in India cheaper than the hyperscalers?
Against list on-demand pricing, yes — by roughly three to one. Against three-year committed hyperscaler pricing, no. AWS Mumbai’s three-year all-upfront reserved rate of about ₹293 per GPU-hour sits below several Indian on-demand rates. Price alone is therefore not a durable position; rupee billing, evidenced residency, channel reach and flexible terms are.
Why does utilisation matter more than price for GPU as a Service?
Because the GPU costs the same whether it runs or not. At a ₹274 sell price, an owned GPU returns roughly ₹137 per owned hour at 50 per cent utilisation, below the approximately ₹160 needed to cover depreciation and interest alone before colocation, power and operations. Around 83 per cent of enterprises running their own GPUs report 50 per cent utilisation or less.
Can you sell part of a GPU in India?
Technically yes, commercially no. NVIDIA’s Multi-Instance GPU supports seven hardware-isolated instances on A100, H100, H200 and B200, but L40S and L4 — India’s volume SKUs — do not support MIG at all. And on the published rate cards checked in September 2026, the smallest listed unit at Yotta, E2E, Jarvislabs, Neysa, Cyfuture, Tata and AceCloud was one whole GPU.
How do Indian GPU providers bill for consumption?
Granularity ranges from per-second at Cyfuture through per-minute at E2E and Jarvislabs to monthly blocks only at AceCloud. Currency is split: Yotta, Cyfuture, Tata and AceCloud publish in rupees; E2E, Jarvislabs and Neysa publish in dollars. Commitment discounts range from 9 per cent at twelve months to 51 per cent at thirty-six.
What Indian tax rules apply to GPU as a Service invoices?
GST at 18 per cent under SAC 998315, e-invoicing mandatory above ₹5 crore of aggregate annual turnover with a 30-day IRP reporting window above ₹10 crore, place of supply determined by the recipient’s location, and reverse charge on capacity bought from foreign providers. Export zero-rating requires payment in convertible foreign exchange with LUT and FIRC documentation.
References
- Press Information Bureau — India’s common compute capacity crosses 34,000 GPUs, with the full L1 rate card
- Yotta Shakti Cloud — published pricing
- E2E Networks — published pricing
- NVIDIA — Multi-Instance GPU supported GPUs and profiles
- NVIDIA — MIG introduction and the multi-tenant isolation argument
- VentureBeat Research — enterprise GPU utilisation survey, June 2026
- Business Standard — India’s AI compute race and the subsidised GPU capability gap
- Solvimon — billing archetypes and revenue recognition for GPU clouds
- Sarvam AI — rupee-denominated per-token pricing
- GST Council — Notification 10/2023 Central Tax, e-invoicing threshold
- FinOps Foundation — State of FinOps 2026
