Article

Sarvam cuts 105B model prices, lowering chatbot running costs

Sarvam’s new 105B rates cut input costs more sharply than output. A worked example shows why chatbot savings depend on usage, not the largest discount.

Editorial illustration for Sarvam cuts 105B model prices, lowering chatbot running costs: a geometric block represents a model release. Not documentary evidence.

Sarvam announced lower prices for its hosted 105B AI models on October 10, 2026, reducing list-rate costs for businesses using them in customer chat and document-text analysis. The biggest reductions apply to text sent to the model; generating answers gets a smaller discount.

Both Sarvam 105B and Sarvam 105B Chat are listed as available at the same new rates on the current pricing page. These are usage charges for software accessing the hosted models, not a consumer chatbot subscription.

What the new prices cover

Tokens are small units of text a model processes or generates. Input includes the question and supplied context; cached input is material whose earlier processing can be reused at a discounted rate.

Per million tokens

Previous rate

New rate

Calculated cut

Uncached input

₹29.28

₹15

48.8%

Cached input

₹10.98

₹5

54.5%

Output

₹73.20

₹60

18.0%

All rates are in Indian rupees. The comparison uses Sarvam’s previous tariff, effective August 5, and its current rate card, checked on October 10. Percentage reductions are calculated from those published prices.

Output includes reasoning tokens—the model’s intermediate processing—as well as the visible answer, according to Sarvam’s billing guidance. A short displayed reply therefore does not necessarily mean a small output charge.

An illustrative bill falls from ₹366 to ₹210

Suppose an application uses 10 million uncached input tokens and one million billable output tokens across a month. Keeping the model and usage unchanged, the published rates give:

  • Previous model charge: (10 × ₹29.28) + (1 × ₹73.20) = ₹366.

  • New model charge: (10 × ₹15) + (1 × ₹60) = ₹210.

That is ₹156 less, or a 42.6% reduction. With one million uncached input tokens and one million output tokens instead, the same comparison is ₹102.48 versus ₹75—a 26.8% reduction. The smaller saving reflects the greater share of spending on output, which received the smallest price cut.

Neither example is a customer invoice or a measured saving. Both assume unchanged model settings and token volumes, no cache hits, and public list prices. They exclude taxes, credits, negotiated discounts and all other application costs, such as speech services, document extraction, hosting and human review.

How to revise a chatbot budget

Use recorded usage to separate uncached input, actual cached input and all billable output, including reasoning. Sarvam’s API usage reference identifies cached tokens within the prompt total: do not count them again as uncached input. Count repeat requests and retries, and do not assume every repeated prompt qualifies for the cache discount.

Then add the rest of the application budget. Sarvam prices speech and document services separately. If those and other costs stay fixed, the percentage saving on the complete application will be smaller than the model-only saving.

For the related question of how workflow design affects spending, RohitAI’s analysis of Asana’s browser-agent cost study examines the difference between cheaper model runs and usable finished work.

The pricing notice does not establish better answers or faster responses. Existing customers can reprice their usage; teams considering a switch still need to assess quality on their own tasks. The notice says the new rates apply now but does not specify the exact billing-switch time or how negotiated contracts are treated.

Methodology: AI-assisted reporting and analysis of Sarvam’s published changelog, rate card and billing documentation. Calculations use public list rates, checked on October 10, 2026. No model testing or customer-invoice review was performed.