Article

Asana reports 76-fold lower model costs in an AI browsing study

Asana’s browser-agent study shows how workflow changes cut estimated model costs—and why businesses need to price complete, usable results.

Editorial illustration for Asana reports 76-fold lower model costs in an AI browsing study: a controlled task flows from input to output. Not documentary evidence.

Asana says it has sharply reduced the model costs of an AI agent that browses websites, in a StackAI report dated October 8, 2026. Highlighted in an OpenAI case study on October 9, the reported 76-fold reduction offers business owners and operations teams a useful example of how automation costs depend on workflow design, not just the model they buy.

It is not a promise of equivalent savings on a customer’s bill. The study estimated model charges for one website task, excluded servers and sandboxes, and found that its cheapest configuration returned a summary instead of the requested table. Asana’s technical report supplies those important qualifications.

Where the reported savings came from

A browser agent is software that visits pages and gathers information on someone’s behalf. In this experiment, it had to collect six fields for each of 32 books in a public demo catalog, then deliver a table, a price total and the cheapest and most expensive books. The study’s methods describe the same task repeated across four models and different settings. The main experiments took place in September, before the October reports.

The company’s published comparison separates improvements to the original model, anonymized as Model B, from the subsequent comparison with OpenAI’s GPT-6.1 Sol:

Configuration

Estimated model cost per run

What the comparison represents

Original Model B setup

At least $36.21

Baseline includes spending on unfinished runs stopped at a limit.

Optimized Model B setup

$1.24

Workflow changes while retaining the original model.

Optimized GPT-6.1 Sol setup

$0.47

A different model and configuration, not a model-price cut.

Dollar amounts are US dollars and are company-reported means of three runs per condition. Asana describes the first reduction as about 29-fold and the additional reduction as 2.6-fold, producing its headline 76-fold comparison. Source: OpenAI’s account of Asana’s results.

The substantial reduction was already present before switching models. Subtracting the published rounded figures gives $34.97 between the original and optimized Model B setups, then $0.77 between optimized Model B and Sol. This arithmetic describes the displayed estimates; it is not a forecast of savings per successfully completed customer task.

The methods and limitations also matter: the cheapest settings were selected after testing twelve configurations per model, and cross-model runs differed in reasoning, tool and cache settings. With three repetitions per main condition, this is evidence about one workflow, not a general model ranking.

Why retaining more history could cost less

The changes focused on retaining browsing history and reusing it between steps. Asana explains that the original agent repeatedly removed screenshots and trimmed older text. Those edits disrupted reuse of earlier processing and could leave the agent revisiting pages.

The relevant mechanism is prompt caching: when the beginning of a new request matches material already processed, the model can reuse that processing. New material still needs work. Editing an earlier part of the request limits how much can be reused.

Asana extended caching to the browsing history, allowed more text to remain, and removed screenshots in batches rather than at every step. The company also reports a warning against a simplistic fix: caching more history without changing frequent screenshot removal increased costs on three of the four models at the larger history budget.

The pricing explains why reuse matters. As checked on October 9, OpenAI’s Sol documentation prices cached input at 5% of ordinary input, while writing it to cache costs 25% more than ordinary input. Paying to cache material is therefore useful only if enough of it is reused later.

That does not make unlimited memory a sensible default. Asana retained limits on steps, tokens and spending, and warns that longer or drifting tasks can behave differently. A buyer should ask for measured results on their own workload, not a universal screenshot-retention setting.

A correct total is not a finished table

The report’s answer evaluation distinguishes encountering the source facts from delivering the requested output. Sol supplied the correct total whenever it answered, but returned a short summary rather than the full table. The authors say final-answer tuning was outside the study’s scope.

For a business needing a reusable catalog, those are different outcomes. A correct total may answer a quick question; it does not provide the rows needed for a spreadsheet or another process. Acceptance should separately cover missing records, factual accuracy and the required format. The report does not establish a cost per fully accepted table.

An illustrative calculation shows why the distinction matters. The $0.77 difference between the two optimized model-cost estimates equals about 92 seconds of work at an assumed $30 an hour: $0.77 ÷ $30 × 3,600. That is not measured review time, and it does not mean Sol needs more review than Model B. It shows how little additional checking or rework would be needed to offset that particular model-cost difference.

What to ask before budgeting for automation

For a pilot involving website research or data collection, turn the study into a few concrete purchasing questions:

  • What exactly counts as finished? For a task like this one, require all 32 rows, all six requested fields, correct arithmetic and a usable table—not just an answer that says the task is complete.

  • Which change produced the improvement? Compare history-management changes on the same model first, then record any different settings when comparing models.

  • What did every attempt cost? Include stopped and failed runs, browser infrastructure, platform charges and human checking or rework. Spread those costs across accepted outputs, not merely runs that returned text.

Asana says the navigation changes have shipped in StackAI. That shipment statement does not establish a $0.47 customer tariff or the exact configuration available under each plan. Request a workload-specific quote and a representative evaluation before treating an experimental model-cost estimate as a budget.

The study used Sol’s Standard service tier. For the separate question of model access and paying for faster responses, see RohitAI’s Sol Ultrafast pricing and access guide.

The practical lesson is to compare the cost of work you can actually use. Asana’s results make a strong case for examining workflow design; the output limitations explain why a cheaper run alone is not enough.

Methodology: AI-assisted reporting and analysis of the company’s published study, its accounts and official documentation. RohitAI did not run or independently replicate the browser-agent experiment. Calculations use the reported rounded estimates and the explicitly stated wage assumption.