An expense workflow can be mostly correct and still create a mess: the right receipt, the right employee, the wrong reimbursement amount. Cheap clicks help only if the finished record survives a check. That is the useful lens for H Company’s September 28 Holo4 release, a family of models trained to work across screens, code and business tools.
The release brings two Holo4 sizes, downloadable weights, paid API access and public execution traces. The obvious reading is that smaller models are catching up at computer use. The more consequential finding is that the cheapest model to call, the easiest model to self-host commercially, and the strongest model on long workflows are not the same choice. Builders have to select a route, not simply a checkpoint.
RohitAI’s read: Holo4 deserves a serious evaluation as computer-use infrastructure, especially where a workflow crosses an API boundary and falls back to a screen. But its evidence argues for routing by task length, recoverability and license—not appointing one low-cost agent to every job. The public traces make that judgment more concrete, including where they complicate the headline numbers.
Start with the license, then choose the worker
There are three distinct weight-license routes. Calling all of them “open source” erases the first deployment decision.
Model | What it contains | Weight license | Deployment implication |
|---|---|---|---|
35B total parameters; roughly 3B active per token; Qwen3.5 MoE architecture | Apache 2.0 | The straightforward permissively licensed Holo4 self-hosting candidate. | |
27B dense model based on Qwen3.8 | CC BY-NC 4.0 | Downloadable weights do not grant commercial self-hosting rights. A paid hosted route is offered separately. | |
30B-A3B model based on NVIDIA Nemotron 3 Nano Omni | NVIDIA Open Model Agreement | A separate backbone and license; evaluate it as its own option. |
NVIDIA’s agreement permits commercial use subject to its conditions. That is different from both Apache 2.0 and the 27B model’s non-commercial weight terms. H lists both Holo4 sizes on its paid Models API; the API offering does not silently rewrite the downloadable license.
This creates a practical tension. A team may prefer the stronger long-workflow results of 27B while wanting to own its serving infrastructure. Those goals do not resolve into a free commercial weight download. Conversely, choosing Apache-licensed 35B-A3B may improve control over deployment while requiring a narrower task queue. License and capability belong in the same procurement conversation.
That extends our Atria Dawn analysis. Open weights give builders more control over the worker. Holo4 adds a sharper question: which worker can you operate under acceptable terms, for the particular jobs you need done?
The valuable handoff is between interfaces
Imagine reconciling supplier invoices. An API retrieves records, a short program checks totals, and a browser handles a legacy portal with no useful integration. A computer-use model becomes valuable at the handoffs: keeping the supplier identity, invoice number and approval state consistent as the interface changes. This is an illustrative workflow, not a reported Holo4 customer deployment.
H’s hybrid-agent example supplies browser tools and a Python runner to one model. Its smaller demonstration reads dates from a page and computes their difference in code. The application still provides and executes the tools; weights alone do not supply a desktop or integration service.
The builder opportunity is to automate the awkward remainder around existing integrations. A product does not need to replace a reliable API with visual clicking. It needs a worker that can continue when one step has no API, then return to a structured route where the result is easier to verify.
H also reports changes to the surrounding execution loop: better memory handling, a shell on the desktop machine, and larger step and time budgets. Its harness account therefore describes a model-plus-system improvement, not an experiment that isolates the weights. If your wrapper drops state at an interface handoff, the checkpoint cannot inherit H’s results by model name alone.
Our working rule: use the most verifiable interface available, and make every handoff carry the same task identity and constraints.
Long jobs separate models that look close on short ones
H’s launch reports 85.2% on OSWorld for 27B and 80.8% for 35B-A3B: a 4.4-point gap. The more demanding OSWorld 2 results diverge much further. Its benchmark authors describe workflows taking humans a median of about 1.6 hours. The two benchmarks are different task sets, not interchangeable measures of one ability.
The following are H’s published launch figures, not independent measurements. OSWorld 2 was a single run per model. Costs are token-derived charges at H’s API rates, not the full cost of operating a computer.
Evaluation | 27B result | 35B-A3B result | Reported cost per attempt: 27B / 35B |
|---|---|---|---|
OSWorld | 85.2% success | 80.8% success | $0.08 / $0.05 |
OSWorld 2: average partial credit | 61.7% | 30.9% | $1.22 / $0.61 |
OSWorld 2: benchmark-defined pass rate | 41.5% | 12.3% | Same runs as the row above |
61.7% partial credit does not mean 61.7% of workflows passed. Partial credit measures progress against a rubric; a pass is a separate criterion. Neither automatically proves that every requirement of a customer’s business process was satisfied.
Our interpretation is not that 35B-A3B is a bad agent. It is that the gap on long, stateful work is too large to choose it solely from the short-task score or token price. A bounded update with an immediate check and an afternoon-long workflow with changing constraints should not enter the same default queue.
A useful routing feature may therefore be recoverability rather than apparent difficulty. A complicated calculation can be cheap to check and rerun. An easy-looking sequence that creates records across three systems may be costly to repair. Route on the cost of a wrong intermediate state, not just how demanding the opening prompt sounds.
The traces are useful because they do not give one tidy answer
H publishes 7,366 trajectories covering actions, observations, grading information and usage. The dataset is sanitized, with some screenshots replaced and some tasks omitted. This is meaningful inspectability, not an unredacted record of everything.
For this article, we checked the pinned trajectory index. Its 106 exported OSWorld 2 rows per model reproduce the launch arithmetic: 44 passes for 27B and 13 for 35B-A3B. But embedded metadata also gives reviewed aggregates excluding 10 and 8 attempts respectively, under rules covering reference-answer access and forbidden methods.
Those reviewed summaries report approximately 61.91% partial credit and 41.67% success for 27B, versus 30.12% and 12.24% for 35B-A3B. We have not reconciled which review policy should govern the launch table. The calculations below explicitly use the launch figures; they do not splice reviewed scores into launch costs.
This is a reason to ask for a machine-readable evaluation manifest, not to throw away the release. A score should travel with its task version, denominator, permitted actions, budget and exclusion policy. Otherwise two people can read the same public artifact and answer different questions while using the same benchmark name.
AutomationBench supplies another caution. Its maintainers exclude 200 simple tasks from the official 600-task score. H’s trace export includes those simple tasks, so its top-line aggregate is not a substitute for the scored benchmark. H also discloses training-data collection overlap with 480 of the 600 public tasks, with 120 held out. A future official private-set result would answer a different and more useful generalization question.
We inspected documents and trace arithmetic, not a running Holo4 deployment. Launch-day research found no substantive independent replication. Public files make scrutiny possible; they do not turn publisher-run evaluations into independent tests.
Half the attempt price can still buy fewer successful runs
Start with the published API rates, checked on September 28. Both models have a documented 262,144-token context and accept up to five images per request.
Model | Fresh input / million tokens | Cached input / million tokens | Output / million tokens |
|---|---|---|---|
holo4-35b-a3b | $0.30 | $0.03 | $2.00 |
holo4-27b | $0.40 | $0.04 | $3.00 |
Those prices make experiments accessible. They do not settle the cost of useful work. Using only the OSWorld 2 launch numbers above, a simple accounting exercise reverses the apparent winner:
Illustrative inference spend per benchmark-defined pass
= mean spend per attempted task / observed pass fraction
Holo4 27B: $1.22 / 0.415 ≈ $2.94
Holo4 35B-A3B: $0.61 / 0.123 ≈ $4.96This is derived aggregate accounting, not a retry-to-success forecast. It spreads the reported spend across the reported successful runs. It assumes compatible launch cost and pass-rate populations, excludes runtime and human costs, and remains provisional given the review-policy ambiguity. Repeating a failed task does not create an independent draw from the benchmark pass rate.
Nevertheless, the purchasing lesson is concrete. An inexpensive worker that needs frequent intervention can cost more per finished record. A stronger worker can justify a higher inference bill if it avoids an expensive recovery. Measure both on the same queue before designing a “cheap first, expensive second” router: the first attempt may leave damage that the second must diagnose.
For the invoice example, count duplicate records, repair minutes and late escalations alongside model spend. A valid completion message is not an accounting control.
Cache reuse is part of the operating model
H documents automatic, best-effort caching of repeated prompt prefixes on the same model. The cached-token count is returned in usage. Cached input costs one tenth of fresh input for both Holo4 models; output is billed separately.
The OSWorld 2 trace metadata reports roughly 86–88% cached prompt tokens across the two models’ runs. That is cumulative usage over many calls, not a claim that a single request exceeds the context window.
Here is a sensitivity example, not a reconstruction of H’s benchmark bill. One million 27B input tokens would cost $0.40 if all were fresh. With 90% cached, the input charge becomes $0.076: $0.04 for fresh tokens plus $0.036 for cached tokens. Output and runtime charges do not disappear.
Our inference: two wrappers using the same weights can have materially different economics without any difference in intelligence. One preserves stable prefixes and useful memory; the other rebuilds context unnecessarily and pays to process it again. Compaction then becomes a tradeoff between retaining task state, reducing tokens and preserving reusable context—not simply shortening the prompt.
Record cache fraction, total input, output, model time and tool time per finished task. A low list price with a poor cache pattern is a different service from the one suggested by a heavily cached evaluation.
“3B active” does not mean a three-billion-parameter download
The self-hosting story has a different constraint. H offers BF16, FP8, NVFP4 and Q4 GGUF for both Holo4 sizes. Its local-serving guide maps the first two to vLLM, NVFP4 to vLLM on Blackwell, and GGUF to llama.cpp.
Q4_K_M release | Language weights | Vision projector | Combined artifact size |
|---|---|---|---|
16.88 GB | 0.93 GB | 17.80 GB | |
21.30 GB | 0.90 GB | 22.20 GB |
These are observed file sizes in decimal gigabytes, not measured RAM or VRAM requirements. Totals use unrounded bytes. Cache, activations, runtime allocations and concurrency add further requirements.
The MoE model activates fewer parameters per token but carries more total weights than the dense 27B model. That distinction separates potential compute efficiency from storage and memory requirements. Neither the active-parameter label nor the file total proves acceptable step latency on your hardware.
H recommends vLLM 0.28 or newer with qwen3_coder tool parsing and qwen3 reasoning parsing. Crucially, the memory and speed charts further down its local guide cover Holo3.1, not Holo4. Do not turn them into Holo4 performance promises.
As in our GGUF test-bench guide, pin the exact checkpoint, quantization, projector, template and runtime together. A successful load is the start of evaluation. Test tool-call validity, visual grounding and whole-workflow quality before concluding that a smaller numerical format preserves the result.
Build a loop that knows the difference between “answered” and “accepted”
Holo4 uses native function calling; the agent-loop compatibility matrix marks structured outputs unsupported for these two models. Register the answer tool as the model’s completion signal. An empty tool-call result is not that signal. This can be an integration change, not just a model-ID replacement.
The same guide specifies normalized 0–1000 coordinates, scaled to the screenshot actually sent, and recommends retaining at most three screenshots while preserving observation wrappers. Durable facts belong in carried-forward content; replayed reasoning is not the application’s memory. Five-image request capacity and a three-screenshot history recommendation describe different limits.
I would keep an application-owned task record outside that conversation:
Task record
identity: invoice, supplier, destination record
constraints: amount, currency, deadline, permitted changes
checkpoint: last independently verified state
unresolved: missing fields, tool failures, user decisions
acceptance: saved values match; no duplicate; status correctThat record should survive compaction and model escalation. Before committing a change, reconcile it against the actual destination state. After the model calls answer, run the acceptance check. The model may be done speaking while the workflow is still incomplete.
Cross-interface tools also need cross-interface permissions. In the illustrative invoice workflow, a blocked submission must remain blocked through the browser, a Python subprocess and an MCP tool. Otherwise the product has restricted a button, not an action. H’s hybrid cookbook explicitly calls for isolated production code execution. This is a design requirement, not a claim of a discovered Holo4 vulnerability.
Keep the policy attached to the intended resource change. A tool name alone does not describe everything that tool can do. Cancellation and spending limits should terminate the job across all available routes, including outstanding tool work.
The first evaluation should be small enough to diagnose
Do not begin with an undifferentiated pile of enterprise tasks. Build a held-out ladder that separates bounded actions, cross-application state transfer and long workflows. Keep the task set private from prompt tuning, and use the same acceptance criteria across hosted and local routes.
Start with reversible work. Use drafts and test records. Compare the destination state against a fixed specification, including things that must remain untouched.
Introduce one interface handoff. Read through an API, calculate in code, then update a GUI. Check whether identifiers and constraints survive each transition.
Exercise recovery. Inject a stale screenshot, a tool timeout and an ambiguous save result. Verify that the agent checks before repeating a potentially duplicating action.
Extend the horizon. Add a late requirement change and enough work to trigger compaction. Score forgotten constraints separately from clicking errors.
Price the whole result. Record inference, runtime, retries, repair and review time. Report accepted outcomes by task class rather than hiding everything in one average.
Fix step, time and spend budgets before comparing models. Then inspect failures: model misunderstanding, missing tool capability, lost memory, execution error and a bad acceptance check require different fixes. A router trained on an unexplained pass/fail label will inherit that confusion.
The public traces are a useful source of failure hypotheses. They should not become the only tasks on which your implementation is tuned and judged.
Questions builders will ask
Can I commercially self-host every Holo4 model?
No. The weight licenses differ: 35B-A3B uses Apache 2.0, while 27B uses non-commercial terms. Holotron4 has NVIDIA’s separate agreement. Hosted API access is a different route from permission to deploy downloaded weights commercially.
Is Holo4 free on the API?
The quickstart lists Holo4 on the paid tier. Its free-model option is Holo3.1 35B, not Holo4. Download availability and free hosted inference are different claims.
Does the release establish frontier-model parity?
No. H’s comparison notes describe different harnesses, effort settings and task subsets. AutomationBench also mixes public-set scores with competitor costs from a private-set leaderboard. Those figures are useful leads for evaluation, not a controlled universal ranking. See the launch methodology notes.
The next advantage may live in the wrapper
My prediction is that useful Holo4 deployments will specialize before they generalize: bounded local workers where control matters, stronger hosted routes where long-horizon competence pays for itself, and explicit escalation before a failed attempt becomes expensive to unwind. That is a forecast, not an announced H product strategy.
The evidence to watch is equally specific: matched-harness results, a clearly reconciled review policy, independent private-task evaluations and measured local performance. Better screenshots of a demo are less informative than a clean explanation of which jobs finish and what their failures cost.
Holo4 gives builders enough access to investigate those questions. Its promise is a worker that can cross the gaps in existing software. The product decision is where to put that worker, what state to preserve, and how to recognize a finished job without asking the same model to grade itself.
