AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What Makes Holo4 A Platform For Generalist Computer-Use Agents? on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

H Company announced Holo4, a series of open-weight agentic models (27B dense and 35B-A3B MoE) that operate software through GUIs, code, MCP and APIs with a single model. The company reports 61.7% on OSWorld 2.0 for the 27B version, trailing only top closed models, and has released every benchmark trajectory for public audit. All headline numbers are self-reported.

H Company has released Holo4, a new series of open-weight agentic models built to operate software through GUIs, code, MCP and APIs using a single model. The series ships in two sizes — a 27B dense model and a 35B-A3B Mixture of Experts model — available on the H Models API and downloadable from Hugging Face. According to the company, Holo4 27B scores 61.7% on OSWorld 2.0, trailing only the strongest closed models while using far fewer parameters and lower cost.

Holo4 is positioned as a generalist computer-use agent: it clicks and types on screens, writes and runs its own code, and calls MCP or API tools, choosing whichever interface fits the task. According to H Company, most agentic models are trained for a single interface — GUI-focused models fail without a screen, while tool-calling models cannot handle applications without APIs. Holo4 runs on desktops, the web, Android, in a code sandbox and against business APIs, with the same model invoked the same way on each platform.

The models were trained through supervised and reinforcement learning on a large set of environments and tasks, including tasks generated by H Company’s Agentic Task Factory. The company reports that Holo4 improves substantially over its Qwen base model, demonstrated in side-by-side examples on professional software such as FreeCAD 3D modeling and Godot game design, run with the same prompt and harness. Both variants are built on Qwen bases — Qwen3.8 27B for the dense model and Qwen3.6 35B-A3B for the MoE variant, according to the company’s benchmark notes.

On benchmarks, the company reports that on OSWorld 2.0, Holo4 27B scores 61.7% and Holo4 35B-A3B scores 30.9%, compared with 81.8% for Opus 5.5, the strongest closed model in the comparison. On AutomationBench for API use, Holo4 was measured in the company’s internal harness (v1.0.6) against public-set scores for other models. Weights are available in FP16, FP8 and GGUF formats, and H Company has open-sourced every trajectory behind its public benchmark scores, viewable at trajectories.hcompany.ai and downloadable from Hugging Face.

At a glance
announcementWhen: announced and released; independent eva…
The developmentH Company announced and released Holo4, open-weight agentic models for generalist computer use, publishing weights, API access and all benchmark trajectories.
At a glance
announcementWhen: announced via Hugging Face and company…
The developmentH Company announced the release of Holo4, a two-model series of open-weight computer-use agents, along with an updated Holotron4 Nano and open-sourced benchmark trajectories.

Open-Weight Agents Closing the Gap

The release matters because open-weight computer-use agents remain rare at this performance level. If Holo4’s reported scores hold up under independent evaluation, businesses could run capable software-automation agents at a fraction of the cost of frontier closed models, with the flexibility of self-hosting or open weights.

The cost-performance gap the company claims — 61.7% on OSWorld 2.0 from a 27B model against 81.8% from a much larger closed model — would represent meaningful progress for smaller, cheaper agents. The multi-interface design also addresses a practical limitation: real business tasks often mix screen work, code and API calls, and single-interface models break at those boundaries.

H Company’s decision to release all benchmark trajectories lets outside parties verify each step, which is more transparency than most closed-model providers offer.

From Holo1 to Holo4

Holo4 builds on H Company’s previous agentic model line and arrives alongside an updated version of Holotron 3, called Holotron4 Nano. The company’s cost comparisons use specific assumptions: Holo4 is priced at H Models API rates for a single run, Qwen costs are calculated at Alibaba Cloud list prices (with cache hits at 20% of input price for the MoE model), and GPT and Opus effort sweeps come from OpenAI launch data. The company cautions that releases, harnesses and task subsets differ across the compared models.

“Real work is not siloed that way, and a single business task can require combining these different approaches.”

— H Company, announcement

Claims Awaiting Independent Verification

All headline benchmark numbers are self-reported by H Company and measured in the company’s own harness, which the company itself notes differs from other models’ releases, harnesses and task subsets. On AutomationBench, other models’ scores come from the public set while costs come from a leaderboard running the private set — a mismatch the company acknowledges. Holo4 has not yet been evaluated on the AutomationBench private set.

The steep score difference between the 27B dense model (61.7%) and the larger 35B-A3B MoE model (30.9%) on OSWorld 2.0 is not explained in the announcement. Real-world reliability on business workflows, beyond curated demo examples, also remains unverified by third parties.

Evaluations and Adoption Watchpoints

H Company says it will report Holo4 results on the AutomationBench private set once that evaluation is complete. Independent benchmark submissions and third-party reproductions — now possible because trajectories and weights are public — will be the next test of the company’s claims. Developers can access the models through the H Models API or download the full collection from Hugging Face.

Key Questions

What is Holo4?

Holo4 is a series of open-weight agentic models from H Company, released in a 27B dense version and a 35B-A3B Mixture of Experts version. A single model is designed to operate software through GUIs, code, MCP and APIs.

How does Holo4 perform on benchmarks?

According to H Company’s own measurements, Holo4 27B scores 61.7% on OSWorld 2.0, with the 35B-A3B variant at 30.9%, compared with 81.8% for the strongest closed model compared, Opus 5.5. These figures have not yet been independently verified.

Where can developers get Holo4?

Both models are available on the H Models API and for download on Hugging Face in FP16, FP8 and GGUF formats, enabling self-hosting.

Why did H Company release the benchmark trajectories?

Releasing every trajectory behind its public benchmark scores, viewable at trajectories.hcompany.ai, allows outside parties to audit each step and attempt independent reproduction — a level of transparency most closed-model providers do not offer.

What is the main caveat about Holo4’s results?

All headline numbers are self-measured in H Company’s own harness, and the company itself notes that releases, harnesses and task subsets differ across the models compared. The gap between the two model sizes on OSWorld 2.0 is also unexplained.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Trends to Watch in 2026: Predictions for the Year Ahead

With AI rapidly evolving in 2026, discover the key trends shaping its future and why they matter for industries and innovators alike.

The European Bet: How Mistral, Aleph Alpha, and Black Forest Labs Are Playing a Different Game

European AI vendors Mistral, Aleph Alpha, and Black Forest Labs are positioning for the EU AI Act enforcement, emphasizing compliance and sovereignty over frontier capabilities.

AI in Healthcare Education: Training the Next Generation of Medical Professionals

Navigating the future of medical training with AI promises personalized learning and immersive simulations, but what ethical dilemmas might arise?

Google and Amazon Fuel Race in AI With Billion-Dollar Investments

AIThis post was created with the assistance of artificial intelligence (AI). As…