Holo4: Generalist Models for Computer-Use Agents
Holo4: Generalist Models for Computer Use Agents Published September 28, 2026 Holo4 is a new series of agentic models available in two sizes: a 27B dense model and a 35B A3B Mix...
By AI Engineering Team
Holo4: Generalist Models for Computer-Use Agents
Published September 28, 2026
Holo4 is a new series of agentic models available in two sizes: a 27B dense model and a 35B-A3B Mixture of Experts model. Both are available through the H Models API. An updated version of Holotron 3, called Holotron4 Nano, is also being released.
Holo4 builds on the previous model and interacts with software through available interfaces, including graphical user interfaces, code, MCP, and APIs. The models were trained with supervised and reinforcement learning across environments and tasks, including tasks generated by the Agentic Task Factory.
The models and related resources include:
- Models: Holo4-27B, Holo4-35B-A3B, and Holotron4-30B-A3B
- Model collection: FP16, FP8, and GGUF versions of Holo4
- Trajectories: A viewer and downloadable dataset
- H Models API: Documentation and quickstart resources
Models for multiple interfaces
Holo4 can click and type on a screen, write and execute code, and call MCP or API tools. It selects the interface that fits the task. Many agentic models are trained for only one interaction method. GUI-focused models cannot operate without a screen, while tool-calling models are limited when an application has no API. Business workflows can require several of these methods in a single task.
Holo4 runs on desktop systems, websites, Android, code sandboxes, and business APIs. The same model can be used across these environments and called through the same interface, without selecting a separate model for each platform.
Holo4 models improve substantially over their Qwen base models. On long workflows in OSWorld 2.0, Holo4 27B scores 61.7%, compared with 81.8% for Opus 5.5. Holo4 35B-A3B reaches 30.9%. The results use substantially fewer parameters and lower task costs than the strongest closed models.
The trajectories behind the public benchmark scores are available for replay and download, allowing each step of an agent run to be examined.
Benchmark performance and task cost
On OSWorld 2.0, which evaluates desktop control, and AutomationBench, which evaluates API use, Holo4 competes with frontier models at a lower cost per task.
Cost-performance methodology
OSWorld 2.0: Costs are estimated from the input and output tokens used in each agentic run. Holo4 uses H Models API rates for a single run. For Qwen3.8 27B, the cost is based on tokens from the authors' run and Alibaba Cloud list prices. Qwen3.6 35B-A3B uses a single run in the same harness, with Alibaba Cloud list prices and cache hits charged at 20% of the input price.
OpenAI launch data supplies the GPT and Opus effort sweeps. Other closed and open-weight results come from the official OSWorld 2.0 leaderboard. Releases, harnesses, and task subsets differ between evaluations. The line in the cost-performance chart connects non-dominated score and cost pairs among closed models, excluding Holo4.
AutomationBench: Holo4, Qwen3.8 27B, and Qwen3.6 35B-A3B were evaluated with AutomationBench v1.0.6 in an internal harness, with scores and costs measured there. Results for other models use public-set scores from the AutomationBench README and cost-per-task figures from the official leaderboard, which uses a private set. Holo4 had not yet been evaluated on that private set.
Examples of professional software tasks
Holo4 was trained on environments and tasks generated by the Agentic Task Factory. The following examples compare Holo4 27B with its base model, Qwen3.8 27B, using the same prompt and harness.
3D modeling: Eiffel Tower
The task required building a FreeCAD model of the Eiffel Tower at a scale of 1 mm to 1 metre, centered on the origin and aligned to the X and Y axes. The tower had to remain square in plan at every height, with specified half-widths at heights 0, 57, 115, and 276, connected by a smooth curve.
The model required four identical square-column legs, three solid square platforms, and a tapered square mast. Each component had to be a closed solid with non-zero volume, while the space between the legs had to remain open.
- Holo4 27B: 84 calls, 1.3 million tokens
- Qwen3.8 27B: 60 calls, 1.0 million tokens
3D modeling: H logo
The task required creating a FreeCAD model of the H company logo, consisting of a filled disc and a block-style sans-serif capital H. Both shapes had to be extruded to the same thickness, separated without overlap, centered vertically, and colored black.
- Holo4 27B: 94 calls, 1.5 million tokens
- Qwen3.8 27B: 118 calls, 1.9 million tokens
Game design: Pac-Man-style game
The task required building and running a Godot game with a rectangular grid maze, pellets in open corridors, a continuously moving player, three pursuing ghosts, a score display, and a lives counter. The game had to reset after a collision and operate indefinitely without keyboard input.
The player was controlled by a simple heuristic: it moved toward the nearest pellet at junctions unless a ghost was nearby, in which case it moved away from the ghost.
- Holo4 27B: 68 calls, 2.4 million tokens, 268 lines of code
- Qwen3.8 27B: 197 calls, 11.4 million tokens, 327 lines of code
How Holo4 was developed
Agentic Task Factory
The internal Agentic Task Factory creates interactive environments and verifiable tasks from documentation alone. Its inputs can include screenshots of real websites and open-source software. It has produced about 10,000 tasks across web applications, MCP servers, and desktop environments. Some tasks are hybrid environments that expose the same state through both a GUI and MCP.
Training
The training process used supervised fine-tuning on 127 billion tokens, followed by two reinforcement-learning experts and a merged model.
Harness
The development team also rebuilt the harness, which executes model actions and manages context across runs that can last hundreds of steps. The redesign incorporated feedback from OSWorld 2.0 performance. Agents identified reasons for task failures, and engineers reviewed the resulting fixes.
The largest changes included a reliable memory system that can track hundreds of steps and a shell running directly on the desktop machine.
For one historical comparison, Opus 5 scored 70.2% and GPT-5.6 Sol scored 66.2% using maximum-effort partial rewards on the v2026.08.08 offline set from OpenAI's launch chart. Other reference scores came from model cards and the official leaderboard. Task releases, subsets, and harnesses varied.
Holotron4 Nano
The post-training stack is designed to adapt to new foundation models and produce agents that generalize across interfaces and environments. As part of the NVIDIA Nemotron Coalition, the stack was applied to the Nemotron 3 Nano Omni model as a follow-up to Holotron 3.
The resulting Holotron4 Nano is a generalist agentic model that improves over Nemotron 3 Nano Omni on GUI workflows and environments exposing MCP, APIs, or coding sandboxes.
The reported gains are absolute percentage-point improvements over Nemotron 3 Nano Omni across five benchmarks. These results indicate that the post-training recipe can transfer to a generalist foundation model and adapt it for agentic tasks. The approach is not limited to a particular model size.
Availability
Holo4 27B and Holo4 35B-A3B are available through the H Models API. Model weights are available in BF16, FP8, NVFP4, and 4-bit GGUF formats, alongside Holotron4 Nano, the smaller model.
Optimized DSpark drafter checkpoints were scheduled for release in the following days to accelerate inference.