H Releases Holo4 Open Models for GUI, Code, and Model Context Protocol
H released Holo4, an open-weight model family featuring a 27B dense and a 35B-A3B Mixture of Experts architecture. The models natively bridge desktop GUI navigation, shell execution, and Model Context Protocol tool calling at significantly reduced inference cost.

Impact: High
Why it matters
You can now run unified computer-use and tool-calling agents locally or via API without needing separate models for GUI navigation and code generation.
TL;DR
- 01Holo4 unifies GUI clicking, shell interaction, and MCP tool calling into a single 27B or 35B model.
- 02Achieves 61.7% on OSWorld 2.0, outperforming Qwen bases while dramatically cutting token consumption.
- 03Model weights are open in GGUF, FP8, and NVFP4 formats for local workstation execution.
Key facts
- Architectures
- 27B Dense, 35B-A3B MoE, Holotron4 Nano
- OSWorld 2.0 Score
- 61.7% (Holo4 27B), 30.9% (Holo4 35B-A3B)
- Supported Interfaces
- Desktop GUI, Web, Android, Code Sandbox, MCP, APIs
- Quantizations Available
- BF16, FP8, NVFP4, 4-bit GGUF
Multimodal Computer-Use Beyond Siloed APIs
Traditional agent pipelines typically rely on one model for screen reading and an entirely different model for calling external APIs. Holo4 combines these capabilities, executing actions across desktop GUIs, code sandboxes, and Model Context Protocol (MCP) tools within the same operational loop.
Benchmark Results and Trajectories
Evaluations across standard agentic benchmarks show high task completion rates at fractions of closed-frontier token overhead:
- OSWorld 2.0: Holo4 27B scores 61.7% (compared to 81.8% for Opus 5.5 and 66.2% for GPT-5.6 Sol), while Holo4 35B-A3B reaches 30.9%.
- Token Efficiency: In a complex automated CAD task in FreeCAD, Holo4 27B completed the design in 68 calls and 2.4M tokens (268 lines of code), compared to base Qwen3.8 27B requiring 197 calls and 11.4M tokens.
- Harness Upgrades: The release includes an updated agent harness offering persistent multi-hundred-step memory and a local desktop shell interface.
Open Weights and Deployment Options
Checkpoints are published directly to Hugging Face across multiple quantization profiles, including BF16, FP8, NVFP4, and 4-bit GGUF. Developers can deploy them locally through llama.cpp backends or query them through the H Models API.
Try it in 2 minutes
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="Hcompany/Holo4-27B",
allow_patterns=["*.gguf"],
local_dir="./holo4-27b"
)python
✓ When to use
- Building autonomous agents that must navigate graphical software while simultaneously calling CLI tools.
- Local deployment of computer-use agents without sending proprietary screen captures to closed cloud vendors.
✕ When NOT to use
- Pure text-based completions or standard conversational chatbots where minimal latency is required.
- Constrained edge microcontrollers lacking at least 16GB VRAM to run quantized 27B models.
What to do today
- Download the 4-bit GGUF variant of Holo4 27B from Hugging Face for local computer-use testing.
- Connect Holo4 to local Model Context Protocol servers to verify combined UI and API tool orchestration.
- Examine the open benchmark trajectories on trajectories.hcompany.ai to understand multi-step failure recovery.
Sources