Skip to content
HomeNewsDigestsConceptsGuidesToolbox
AboutSubscribeUA
Subscribe

AI Today Brief

The daily AI-engineering brief. Built in public. EN · UA.

XTelegramLinkedInYouTubeRSS

Follow AI Today Brief on LinkedIn for daily AI-engineering updates and the weekly “5 shifts that changed how developers work” PDF.

Explore

NewsDigestsConceptsGuides

Company

SubscribeAdvertiseAbout

Legal

Editorial policyAI disclosurePrivacyTerms

© 2026 AI Today Brief. All rights reserved.

  1. Home/
  2. News/
  3. Models & research/
  4. H Releases Holo4 Open Models for GUI, Code, and Model Context Protocol
Models & research

H Releases Holo4 Open Models for GUI, Code, and Model Context Protocol

H released Holo4, an open-weight model family featuring a 27B dense and a 35B-A3B Mixture of Experts architecture. The models natively bridge desktop GUI navigation, shell execution, and Model Context Protocol tool calling at significantly reduced inference cost.

September 28, 2026· 6 min read
OKCurated by Oleksandr Kuzmenko, AI Product Engineer·Updated September 28, 2026·Sources cited on every story
AI-assisted · editor-reviewed·How we use AI
H Releases Holo4 Open Models for GUI, Code, and Model Context Protocol

Impact: High

Why it matters

You can now run unified computer-use and tool-calling agents locally or via API without needing separate models for GUI navigation and code generation.

TL;DR

  • 01Holo4 unifies GUI clicking, shell interaction, and MCP tool calling into a single 27B or 35B model.
  • 02Achieves 61.7% on OSWorld 2.0, outperforming Qwen bases while dramatically cutting token consumption.
  • 03Model weights are open in GGUF, FP8, and NVFP4 formats for local workstation execution.

Key facts

Architectures
27B Dense, 35B-A3B MoE, Holotron4 Nano
OSWorld 2.0 Score
61.7% (Holo4 27B), 30.9% (Holo4 35B-A3B)
Supported Interfaces
Desktop GUI, Web, Android, Code Sandbox, MCP, APIs
Quantizations Available
BF16, FP8, NVFP4, 4-bit GGUF

Multimodal Computer-Use Beyond Siloed APIs

Traditional agent pipelines typically rely on one model for screen reading and an entirely different model for calling external APIs. Holo4 combines these capabilities, executing actions across desktop GUIs, code sandboxes, and Model Context Protocol (MCP) tools within the same operational loop.

Benchmark Results and Trajectories

Evaluations across standard agentic benchmarks show high task completion rates at fractions of closed-frontier token overhead:

  • OSWorld 2.0: Holo4 27B scores 61.7% (compared to 81.8% for Opus 5.5 and 66.2% for GPT-5.6 Sol), while Holo4 35B-A3B reaches 30.9%.
  • Token Efficiency: In a complex automated CAD task in FreeCAD, Holo4 27B completed the design in 68 calls and 2.4M tokens (268 lines of code), compared to base Qwen3.8 27B requiring 197 calls and 11.4M tokens.
  • Harness Upgrades: The release includes an updated agent harness offering persistent multi-hundred-step memory and a local desktop shell interface.

Open Weights and Deployment Options

Checkpoints are published directly to Hugging Face across multiple quantization profiles, including BF16, FP8, NVFP4, and 4-bit GGUF. Developers can deploy them locally through llama.cpp backends or query them through the H Models API.

Try it in 2 minutes

from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="Hcompany/Holo4-27B",
    allow_patterns=["*.gguf"],
    local_dir="./holo4-27b"
)

python

✓ When to use

  • Building autonomous agents that must navigate graphical software while simultaneously calling CLI tools.
  • Local deployment of computer-use agents without sending proprietary screen captures to closed cloud vendors.

✕ When NOT to use

  • Pure text-based completions or standard conversational chatbots where minimal latency is required.
  • Constrained edge microcontrollers lacking at least 16GB VRAM to run quantized 27B models.

What to do today

  • →Download the 4-bit GGUF variant of Holo4 27B from Hugging Face for local computer-use testing.
  • →Connect Holo4 to local Model Context Protocol servers to verify combined UI and API tool orchestration.
  • →Examine the open benchmark trajectories on trajectories.hcompany.ai to understand multi-step failure recovery.
#Holo4#Holotron4 Nano#Qwen#Nemotron#FreeCAD#Model Context Protocol

Sources

  • Hugging Face Blog: Holo4 Announcement
  • H Company Newsroom: Holo4 Details
ShareShare on XShare on LinkedIn
← Previous storyPaperclip Open-Sources Control Plane for Multi-Agent Workflows and BudgetsNext story →Fireworks AI Releases Ember-1: About 40% Fewer Tokens Than Kimi K3

Related stories

  • Models & researchNVIDIA Nemotron 3 Diarization Separates Eight Concurrent Speakers in Real Time
  • Models & researchStepFun Step 5 Preview: 600B Sparse MoE Agent Model with Claude Code Support
  • Models & researchAlibaba Launches Qwen3.8-LiveTranslate Realtime WebSocket Interpretation Model
  • Models & researchGoogle Launches Gemini 3.8 Live and Extended Thinking Voice Models

Email digest

Get the morning AI brief

One email a day — the stories that matter for engineers, founders and tech leads. Human-edited, with links to primary sources.

  • ✓120+ sources scanned daily
  • ✓Edited by a human
  • ✓1 email per day
  • ✓EN + UA

By subscribing you agree to the privacy policy.