Lifestyle

H Company Releases Holo4, an AI Agent Model It Says Can Operate Screens, Code and APIs

On September 28, 2026, H Company unveiled Holo4, a family of AI 'agent' models in two sizes (27B and 35B-A3B). The company says one model can click through apps on screen, write code and call software tools, which could matter to businesses automating their work. Here is what was announced, the company's own benchmark figures, and what to keep in mind.

About 7 min read

H Company Releases Holo4, an AI Agent Model It Says Can Operate Screens, Code and APIs
Image: Mokaair (Original editorial artwork)

What H Company Released

On September 28, 2026, H Company announced Holo4 on the Hugging Face blog and described it as a new family of agent models. An agent model is an AI system built to carry out tasks step by step, not just answer questions. Holo4 comes in two sizes: a 27B dense model and a 35B-A3B Mixture of Experts model. Both are available through the H Models API, the company's paid service for using its models. The company also released Holotron4 Nano, an updated version of Holotron 3.

The numbers in the names refer to parameters, the internal values a model learns during training. "27B" means about 27 billion. A "dense" model uses all of its parameters for every step. A Mixture of Experts model is split into specialised parts and uses only some of them at a time, which can make it cheaper to run. H Company lists the full Holo4 collection on Hugging Face in FP16, FP8 and GGUF formats, which are different file versions suited to different hardware. The company says Holo4 improves significantly over the Qwen models it is based on, and names Qwen3.8 27B as the base model for Holo4 27B.

H Company Releases Holo4, an AI Agent Model It Says Can Operate Screens, Code and APIs
Mokaair editorial verification flow · Image: Mokaair (Original editorial artwork)
Read the full description

Sources are collected, independently checked, then reviewed by Jev.

What "One Model for Every Interface" Means

Software can be operated in several ways. There is the GUI (graphical user interface: the windows, buttons and menus a person clicks). There is code. And there are APIs and MCP, which are ways for programs, including AI models, to call other software directly without a screen. H Company says most agent models are trained for only one of these. Models focused on graphical interfaces cannot work without a screen, while models that favour tool calling get stuck on applications that have no API. The company says Holo4 can click and type on screen, write and run its own code, and call MCP or API tools, choosing the method that fits each task.

According to H Company, Holo4 runs on desktop, web, Android, code sandboxes and business APIs. It is the same model, called the same way, in every setting, so there is no need to pick a different model for each platform.

Benchmark Figures Published by H Company

A benchmark is a standard set of test tasks used to compare AI models. H Company calls OSWorld 2.0 one of the hardest academic benchmarks for controlling a desktop computer, and AutomationBench one of the hardest for API use.

All figures come from H Company's Hugging Face post; the company notes that model versions, test harnesses and task subsets differ.
ModelOSWorld 2.0 score (compiled by H Company)Notes
Holo4 27B61.7%H Company's own model
Holo4 35B-A3B30.9%H Company's own model
Opus 5.581.8%Listed by H Company for comparison
Opus 570.2%H Company citing an OpenAI release chart
GPT-5.6 Sol66.2%H Company citing an OpenAI release chart

H Company says Holo4 trails only the strongest closed models on long workflows, while using orders of magnitude fewer parameters and costing far less. It says costs are estimated from the input and output tokens of each agent run. Tokens are the small chunks of text a model reads and writes, and they are what usage is billed by. Holo4 is priced at H Models API rates. For AutomationBench, H Company says Holo4's scores and costs were measured in its internal harness, and that it will report Holo4's results on the benchmark's private test set once they have been evaluated.

Real-World Task Examples and Training Approach

H Company gave Holo4 27B and its base model Qwen3.8 27B the same prompts and the same harness on professional software. It then published how many calls and tokens each one used:

  • Building a Pac-Man game in Godot: Holo4 27B used 68 calls, 2.4M tokens and 268 lines of code; Qwen3.8 27B used 197 calls, 11.4M tokens and 327 lines of code.
  • Modeling the Eiffel Tower in FreeCAD: Holo4 27B used 84 calls and 1.3M tokens; Qwen3.8 27B used 60 calls and 1.0M tokens.
  • Modeling an H logo in FreeCAD: Holo4 27B used 94 calls and 1.5M tokens; Qwen3.8 27B used 118 calls and 1.9M tokens.

On training, H Company says Holo4 learned through supervised learning (from worked examples) and reinforcement learning (from trial and feedback). It trained on a large number of environments and tasks, including tasks generated by the company's Agentic Task Factory. H Company says this tool has so far produced about 10,000 tasks across web applications, MCP servers and desktop environments. The company also says it rebuilt its harness, which it describes as the loop that carries out the model's actions and manages what it remembers over hundreds of steps. The biggest changes were a reliable memory that can keep track of hundreds of steps, and a shell (a command line) on the desktop machine itself.

What It Means for General Readers

A "computer-use agent" is an AI that can look at a screen, click buttons and type text like a person, or switch to code and APIs to get work done. If H Company's claims hold up, businesses may one day be able to automate workflows that span websites, desktop software and internal systems without preparing a separate model for each platform. For now, this kind of technology mainly affects workplace automation tools rather than everyday consumer products.

It is also worth remembering that an AI operating a computer directly can take real actions. Any organization considering such agents should first test them in an isolated environment, limit the accounts and data they can access, and keep a human review step for important operations. This article is a news summary only and does not constitute purchasing advice.

Frequently asked questions

Who launched Holo4, and when was it released?

H Company announced it on the Hugging Face blog on September 28, 2026.

What versions of Holo4 are there?

According to H Company, Holo4 comes in two sizes, a 27B dense model and a 35B-A3B Mixture of Experts model, both available through the H Models API. The company also released Holotron4 Nano.

Is Holo4 really cheaper and better than other models?

H Company says Holo4 competes with frontier models at far lower cost. In its own OSWorld 2.0 figures, however, Holo4 27B's 61.7% is still below Opus 5.5's 81.8%. The company compiled these figures itself and notes that test conditions differ across models. They have not been independently verified.

Can I check these benchmark results myself?

H Company says it open-sources every trajectory (step-by-step record) behind its public benchmark scores. You can replay them at trajectories.hcompany.ai or download them from Hugging Face.

Which platforms can Holo4 run on?

H Company says Holo4 runs on desktop, web, Android, code sandboxes and business APIs, using the same model called in the same way.

Browse the latest news in this topic

Latest travel guides

Sources

Lifestyle