A harness that never makes the model wait
unreal-agent has been climbing GitHub’s trending list all week, at roughly 1.9k stars four days after it appeared, and the pitch behind it is simple enough to repeat in one line: the same model, doing the same work, up to 40 percent cheaper, because the harness never makes the model wait on a tool call. It is MIT licensed, written almost entirely in Go, and it comes from a small team calling themselves Unreal Labs.
The claim sounds like marketing until you look at where the tokens actually go in a typical agent run. Most harnesses run tools synchronously. The model issues a tool call, the harness blocks, and while the tool grinds away, the harness burns tokens on waiting behavior: heartbeat messages, status checks, polling loops. Unreal Labs argues those tokens are a third cost variable that sits entirely under harness control, and their numbers say removing them is worth real money.
What async-first actually means
The mechanics, from the launch post: the moment a tool call is issued, the harness appends an in-progress record to an event log. The tool runs in the background. When it finishes, the harness appends the result and only then calls the LLM again. The model never polls, never heartbeats, never burns a turn asking “is it done yet?”
There is real engineering underneath. Tool translators validate synchronously on the coordinator’s event loop with no I/O, so nothing suspends the loop. Sessions are append-only and forkable, so you can branch a run without copying state. The operation execution layer is swappable and durable. One friction point the team documents honestly: carrying one in-progress and one final tool result in a single context is underspecified in the OpenAI Responses API docs, and some non-OpenAI inference providers reject the pattern. They are openly asking providers to standardize it, which tells you how new this design space is.
The benchmark numbers, with the usual caveat
All of the following come from the vendor, run on Harbor with reproduction jobs published, so treat them as claims rather than measurements. On Terminal-Bench 4.0 with GPT-6 Astra xhigh, unreal-agent matched Codex at a 57.9 percent pass rate but spent $1,428 against Codex’s $2,350, about 39 percent cheaper. On DeepSWE 1.1 it scored 72.4 percent at $1,367 against Codex’s 69.0 percent at $1,633. On SWE-Atlas Codebase QnA it was 65.8 percent at $936 versus 63.3 percent at $1,303. The “up to 40 percent” headline is the ceiling; the published tables land between 16 and 39 percent depending on workload. An independent writeup by ThakiCloud reproduced the structure and the Terminal-Bench figure, which is something, though one reproduction is not an industry audit.
The savings come from two places: a minimal harness footprint (simple prompts, token-optimized tool results, no sub-agent machinery) and more tool work packed into each model turn. That combination matters most on expensive frontier models, where every wasted heartbeat token is billed at premium rates.
Taking it for a spin
Installation is one command if you have Go:
go install github.com/unreallabsai/unreal-agent/cmd/unreal-agent-runner@latest
Set UNREAL_HARNESS_LLM_PROVIDER and UNREAL_HARNESS_LLM_MODEL, then run unreal-agent-runner -p 'your prompt here'. Providers include ollama, openai, openai-codex, openrouter and fireworks, so you can point it at a local model and watch the async tool handling without spending anything on API calls. The repo is organized as a harness library, standalone executables, and a Harbor-compatible benchmark runner, so the benchmark claims are at least reproducible from the same codebase.
Reception and the harness-as-research argument
The launch hit the front page of Hacker News with 246 points and 124 comments, and the discussion split two ways. Half engaged seriously with harness design as a research area, citing the team’s reference to the HarnessTax paper on how much the harness itself matters for coding agent outcomes. The other half asked why anyone would name a coding tool “Unreal” when Epic Games owns one of the most famous trademarks in software. That second point is not a joke; it is a genuine positioning risk for the project.
The team also takes a position that will annoy some people: they argue CLI-oriented SDKs carry production lifecycle assumptions that do not fit, and they prefer enforcing security through sandbox constraints and access tokens outside the harness rather than harness-level hooks. Reasonable people will disagree with that, and saying so in the launch post is a confident choice for a three-contributor project.
Worth watching, not yet worth switching
The honest read: the mechanism is genuinely clever, the cost framing is right, and prompt cache preservation plus background tool execution is the kind of systems work that actually moves agent economics. It is also a very young project from a very small team, the benchmarks are vendor-produced, and the savings range is wide enough that your workload may see the low end. If you operate coding agents on frontier models, the useful move this week is not to switch harnesses, it is to measure how many tokens your current harness spends on waits and polls. That number tells you whether async tool execution is worth your attention. If it is large, unreal-agent just made the case that it is.