Microsoft took the wraps off a public preview of the CodeAct pattern and Hyperlight containers for the Microsoft Agent Framework this week. Two features that together change how agents execute code under the hood, and they solve a problem that anyone who has built a serious agent has felt: the death by a thousand tool calls.

What CodeAct actually does

Most agents today work by chaining tool calls. The LLM decides it needs to do something, calls a function, waits for the result, looks at it, decides the next step, calls another function. Each round trip burns tokens and adds latency. If you have ever watched an agent debug trace scroll past, you have seen how many of those calls are wasted: checking if a file exists before reading it, checking if a service is running before calling it, verifying an action succeeded by calling the same status endpoint three times.

CodeAct collapses all of that into a single executable code block. The agent writes a complete script containing every step of its plan, runs it in one shot, and gets back the final result. Think of it as the difference between asking someone to build a chair one nail at a time and reporting back after each nail, versus handing them the blueprint, all the materials, and saying “come back when the chair is done.” The agent still plans. It just executes the plan as a single unit instead of crawling through it step by step.

Microsoft reports a 70 percent reduction in end-to-end latency on representative workloads. That is a big number. It comes from eliminating the model-to-tool round trips that eat up most of the execution time in agentic workflows. For context, OpenAI introduced a similar CodeAct approach in early 2025 with their agents SDK, and Anthropic’s Claude Code uses a version of the same concept. Microsoft’s implementation is notable because it ties directly into the Agent Framework’s existing harness, so you don’t need to restructure your agent.

Hyperlight micro-VMs for sandboxed execution

The obvious risk with code execution agents is letting a model run arbitrary code anywhere near your production data. CodeAct makes this worse because the agent runs a bigger, more complex script all at once rather than making small, auditable tool calls. Microsoft addresses this with Hyperlight, a micro-VM technology that sits well below the container layer.

Instead of spinning up a full Docker container, which takes seconds and carries a large surface area including an entire OS userland, Hyperlight boots a tiny virtual machine in milliseconds. Each agent gets its own isolated micro-VM that is destroyed after execution. No persistence, no lateral movement, no shared state. These are not containers. Hyperlight uses hardware virtualization extensions to create a lightweight sandbox at the VMM level. The attack surface is dramatically smaller than a container runtime, and the startup cost is close to negligible.

For developers, this means agents can safely run generated code without the operational overhead of managing sandbox environments. You do not need to configure Docker-in-Docker setups, manage container image builds, or worry about container escape vulnerabilities. The micro-VM is ephemeral and fully isolated by the hypervisor.

The feature requires host virtualization support: KVM on Linux, WHPX on Windows. On Azure VMs this is available out of the box. On-premises and edge deployments may need to check their host configuration before enabling the feature.

Putting it together in Agent Harness

CodeAct and Hyperlight are both configured through the Agent Harness, which reached GA earlier this month as the production runtime for agents built with the Microsoft Agent Framework. You enable CodeAct as a capability on your agent definition and select the Hyperlight backend for sandboxed execution.

For .NET developers, the Microsoft.Agents.AI.Hyperlight NuGet package provides the integration. Python developers configure it through the Agent Harness configuration YAML. The framework handles dispatching the code block to the micro-VM, collecting results, surfacing errors, and reporting back to the agent. If the micro-VM execution fails, the agent receives the error output and can respond accordingly, just like it would with a failed tool call.

What this means for agent development

The CodeAct pattern addresses a real pain point that gets worse as agents get more capable. Simple agents with two or three tool calls do not benefit much from CodeAct. The overhead of the pattern itself may even make them slower. But as agents grow to handle complex multi-step tasks with 20, 30, or 50 tool calls, the savings compound dramatically.

There is a caveat worth flagging. Giving an agent the ability to generate and execute arbitrary code is powerful, but it also increases the blast radius of a prompt injection attack. If a malicious input tricks the agent into generating code with side effects, the micro-VM isolation is your last line of defense. Hyperlight handles this well in theory. In practice, developers should apply least-privilege principles to what the agent can access, rather than relying solely on sandbox isolation. Consider what credentials the agent holds, what data sources it can reach, and whether the tasks it performs are reversible.

How to try it

CodeAct with Hyperlight is available in public preview today. The steps to get started are:

The features work with both new and existing agents, so you can test them on a development agent before rolling out to production.

Leave a Reply

Your email address will not be published. Required fields are marked *