Microsoft has taken the Troubleshooting Agent in Azure Copilot to general availability. It lives in the Azure portal, in both the Copilot chat experience and the Support + Troubleshooting blade, and its job is to take a vague complaint like “my VM rebooted itself last night” and turn it into a diagnosis, a fix, or a support ticket with the evidence already attached.
Deep troubleshooting for Azure Compute and Azure Kubernetes Service is also generally available alongside it. Azure Local and Microsoft Entra troubleshooting are in public preview. If you don’t see the agent in your portal, an administrator may have it disabled; tenant admins can toggle access to Copilot agents.
How the agent works
The flow runs in five stages. You describe the problem in natural language, in the context of a resource, resource group, or subscription. The agent scopes the issue, identifying the affected resource and the product area, and only asks clarifying questions when it genuinely needs them. Then it diagnoses: it pulls health signals, resource diagnostics, and where they exist, product-specific diagnostic skills. If it finds a root cause it either recommends or applies a fix, sometimes as a one-click remediation. If it can’t resolve the problem, it escalates by creating a prepopulated support request or connecting you to a live support engineer, with the full session context attached so you don’t repeat yourself.
Two things in that flow matter more than the marketing suggests. First, the diagnostics are read-only, and any actual change to your environment comes with a recommendation you review before it happens. The agent works within your identity and Azure RBAC, so it can only see and do what your account can. Second, the escalation path is the part most teams will end up using. Anyone who has opened a Azure support ticket knows the first twenty minutes are spent re-explaining the problem and pasting diagnostic output. The agent assembling that package automatically is a real time saving.
What the Compute and AKS coverage looks like
For Azure Virtual Machines, Virtual Machine Scale Sets, and Compute Fleet, the agent handles the familiar failure modes: unexpected restarts, RDP and SSH connectivity failures, sustained CPU or disk performance problems, boot failures, allocation and deployment errors, VM agent health, unhealthy scale-set instances, and availability zone configuration issues.
The AKS story is broader, which reflects where most of the operational pain actually sits these days. The agent can investigate CrashLoopBackOff and startup failures, stalled deployments, pod scheduling and capacity constraints, service discovery and connectivity issues, scaling behavior, upgrade regressions, out-of-memory terminations, image pull failures, and general cluster performance degradation. Anyone who has stared at a pod stuck in CrashLoopBackOff at 2am knows the value of a tool that checks the cluster, explains what it found, and points at remediation docs in one pass.
Cost and limitations
There is no separate charge for the agent. No license, no per-query fee. You still pay for the underlying resources and your existing support plan works as before.
The limitations are worth reading before you lean on it. Automatic mitigation isn’t available for every issue or resource type, and in those cases you get step-by-step instructions instead of a one-click fix. Diagnostic depth varies by service; the product-specific integrations for Compute and AKS are the deep ones today, and coverage is expanding. And the agent works from the diagnostic data available to it, which means it can miss things a human with broader access would catch. Review the recommendations before applying anything, same as you would with a colleague’s pull request.
Why this matters
Support and troubleshooting is one of the more sensible places to put an AI agent, because the workflow is already structured: describe the symptom, gather evidence, form a hypothesis, apply a fix or escalate. That’s a state machine, not an open-ended reasoning problem, and grounding the agent in live resource diagnostics rather than generic documentation is what makes it useful rather than a search box with better manners.
Setup is deliberately boring. Open Azure Copilot in the portal, pick Troubleshooting from New chat, and describe the problem, naming the affected resource if the chat isn’t already scoped to it. That’s the whole onboarding story, and the fact that no configuration or data pipeline is required is probably why this shipped as a portal feature rather than a separate product SKU.
It’s also a sign of where the Copilot agents are heading. Microsoft has been shipping these one at a time, and troubleshooting was the obvious first candidate because the payoff is measurable: fewer support tickets, faster time to diagnosis, and less context-switching between portal blades. The Entra and Azure Local previews in the same announcement hint at where the next coverage is going.
If you run Azure estates, this is worth enabling and testing against your last three incidents. If it had diagnosed them faster than your on-call rotation did, you have your answer about whether to leave it on. If it flailed on your exotic networking setup, you’ve lost ten minutes and learned the boundary of today’s coverage.