Anyscale on Azure went generally available on October 7, closing out a run that started with a private preview in November 2025 and a public preview at Microsoft Build this June. The short version: managed Ray on Kubernetes, but the Kubernetes cluster is yours, sitting in your Azure subscription, and the bill lands on your normal Azure invoice.

That last part is what separates this from most managed AI platform announcements. Anyscale operates the control plane in its own Azure tenant: scheduling, monitoring, the web console. The data plane, meaning your actual workloads, container images, and datasets, runs on Azure Kubernetes Service inside the customer’s subscription. For teams that spent the last year arguing with security reviewers about sending training data to third-party SaaS, that split is the whole pitch.

What actually shipped

The Microsoft Learn overview describes it as an Azure Native Integration, which in practice means the usual integration points you would expect: sign-in through Microsoft Entra ID SSO, governance through Azure RBAC and Azure Policy with three built-in Anyscale roles, storage on Azure Blob Storage or Data Lake Storage, custom images from Azure Container Registry, secrets in Key Vault, and client access through Azure Load Balancer. GA covers twelve regions including East US, West Europe, Sweden Central, and Australia East.

One honest limitation worth flagging early: this is AKS-only. There is no VM stack, and Anyscale-hosted clouds are not part of the Azure offering. If your team wanted the lighter managed option without running any Kubernetes yourself, this product is not that. Ray workloads also need at least 4 vCPUs per worker node, with the quickstart suggesting Standard_D4s_v5 as a starting point.

The pricing question

Billing is pay-as-you-go on a single Azure invoice. The published rates for the Anyscale runtime layer look like Compute at $0.013 per hour and Memory at $0.002 per hour, with GPU attach rates like H100 80GB at $2.134 per hour and T4 at $0.145 per hour, stacked on top of standard AKS compute, storage, and networking charges. Anyscale’s launch post also confirms spend counts toward an existing Microsoft Azure Consumption Commitment, so no separate procurement exercise.

My take: that is genuinely convenient for FinOps and genuinely dangerous for anyone who forgets that autoscaling GPU pools and idle time are the two classic ways managed compute bills balloon. Model your expected GPU hours before committing. The pricing page is detailed enough to do it on paper first.

Why Ray, and why now

Ray has quietly become the default runtime for the parts of the AI loop that are not inference: data preprocessing, distributed training, post-training work like RLHF, and batch scoring. Anyscale was founded by the team that created Ray, and the project now sits under the PyTorch Foundation. The Azure-native version ships Anyscale’s optimized Ray-compatible runtime, formerly RayTurbo, which the AKS team has previously described as scaling better than open source KubeRay on AKS.

The strategic read is fairly clear. Microsoft gets to keep owning the Kubernetes substrate while a partner owns the AI framework layer, and Anyscale gets distribution into enterprises that would never sign a separate contract for compute. It also positions owning your learning loop on open source Ray against renting inference by the token, which is a real fork in the road for anyone architecting an AI platform right now. Running everything on one shared GPU pool beats stitching together separate stacks per stage, both for utilization and for sanity.

Named early customers include Xoople, doing satellite imagery intelligence, and Wayve, for autonomous driving training and deployment. Both are heavy distributed-Python shops, which tracks.

Caveats before you click deploy

The control plane runs in Anyscale’s Azure tenant, and while the data plane stays in yours, some operational metadata obviously flows to the platform side. Teams with strict data residency requirements should read the FAQ closely and check what leaves the tenant. GPU SKU availability also varies by region, so check quota in your target region early; vCPU quota requests can take longer than the deployment itself.

If you are already on AKS, the realistic comparison set is this product versus self-managed Ray on KubeRay versus Azure Machine Learning. Anyscale on Azure wins on cluster lifecycle management and the optimized runtime, at the cost of a per-hour platform meter and a second vendor’s control plane to trust. Self-managed KubeRay is cheapest on paper and most expensive in engineer-hours. Azure ML makes sense if your loop is mostly model training with less custom distributed Python. There is no universally right answer, but at least now there are three documented ones.

GA includes a production SLA and support from both Microsoft and Anyscale, and the service inherits the underlying AKS SLA. Deployment is a portal-driven Azure resource creation away if you want to poke at it today.

One practical workflow note for anyone migrating from Anyscale-hosted clouds: existing Ray job definitions and cluster configs port over conceptually, but environment variables, storage mounts, and secrets need re-wiring to their Azure equivalents, so budget a day or two of plumbing rather than assuming a config-file swap. Teams running hybrid fleets, some workloads on Anyscale’s own cloud and some on Azure, should also confirm how their observability tooling reaches the data plane in each location. None of this is hard, but it is the kind of work that eats a sprint if nobody owns it. Start with one low-stakes batch job, watch the Azure cost meters for a week, then move the training workloads that actually matter.

Leave a Reply

Your email address will not be published. Required fields are marked *