What is Prepared Image Specification?
If you run AI training jobs, GPU inference, or large Windows workloads on AKS, you have probably noticed the startup lag. A new node joins the cluster, and then it spends minutes pulling container images before it can actually run anything. For autoscaling events, those minutes add up. Azure’s Prepared Image Specification, now available in public preview, aims to fix this by baking container images directly into the node image before the node provisions.
The idea is straightforward. Instead of letting each node download the same images from the registry every time it starts, you define which images you need in a customImage resource, and AKS pre-pulls them into the node image during the build phase. When the node boots, the images are already there. No pull wait, no registry latency, no bandwidth competition between new nodes during scale-out events.
Why this matters for AI and GPU workloads
AI training and inference workloads are particularly sensitive to node startup time because they tend to run on expensive GPU-enabled nodes. A node that sits idle for five minutes pulling images while an A100 GPU waits is literally burning money. Prepared Image Specification directly reduces that idle time.
The same applies to Windows workloads. Windows container images are notoriously large often multiple gigabytes. If you are running a Windows node pool that autoscales, the image pull time can dominate your scale-out duration. Pre-pulling those images into the node image means Windows nodes come online ready to serve.
How it works in practice
You define a customImage resource in a resource group and specify the container images you want pre-pulled, any OS-level customizations, and optional startup scripts. AKS uses this specification to build a prepared node image that it uses when provisioning new nodes in the cluster.
Key details from the preview:
- The customImage resource lives outside the cluster, in its own resource group, which means you can share prepared images across clusters
- You can specify multiple container images in a single specification. AKS pre-pulls them all during the image build phase
- OS customizations like kernel parameters, package installs, or registry settings can be included in the specification
- The prepared images are versioned, so you can track which base image and container image versions went into each build
- When you update the specification, AKS rebuilds the prepared image. Existing nodes keep running on the old image until they are replaced through your normal update strategy
Comparing with existing approaches
There are other ways to solve the cold-start image pull problem. DaemonSets that pre-pull images on every node. Custom node images built with Azure Image Builder. Third-party tools like kube-fledged or Spegel for peer-to-peer image distribution.
Each has trade-offs. DaemonSets work but compete for bandwidth during scale-out events. Custom node images give you full control but require maintaining a separate image build pipeline. Peer-to-peer distribution helps with large fleets but adds operational complexity.
Prepared Image Specification sits in the middle. You get the benefits of pre-pulled images without maintaining a separate image build pipeline. AKS handles the image build and node provisioning integration. The trade-off is that you are locked into Azure’s implementation, so if you run hybrid or multi-cloud, the approach does not port easily.
Practical considerations
A few things to think about before adopting this:
- Start with your most expensive node pools. GPU-backed and Windows node pools get the most benefit from reduced startup time. General-purpose Linux node pools with small container images may not see enough improvement to justify the setup.
- Version your prepared images. When you update container image tags in the specification, AKS rebuilds the prepared image. Existing nodes keep running with the old image until they cycle. Make sure your deployment strategy accounts for this lag.
- Monitor the build time of your prepared images. If your specification includes many large container images, the build itself takes time. The total time saved on node startup needs to exceed the initial build time for the approach to be worth it.
- The preview currently requires defining customImage resources in dedicated resource groups. Plan your naming and organization strategy before scaling to multiple clusters.
What this means for cluster operations
For teams managing large AKS clusters with autoscaling node pools, Prepared Image Specification changes the economics of scale-out events. When a traffic spike triggers a scale-out, new nodes start contributing to the workload immediately instead of spending minutes pulling images. This means you can set more aggressive autoscaling thresholds because the penalty for scaling out is lower.
It also simplifies the node readiness story. Currently, operators need to account for image pull time when setting readiness probes and pod disruption budgets. With pre-pulled images, the time from node creation to pod scheduling is more predictable, which makes capacity planning easier.
The preview is worth evaluating if GPU or Windows node startup time is a pain point in your current setup. Even a modest reduction in startup time translates to real cost savings when you are paying for expensive hardware that sits idle during boot.