Kubernetes AI platforms need workload classes before autoscaling

Platform engineering SeedlingPlanted Sep 2026

Kubernetes AI platforms need workload classes before autoscaling. A single policy cannot responsibly scale an interactive inference service, a queue-driven agent worker, a fine-tuning job, and a stateful retrieval engine merely because all four run in pods. Their demand signals, startup costs, interruption tolerance, storage obligations, and scarce resources differ. Autoscaling before classification turns those differences into noisy metrics and expensive incidents.

EKS exposes a rich substrate: managed node groups, Karpenter NodePools, KEDA ScaledObjects and ScaledJobs, GPU-specific pools, resource quotas, pod disruption budgets, topology spread, and storage classes that wait for a consumer before selecting an availability zone. The mistake is treating that catalogue as a menu to configure workload by workload. I would first define a small platform vocabulary that states which combinations are supported and what operating contract each one receives.

An interactive inference class might require warm capacity, accelerator affinity, model-readiness probes, topology spread, and strict scale-down protection. A batch-inference class can accept queue-driven ScaledJobs, cheaper capacity, and bounded retries. An API-bound agent class should scale on queue age or deadline risk rather than CPU, while a long-running agent class needs external checkpoints and disruption rules that protect in-flight work. Stateful data services need volume and availability semantics that should not be inferred from a deployment label.

The class must reach both scheduling and identity. A GPU class selects the allowed instance families, taints, resource ceilings, and Karpenter NodePool. A tenant-sensitive class adds namespace quotas, default-deny network policy, and a narrowly bound service identity through Pod Identity or IRSA. A public-serving class may attach load-balancer readiness gates and graceful deregistration. This makes “where can this run?” and “what may it reach?” part of the same declared contract instead of unrelated YAML assembled by each team.

Only then do autoscaling signals become interpretable. KEDA queue lag is meaningful for a homogeneous worker class; vLLM waiting requests can represent pressure for a warm inference class; CPU may be appropriate for a genuinely compute-bound service. Scale-to-zero is sensible for some batch workers and disastrous when model loading takes minutes against an interactive latency objective. The signal is not intrinsically good or bad. It is valid only relative to the workload contract it is supposed to protect.

Classes also make cost governance honest. Baseline capacity can use savings commitments, burst classes can diversify across Spot families, and expensive accelerator classes can carry explicit namespace quotas and NodePool limits. Kubecost can then attribute spend to a stable operating category rather than a shifting collection of deployments. Without that layer, consolidation may look efficient while evicting high-checkpoint-cost work or placing latency-critical inference behind a cold-start boundary.

There is one boundary: a small platform with two predictable services does not need a grand taxonomy or a new custom resource. Labels, documented defaults, and two reviewed deployment templates may be enough. The point is not to create a platform product prematurely. It is to make the differences explicit before an autoscaler encodes accidental assumptions about them.

I would approve an autoscaling design only after the workload declares latency objective, demand shape, state location, interruption budget, accelerator need, tenancy boundary, and startup cost. Kubernetes already provides most of the mechanisms. Workload classes supply the missing decision layer that tells those mechanisms which trade-offs are legitimate. Autoscaling can then respond to a known kind of work instead of treating every pod as an interchangeable consumer of CPU.