writing
Gang Scheduling in Kubernetes 1.37: A Deep Dive into All-or-Nothing Placement
Platform Engineering··15 min read

Gang Scheduling in Kubernetes 1.37: A Deep Dive into All-or-Nothing Placement

Six of eight workers hold a GPU and wait for two siblings that will never arrive. That is kube-scheduler working as designed. Gang scheduling went beta in Kubernetes 1.37, so here is what actually changed, down to the plugin code.

Share

Key Takeaways

  • kube-scheduler's unit of work is one pod, so an eight-pod training job is eight independent decisions. Partial placement is the designed outcome, not a bug.
  • Two eight-pod jobs on twelve GPUs each win six and neither yields. Allocation reads 100%, utilisation reads zero, and no component is misbehaving.
  • Gang scheduling reached beta in Kubernetes 1.37 through KEP-4671, with Workload and PodGroup in scheduling.k8s.io/v1beta1 behind the GenericWorkload feature gate.
  • Beta does not mean on. GenericWorkload ships with Default: false in 1.37, and there are two switches, not one: the gate on the apiserver, scheduler and controller manager, plus --runtime-config=scheduling.k8s.io/v1beta1=true, because new beta group versions are served only when asked for.
  • All-or-nothing is enforced at three points in the plugin: PreEnqueue keeps a gang out of the queue until minCount pods exist, PlacementFeasible kills a group cycle that can no longer reach minCount, and Permit parks every placed pod for up to five minutes until the quorum releases them all to bind at once.
  • minCount is a placement-time guarantee only. PodGroupInitiallyScheduled is terminal once True, so a gang that later loses pods to eviction is never re-gathered.
  • Upstream lists fairness and multiple queues as non-goals, so Kueue keeps quota and admission and Volcano keeps DRF and GPU sharing. What you can retire is a scheduler replacement you installed purely to get all-or-nothing.

A training job asks for eight workers, one H100 each. Six of them land. The remaining two sit Pending, because the two GPUs they need are held by six workers of somebody else's job, which is also waiting for two more.

Your allocation dashboard says the cluster is full. Your utilisation dashboard says nothing is computing. Both are correct. Twelve accelerators are rented, reserved, and burning money while a torch.distributed process group that will never form waits on a rendezvous that will never complete.

Nothing in the cluster is broken. This is the scheduler doing precisely what it was built to do.

The scheduler is doing exactly what you asked

kube-scheduler's unit of work is one pod. It takes a pod off a queue, filters nodes that cannot run it, scores the ones that can, picks the best, binds. Then it takes the next pod. The loop is beautiful for the workload Kubernetes grew up on, which is a stateless replica that is useful the instant it starts and does not care whether its siblings exist yet.

A distributed training job is the opposite workload. Worker 3 is worth nothing without workers 0 through 7. It will hold its GPU, open its NCCL connections, and block.

So when you submit an eight-pod job, you are not handing the scheduler one decision with eight parts. You are handing it eight unrelated decisions, and it will happily satisfy six of them. Add a second job and the failure gets properly expensive:

architecture-beta
    group cluster(logos:kubernetes)[Twelve GPU nodes]

    service helda(logos:nvidia)[Six GPUs held by Job A] in cluster
    service waita(server)[Job A needs two more] in cluster
    service heldb(logos:nvidia)[Six GPUs held by Job B] in cluster
    service waitb(server)[Job B needs two more] in cluster

    helda:B --> T:waita
    heldb:B --> T:waitb
    waita:R --> L:heldb
    waitb:L --> R:helda

Neither job can start. Neither job will release what it holds, because from the kubelet's point of view those six pods are running fine. The deadlock is stable, and it survives until a human notices the utilisation graph and deletes something.

This is the problem gang scheduling solves, and it is why the batch and AI corner of the ecosystem has spent years running something other than the default scheduler.

What we installed instead

Three answers have been in production for a long time, and they sit at different layers.

Volcano replaces kube-scheduler outright with an action and plugin pipeline, where a gang plugin enforces minAvailable inside the allocation loop. You get all-or-nothing, plus DRF fairness, queues and GPU sharing. You also get a second scheduler to operate, and every workload you point at it leaves the code path the rest of your cluster uses.

Apache YuniKorn also replaces the scheduler, and implements gangs by first binding placeholder pods that reserve the exact shape of each task group, then swapping the real pods into those reservations. It works, and it means your cluster briefly runs pods whose only job is to hold space.

Kueue stays out of placement entirely. It sits above the scheduler as an admission layer, holding a Workload object in a ClusterQueue until the quota exists, then unsuspending the Job and letting the stock scheduler place the pods one at a time. This is excellent quota governance and it is not gang scheduling: once Kueue admits the job, nothing stops the scheduler from placing six of eight pods and stalling.

There was also the coscheduling plugin in sigs.k8s.io/scheduler-plugins, which is the closest thing to a native answer and the direct ancestor of what finally shipped. It needed its own PodGroup CRD and a scheduler built with the out-of-tree plugin registry.

Every one of those is a reasonable choice with a real operational bill. The interesting change is that for the narrow case of "these pods must land together", you may no longer need to pay it.

What landed in the API

KEP-4671 put gang scheduling in kube-scheduler itself. It went alpha in Kubernetes 1.35 as scheduling.k8s.io/v1alpha2, reached beta in 1.37 as scheduling.k8s.io/v1beta1, and is targeting v1 in 1.38.

There are two objects, both defined in staging/src/k8s.io/api/scheduling/v1beta1/types.go. A Workload is the declarative shape, written by whatever controller owns your job:

apiVersion: scheduling.k8s.io/v1beta1
kind: Workload
metadata:
  name: llm-finetune
spec:
  podGroupTemplates:
    - name: workers
      schedulingPolicy:
        gang:
          minCount: 8

A PodGroup is the runtime instance the scheduler actually reasons about, with the policy copied down from the template. Pods join one through a dedicated pod spec field:

spec:
  schedulingGroup:
    podGroupName: llm-finetune-workers

spec.schedulingGroup.podGroupName is immutable, which the plugin leans on: the code notes that because the field cannot change, it never has to subscribe to pod update events, only adds.

The policy itself is a union with exactly two members, and the choice is permanent:

type PodGroupSchedulingPolicy struct {
	Basic *BasicSchedulingPolicy `json:"basic,omitempty"`
	Gang  *GangSchedulingPolicy  `json:"gang,omitempty"`
}

BasicSchedulingPolicy is an empty struct whose mere presence means "schedule these the normal way". GangSchedulingPolicy carries one field, minCount, and that field stays mutable after creation so a workload can be scaled. The API comment is refreshingly honest about what that means in practice: updates to minCount may not land in the current scheduling cycle, and they never retroactively affect pods that are already scheduled.

Two numbers worth knowing before you design around this: a Workload is capped at 8 pod group templates, and a composite hierarchy is capped at a tree depth of 4.

Beta does not mean on, and the gate alone is not enough

This is the detail most release coverage skips. In pkg/features/kube_features.go, the gate reads:

GenericWorkload: {
	{Version: version.MustParse("1.35"), Default: false, PreRelease: featuregate.Alpha},
	{Version: version.MustParse("1.37"), Default: false, PreRelease: featuregate.Beta},
},

Beta, and Default: false. Nothing described in this post happens on a stock 1.37 cluster. GenericWorkload is one gate covering what used to be the separate GangScheduling and WorkloadAwarePreemption gates, and you enable it yourself on three components, because all three have a part to play: the apiserver serves the objects, the scheduler honours them, and the controller manager runs the finalizer controller that keeps a PodGroup from vanishing under its pods.

There is a second switch, and it is the one that will waste your afternoon. Enabling the feature gate does not make the apiserver serve the API, because the group version is turned off independently in pkg/controlplane/instance.go:

betaAPIGroupVersionsDisabledByDefault = []schema.GroupVersion{
	storageapiv1beta1.SchemeGroupVersion,
	networkingapiv1beta1.SchemeGroupVersion,
	resourcev1beta1.SchemeGroupVersion,
	resourcev1beta2.SchemeGroupVersion,
	schedulingapiv1beta1.SchemeGroupVersion,
}

New beta APIs have been off by default since 1.24, and this one is no exception. Set the gate without the runtime config and kubectl get podgroups tells you the server has no such resource, which looks exactly like a broken feature gate. So you need both:

# kube-apiserver
--feature-gates=GenericWorkload=true
--runtime-config=scheduling.k8s.io/v1beta1=true
 
# kube-scheduler
--feature-gates=GenericWorkload=true
 
# kube-controller-manager
--feature-gates=GenericWorkload=true

For a local cluster, kind sets feature gates across all control plane components for you, so the whole thing is a config file:

kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
featureGates:
  GenericWorkload: true
runtimeConfig:
  "scheduling.k8s.io/v1beta1": "true"
nodes:
  - role: control-plane

Add CompositePodGroup: true and TopologyAwareWorkloadScheduling: true, plus "scheduling.k8s.io/v1alpha3": "true", if you want the hierarchical and topology pieces. Both are alpha, so that group version is off by default too.

Then confirm the API is actually being served before you debug anything else:

kubectl api-resources --api-group=scheduling.k8s.io

podgroups and workloads should both be listed. Once they are, you are done configuring the scheduler. There is no KubeSchedulerConfiguration to write, because the default plugin set in pkg/scheduler/apis/config/v1/default_plugins.go wires itself up from the gate:

if utilfeature.DefaultFeatureGate.Enabled(features.GenericWorkload) {
	applyGangScheduling(config)
}

which appends the GangScheduling plugin to the MultiPoint enabled list. The plugin registers itself across four extension points at once, and that is the whole mechanism.

How it actually works

The plugin is about 520 lines in pkg/scheduler/framework/plugins/gangscheduling/gangscheduling.go, and its type assertions tell you the entire design before you read a single function:

var _ fwk.EnqueueExtensions = &GangScheduling{}
var _ fwk.PreEnqueuePlugin = &GangScheduling{}
var _ fwk.PermitPlugin = &GangScheduling{}
var _ framework.PlacementFeasiblePlugin = &GangScheduling{}

All-or-nothing is not enforced in one place. It is enforced as a gate at the front of the queue, a feasibility check inside the placement cycle, and a barrier just before binding.

architecture-beta
    group sched(logos:kubernetes)[kube-scheduler]

    service pg(database)[PodGroup minCount eight]
    service gate(server)[PreEnqueue quorum gate] in sched
    service cycle(server)[Pod group placement cycle] in sched
    service permit(server)[Permit barrier] in sched
    service bind(logos:nvidia)[Bind all eight at once] in sched

    pg:R --> L:gate
    gate:R --> L:cycle
    cycle:R --> L:permit
    permit:B --> T:bind

PreEnqueue keeps the gang out of the queue

The first gate runs before a pod is allowed into the active queue at all. If a pod has no schedulingGroup, the plugin returns immediately and nothing changes. If it has one, the plugin looks the group up and counts:

allPodsCount := podGroupState.AllPodsCount()
if allPodsCount < int(policy.Gang.MinCount) {
	return fwk.NewStatus(fwk.UnschedulableAndUnresolvable,
		"waiting for minCount pods from a gang to appear in scheduling queue")
}

That is the cheapest and most valuable part of the whole feature. If your controller has created three of eight pods so far, none of the three gets considered. The scheduler does not place them, so they cannot reserve GPUs, so they cannot contribute to the deadlock above. If the PodGroup object itself does not exist yet, the pods are held with a similar status until it appears.

Note the status is UnschedulableAndUnresolvable, not plain Unschedulable. These pods are parked, not retried in a hot loop.

PlacementFeasible kills a doomed cycle early

In 1.37 the scheduler gained a pod group scheduling cycle: rather than evaluating one pod and binding it, it works through a group and accumulates proposed assignments. PlacementFeasible is consulted as that cycle progresses, and it does arithmetic on what is left:

if remaining+scheduled < minCount {
	return fwk.NewStatus(fwk.Unschedulable,
		fmt.Sprintf("minCount (%d) cannot be satisfied: %d scheduled, %d remaining", minCount, scheduled, remaining))
}
 
if scheduled < minCount {
	return fwk.NewStatus(fwk.Wait,
		fmt.Sprintf("minCount (%d) is not yet satisfied: %d scheduled, %d remaining", minCount, scheduled, remaining))
}

Once the pods still to be evaluated plus the pods already placed cannot reach minCount, there is no point evaluating the rest, so the cycle terminates. For a basic group, getMinCount returns 1 and the check degenerates into ordinary behaviour.

Alongside it, when the TopologyAwareWorkloadScheduling gate is on, a scoring plugin called PodGroupPodsCount scores whole candidate placements rather than individual nodes, favouring the placement that fits more of the group's pods. That is a genuinely new shape for this scheduler: the thing being scored is an assignment of a set of pods, not a node for one pod.

Permit is the actual barrier

Pods that survive placement reach the Permit extension point, which is the last stop before binding. Every gang member parks here:

scheduledPodsCount := podGroupState.ScheduledPodsCount()
if scheduledPodsCount < int(podGroup.Spec.SchedulingPolicy.Gang.MinCount) {
	unscheduledPods := podGroupState.UnscheduledPods()
	pl.handle.Activate(logger, unscheduledPods)
	return fwk.NewStatus(fwk.Wait, "waiting for minCount pods from a gang to be scheduled"), permitTimeoutDuration
}
 
assumedPods := podGroupState.AssumedPods()
for podUID := range assumedPods {
	waitingPod := pl.handle.GetWaitingPod(podUID)
	if waitingPod != nil {
		waitingPod.Allow(Name)
	}
}

Two things are happening there. A pod that arrives without quorum returns Wait and, on its way, calls Activate on its own unscheduled siblings to push them forward in the queue. The gang pulls itself along. Then the pod that completes the quorum walks every assumed pod in the group and calls Allow, releasing the whole set to bind together.

The timeout is not configurable:

// permitTimeoutDuration defines how long the gang pods should
// wait at the permit stage for a quorum before being rejected.
permitTimeoutDuration = 5 * time.Minute

Five minutes. A gang that gets most of the way there and stalls is rejected and recycled rather than holding assumed capacity indefinitely. Worth knowing when you are debugging a job that keeps almost starting.

Waking a stalled gang

The last extension point is EventsToRegister, which decides what cluster changes are worth re-examining a parked gang for. There are four, each with a queueing hint:

  • an unscheduled pod being added, which might be the one that completes the quorum
  • an assigned pod being added, same reason
  • a PodGroup being added, which unblocks pods that were waiting for their group object
  • a PodGroup being updated, specifically handled by isSchedulableAfterPodGroupUpdated, because a decrease in minCount can make a gang that was too large suddenly satisfiable

Scaling a workload down is therefore a first-class way to unstick it, and the scheduler notices without anyone restarting anything.

The parts that are not in the scheduler

Two pieces live elsewhere, and they exist for the same reason: a PodGroup must not disappear from under its pods.

An admission plugin at plugin/pkg/admission/scheduling/podgroupprotection stamps a finalizer on every newly created PodGroup. A matching controller in pkg/controller/scheduling/podgroupprotection removes that finalizer only once the group is being deleted and no active pod still references it.

What is deliberately absent is a controller that creates PodGroup objects from your Job. Upstream lists "take away responsibility to create pods from controllers" as a non-goal, so the wiring belongs to the workload controller: JobSet, Kueue, a training operator, or your own.

What it refuses to do

The KEP's non-goals are short and they matter more than the feature list: no fairness, and no multiple workload queues. Both are explicitly left to Kueue and Volcano.

So the layering is clean rather than competitive. Quota and ordering live above the scheduler. All-or-nothing placement now lives inside it.

architecture-beta
    group k8s(logos:kubernetes)[In tree path]

    service job(logos:pytorch-icon)[Training job]
    service kueue(cloud)[Kueue admits on quota]
    service pg(database)[PodGroup minCount] in k8s
    service gang(server)[GangScheduling plugin] in k8s
    service volcano(server)[Volcano owns both layers]

    job:R --> L:kueue
    kueue:R --> L:pg
    pg:R --> L:gang
    job:B --> T:volcano

Which means the decision is now narrower than "which batch scheduler do we run":

  • You run Kueue for team quota, queueing and job ordering. Nothing here replaces it, and the combination of Kueue for admission plus in-tree gangs for placement closes the gap Kueue alone always had.
  • You run Volcano or YuniKorn if you need DRF fairness, GPU sharing from the scheduler, or their topology and preemption behaviour. They still do more.
  • You can retire a scheduler replacement that you installed only to get all-or-nothing. That was always a heavy bill for one feature.

Before you turn the gate on

Five things I would want a teammate to know.

minCount is a placement-time guarantee, not a runtime invariant. The docs say it plainly: the scheduler never admits fewer than minCount pods during initial placement, but the running count can fall below it afterwards if pods are deleted or evicted. The PodGroupInitiallyScheduled condition is terminal once it goes True and never reverts. Keeping a gang alive is still your controller's job.

There is no migration from alpha, and the gate was renamed under you. The scheduling.k8s.io/v1alpha2 API is removed entirely, so v1alpha2 objects from a 1.35 or 1.36 cluster must be deleted before you upgrade. Backward conversion is not supported on a downgrade either. Separately, because GangScheduling was merged into GenericWorkload, operators have to remove GangScheduling from their feature gate configuration when upgrading, and manually re-enable it if they ever roll back to 1.36.

Every pod in a gang must share the same .spec.schedulerName. The scheduler validates this, and a mismatch rejects the entire group as unschedulable rather than just the odd pod out.

Placement is only guaranteed for the simple case. The scheduler cannot exhaustively analyse every placement permutation, and the KEP is explicit about the tiers. For a homogeneous group with no inter-pod dependencies, the algorithm is expected to find a placement whenever one exists. For heterogeneous groups, and for groups using affinity, anti-affinity or topology spread, a valid placement is not guaranteed even when one theoretically exists. Rejection messages are meant to say when that is the likely cause.

The hierarchical and topology pieces are still alpha. CompositePodGroup, which is how you express a driver group and a worker group that must land as one unit, is scheduling.k8s.io/v1alpha3 behind its own gate, and it depends on both GenericWorkload and TopologyAwareWorkloadScheduling, which is itself alpha since 1.36. If your workload is one homogeneous group of workers, you are on the beta path. If it is a Spark driver plus executors across racks, you are on the alpha path.

The short version

The reason your GPUs were allocated and idle was never a bug. It was a scheduler whose unit of work is one pod being handed a workload whose unit of work is eight.

Kubernetes 1.37 finally gives that scheduler a way to be told. Three gates in one 520-line plugin, two API objects, one feature gate you have to set yourself, and a five-minute timeout you should probably know about before your first 2am page.

Check what your batch scheduler is actually earning. If the answer is "all-or-nothing, and nothing else", there is now a smaller way to get it.

References

Primary sources for everything above, all pinned to the 1.37 release branch.

Specification and docs

Source

The alternatives

Chamod Shehanka
Author

Chamod Shehanka

Software Engineer II at Circles building cloud-native systems with Go and Kubernetes. CNCF & CD Foundation Ambassador, Jenkins GSoC mentor, and lead of Kubernetes Sri Lanka & GDG Sri Lanka.

Keep reading