A nodeSelector Is a Capacity Request
Somewhere in a cluster I look after, an autoscaler had provisioned a full compute node to run two small controller pods. Nothing else would ever schedule onto it. Nothing else could.
The cause was two lines of YAML that had been copied, correctly, from a cluster where they cost nothing.
Two lines, copied correctly
We run the same ingress controller on four clusters. Two of them run a meaningful amount of
arm64 workload and keep arm64 capacity around permanently. On those two, someone had pinned
the controller to arm64 — a nodeSelector on the architecture label, plus the matching
toleration for the arm64 taint. Sensible: the capacity is there, it’s cheaper per unit, the
pods fit in the gaps.
When the controller was rolled out to the other two clusters, the values file was copied. Of course it was. That’s how you get consistency, and the diff between the four files looked right in review precisely because the lines matched.
But those two clusters have no arm64 workloads. The node pool exists — the platform configuration is uniform — and until that moment nothing had ever asked it for anything.
What a scheduling constraint means to an autoscaler
Here’s the part that’s easy to say and easy to forget under review.
To the Kubernetes scheduler, a nodeSelector is a filter: it narrows the set of existing
nodes a pod may land on. If none match, the pod stays Pending. That’s the mental model most
people carry, and in a fixed-size cluster it’s the whole story — you notice immediately,
because the pod doesn’t run.
To a provisioning autoscaler, the same nodeSelector is an order. A pending pod with
unsatisfiable constraints is not an error condition; it’s a work item. The autoscaler reads
the constraint, computes the cheapest instance that satisfies it, and buys one.
So the failure mode inverts. In a fixed cluster, an over-constrained pod is loudly broken. In an autoscaled cluster, it is silently expensive. Everything comes up green. The deployment reports Available, the controller serves traffic, no alert fires. The only evidence is a node in the list with one workload on it and an instance-hours line on the bill.
In our case one cluster got a general-purpose arm64 node and the other got a memory-optimized one — roughly four vCPU and thirty-odd gigabytes of RAM — to host two pods whose combined requests were a rounding error against that.
Nothing needed it
The controller image publishes both linux/amd64 and linux/arm64 in its manifest list.
There is no native extension, no compiled sidecar, no architecture-specific behaviour. The
pin wasn’t buying compatibility or performance; it was an artifact of where the manifest was
first written.
Deleting the two lines rescheduled both replicas onto existing nodes, and the autoscaler deprovisioned the now-empty node in each cluster. That’s the whole fix. No image change, no version bump, no migration.
Why review didn’t catch it
This is the bit worth generalizing, because the review was not lazy.
The correctness of a scheduling constraint is a property of the cluster, not of the
manifest. You cannot evaluate nodeSelector: {kubernetes.io/arch: arm64} by reading it.
The same characters are free on one cluster and cost you a node on another, and the file
gives you no way to tell which. Review compares the change against its neighbours, the
neighbours match, and matching is exactly the wrong signal here.
The related conflation, which I’ve now watched several people make including myself:
- A toleration is permission. It says “I’m allowed onto nodes with this taint.” Adding one to a pod that never encounters the taint changes nothing and costs nothing.
- A nodeSelector (or a required node affinity) is demand. It says “I will not run anywhere else.” Adding one to a cluster that can’t satisfy it obliges someone to go get capacity.
They’re usually written adjacent to each other, in the same block, by the same person, in a single copied hunk. One of them is inert and one of them spends money.
The same asymmetry applies to anything else that narrows placement: zone pins, dedicated
node-pool labels, GPU or accelerator selectors, instance-family requirements, and
topologySpreadConstraints with whenUnsatisfiable: DoNotSchedule — which can force a new
node purely to satisfy a spread rule.
The question to ask
When copying a workload’s manifest to another cluster, go through every placement constraint and ask one thing:
Does this cluster already run something that satisfies this? If not, I am not writing a preference — I am placing an order.
If the answer is no and the workload doesn’t actually require the constraint, delete it. If the workload does require it, that’s fine, but now you know you’re adding a node and you can size the decision honestly.
And after you remove one: check that the node actually goes away. Consolidation isn’t instant, and anything else that drifted onto that node — or any pod on it carrying a “don’t disrupt me” annotation — will keep it alive and keep you paying for a fix you think you already landed.
Takeaways
- In an autoscaled cluster, a scheduling constraint is a purchase order. The scheduler filters; the provisioner procures. Same YAML, opposite consequences.
- Over-constrained pods fail loudly on fixed clusters and quietly on elastic ones. If your intuition was formed on fixed clusters, it is now inverted and nothing will tell you.
- Tolerations are free, selectors are not. They travel together in copied blocks. Treat them differently.
- A constraint can’t be reviewed in isolation from the cluster it lands on. “It matches the other environments” is the specific reasoning that produces this bug.
- Copying manifests between clusters needs a per-constraint pass, not a per-file comparison — the file being identical is the thing that hides it.
- Verify the node disappears after you unpin. Otherwise you’ve paid for the fix twice: once in the incident and once in the node that never left.