OpenShift AI GPU entitlement, counted at the worker core.
OpenShift AI GPU entitlement mechanics in 2026 read on the worker cores of accelerator capable nodes rather than on the GPU device itself. The counting trap that produces most of the variance at renewal is the assumption that the contract reads on the count of GPUs in the fleet; in reality it reads on the cores of the nodes those GPUs sit inside. A node with two large GPUs and ninety six worker cores carries a larger line than a node with eight smaller GPUs and forty eight worker cores even though the second node has four times the accelerator count. The sizing decision drives the entitlement.
The accelerator capable worker core, not the GPU device.
OpenShift AI runs on OpenShift Container Platform and provisions accelerator workloads to the subset of worker nodes that carry GPU hardware. The Red Hat subscription line reads on those worker nodes the same way the broader OpenShift platform line reads on every worker node: on the worker cores reported to the cluster scheduler, in pair units of physical cores, with the standard counting rules that apply across the platform1. The GPU device sitting inside the node is invisible to the Red Hat subscription line. The cores on the node are visible.
The consequence is that the node sizing decision drives the entitlement. A buyer who places eight NVIDIA A100 accelerators inside a node with forty eight worker cores reads the AI line on forty eight cores. A buyer who places two NVIDIA H100 accelerators inside a node with ninety six worker cores reads the AI line on ninety six cores. The second buyer has fewer accelerators in the fleet but carries a larger subscription footprint per node, and the AI workload that runs on the larger node carries a larger Red Hat line even though its raw accelerator count is lower.
The relationship between accelerator count and core count is set by the node vendor and the platform team rather than by Red Hat. A dense GPU node from one vendor pairs eight accelerators with a comparatively modest CPU floor; a different node from a different vendor pairs four accelerators with a larger CPU floor. The buyer who optimises the node sizing for accelerator density reads a different OpenShift AI line than the buyer who optimises for general purpose CPU capacity on the AI nodes.
The pool definition, and the boundary trap.
OpenShift AI counts on the AI pool, the named subset of worker nodes that hosts accelerator workloads. The pool boundary in the contract record is where the renewal arithmetic turns. A pool defined as "every node carrying a GPU operator" reads larger than a pool defined as "every node under active AI workload assignment", and the difference between those two readings is often the single largest movable quantity at the renewal table2.
The boundary trap arises because the GPU operator can be installed on a node well before the node hosts a real AI workload. A platform team that provisions ten future GPU nodes for an expected training campaign, installs the operator across all ten, and then runs the training campaign on only four of them has a pool defined by operator presence at ten and a pool defined by active workload at four. The audit reads against operator presence absent a contract clause that says otherwise. The mitigation is a contract clause that explicitly defines the AI pool as the nodes under active workload assignment.
The pool boundary also matters across mixed clusters. A cluster that runs CPU only application workloads and a small AI subset reads the AI line on the AI subset rather than the full cluster. The audit reads the AI pool through node labels or node selectors that the operational team applies to the AI nodes. The contract should name the node label that defines the AI pool and the buyer should retain the label history as part of the audit evidence record3.
Shared scheduling, and the hyperthread question.
Shared scheduling on the AI pool produces a second layer of counting nuance. Where a single GPU node hosts multiple smaller AI workloads through GPU time slicing or NVIDIA Multi Instance GPU partitioning, the underlying node still reads on its physical core count rather than on the fractional accelerator slices carved out of the GPU. The buyer who carves a single H100 into seven Multi Instance GPU partitions and serves seven small models from one node still pays the AI line on the worker cores of that one node. The accelerator partitioning does not multiply the entitlement.
Hyperthreading raises a separate question that mirrors the broader OpenShift counting rules. The Red Hat platform line on OpenShift reads on physical cores rather than on logical hyperthreads. The same convention applies on the AI pool. A node that reports ninety six logical CPUs to the kernel and forty eight physical cores reads the AI line on forty eight in the standard reading. The buyer who reads against the logical CPU count overstates the entitlement and pays a line that the contract does not require.
For the broader counting mechanics across the OpenShift estate, see the related core counting and hyperthreading article in the practice library. The conventions apply identically to the AI pool with the additional layer that the AI pool is a defined subset of the larger fleet rather than the fleet itself.
Three counting traps on the GPU node fleet.
Three counting traps produce most of the exposure observed across OpenShift AI engagements in the trailing twelve months.
The first trap is the assumption that the line reads on the GPU count rather than on the core count. A buyer who forecasted the AI line on a model of forty GPUs at a per GPU rate signs against an actual line of forty cores per node times the node count, which on dense AI nodes can be a materially larger number than the per GPU forecast suggested. The mitigation is a core level forecast at signature that captures the node sizing decision rather than the accelerator count alone.
The second trap is the operator installed pool with no active workload. The pool reads on the nodes that carry the operator absent a contract clause that scopes the pool to active workload assignment. The mitigation is the pool boundary language described in section two and a quarterly inventory of nodes under active workload assignment versus nodes with the operator installed for capacity planning.
The third trap is the hyperthread overcounting on AI nodes. A buyer who reads the AI line on logical CPU count rather than physical core count pays a line approximately twice the size of the contract reading. The recovery at audit when the platform team can produce the physical core count is typically the full hyperthread differential, which on dense AI nodes can amount to thirty to fifty percent of the AI line.
| Pattern | Frequency | Reading |
|---|---|---|
| Core level forecast captured node sizing | 3 of 8 | Pays |
| Pool boundary tied to active workload | 2 of 8 | Pays |
| Per GPU forecast against per core actual | 2 of 8 | Traps |
| Operator installed but not workload bearing | 1 of 8 | Traps |
| Hyperthread overcount on AI pool | 0 of 8 | Latent |
Reading GPU entitlement against the node sizing.
OpenShift AI GPU entitlement is read on the accelerator capable worker cores in the named AI pool. The reading at signature should reflect the node sizing decision, the pool boundary, the active workload assignment, and the physical core convention. The buyer who signs against a per GPU model pays a line that bears no consistent relationship to the contract reading. The buyer who signs against the worker core count on each AI node pays a line that matches what the audit reads.
The discipline at signature sets four protections that hold across the term. The AI pool is defined by node selector and node label in the contract record. The node sizing assumption is captured at the worker core level rather than the GPU device level. The active workload boundary is named, with a clause that operator installation alone does not enrol a node in the paid pool. The physical core convention is referenced explicitly so the audit cannot read the line on logical CPU count.
For the broader cross product reading, see the OpenShift practice hub, the LLM workload economics read for the workload calendar across training, fine tuning, and inference, the ACM pricing read for governance across the AI and non AI cluster fleet, the ACS pricing read for security on the same fleet, and the RHEL on IBM Power read for the parallel accelerator counting conventions on Power hardware. For the engagement protocol, see subscription assessment and contact.
Notes & references
- 1. Red Hat OpenShift subscription guide and OpenShift AI documentation, accessed across 2025 and 2026. The platform line reads on worker cores; OpenShift AI adds the AI line on the accelerator capable subset of those cores rather than on the accelerator device itself.
- 2. Pool definition in the OpenShift AI contract reading turns on node selectors and node labels that the platform team applies to the AI nodes. A contract clause that defines the AI pool as active workload assignment rather than operator presence is the single most movable mitigation at signature.
- 3. Node labels and node selectors are retained as part of the OpenShift cluster state and the etcd backup record. The labels at the audit reading should be supported by a label history that captures the AI pool boundary across the trailing twelve months.
- 4. Practice observation across eight OpenShift AI engagements settled in the trailing twelve months. The most common forecast error is the per GPU model applied against a contract reading that accrues on physical worker cores.
- 5. Concession bands and trailing twelve month figures refer to the practice observation across signed contracts. The eighty two percent audit exposure reduction in marginalia is the trailing twelve month average across defenses settled.
Preparing a response? The practice keeps a one-page Red Hat audit response checklist — what to acknowledge, what to preserve, and what not to volunteer in the first fourteen days after the letter arrives.