OpenShift AI on ROSA, priced through the marketplace.
OpenShift AI on ROSA pricing reads as three stacked layers on the AWS marketplace bill. The ROSA cluster base is the OpenShift entitlement priced at vCPU per hour on the marketplace; the Red Hat OpenShift AI operator is an additional entitlement priced against the AI workload footprint; the GPU instance hours are AWS infrastructure charges that move with the data science team's actual training and inference cadence. The buyer who sizes all three layers against the workload signs an envelope that reads cleanly; the buyer who sizes only one or two pays the missing layer at the next marketplace invoice.
The three layers of the bill.
Red Hat OpenShift Service on AWS is the hosted OpenShift offering Red Hat operates jointly with Amazon. ROSA is sold through the AWS marketplace under either the classic deployment model with dedicated control plane VMs the buyer pays for or the hosted control plane model where the control plane sits on Red Hat managed infrastructure and the buyer pays only for worker capacity. The ROSA bill is metered in vCPU per hour against the cluster's running worker footprint and consolidated through the buyer's AWS marketplace invoice1.
Red Hat OpenShift AI is the renamed Red Hat OpenShift Data Science product. RHOAI bundles the Jupyter notebook stack, the model serving stack with KServe and the OpenVINO and TGIS backends, the pipeline orchestration with Kubeflow Pipelines, and the operator that ties the components into the OpenShift cluster. On ROSA, RHOAI is an additional entitlement that prices against the AI workload footprint, typically expressed as a node level entitlement on the worker nodes that schedule RHOAI managed pods or as a flat platform charge2.
The third layer is the GPU instance bill. ROSA worker capacity is provisioned on EC2 instance types that the cluster machine pool names. GPU work runs on p3, p4, p5, g4, g5, or g6 instance families depending on the model size and the precision target. The GPU instance hours flow through the AWS portion of the marketplace bill at the instance's on demand or reserved rate. The data science team's training cadence drives this layer more than the platform engineering team's planning does.
Where the three layers drift apart.
The three layers drift apart in characteristic patterns that the practice observes across RHOAI on ROSA engagements. The first drift is the worker pool that scales for inference traffic without resizing the RHOAI entitlement. A buyer that signed RHOAI against a planned twenty node AI worker pool and added eight nodes during an inference traffic surge has scaled the OpenShift footprint and the GPU bill but not necessarily the RHOAI line. The audit reads the AI worker pool against the RHOAI envelope and surfaces the eight node delta.
The second drift is the GPU experiment that ran for thirty hours on a p4d instance during a hackathon and was never decommissioned. The instance kept running because the autoscaler did not see it as idle on the resource graph. The GPU hours show up on the AWS bill at the on demand rate, which is the most expensive rate the instance carries. The RHOAI entitlement is unchanged because the instance was already in the worker pool envelope; the drift is purely an AWS infrastructure cost. A buyer who runs RHOAI on ROSA needs a hackathon decommission playbook more than a buyer who runs RHOAI on bare metal3.
The third drift is the data scientist who provisions a notebook workbench on an oversized GPU instance for prototype work that does not need the silicon. The workbench pattern in RHOAI defaults to a node selector that prefers GPU nodes when GPUs are requested. A request for one GPU on a g5.48xlarge instance reads at the bill as a 192 vCPU instance for the prototype, even if the prototype's compute footprint would fit on a g4dn.xlarge. The mitigation is a workbench profile that constrains the instance family for prototype work and reserves the larger instance families for training runs that demonstrate the need.
| Workload | ROSA layer | RHOAI layer | GPU layer |
|---|---|---|---|
| Inference at steady traffic | Stable | Stable | Stable |
| Inference at surge | Scales | True up | Scales |
| Training run, 30 hours, p4d | Stable | Stable | Spikes |
| Hackathon left running | Stable | Stable | Burns |
| Oversized workbench prototype | Stable | Stable | Wastes |
Reserved instances and the marketplace concession.
The RHOAI on ROSA bill compresses with two structural levers that the buyer should price at the procurement table. The first lever is the AWS reserved instance commitment on the GPU instance families the data science team uses. A one or three year reservation on a g5 or p4 instance family at standard utilisation can reduce the GPU bill by thirty to fifty percent against the on demand rate, depending on the term and the family. The reservation is procured separately from the RHOAI contract and is owned by the cloud finance function, but the procurement narrative should include it.
The second lever is the Red Hat marketplace concession against the consolidated ROSA and RHOAI line. A buyer with sufficient AWS marketplace commit can negotiate a concession against the Red Hat line that flows through the AWS Enterprise Discount Programme. The concession is not advertised at the marketplace shelf price; it is procured directly with the Red Hat field organisation and routed through the marketplace billing channel. The concession band on a multi year ROSA plus RHOAI plus marketplace commit is materially better than the band on a one year on demand procurement, and the procurement function should price both options at the negotiation table.
The reserved instance and marketplace concession together can shift the bill profile from a volatile on demand month to a predictable run rate that the procurement function can defend at quarterly review. The trade off is committed capacity. A buyer who reserves a GPU instance family the data science team subsequently migrates away from is paying for an unused commit through the term4.
The procurement narrative at signature.
The procurement narrative that the buyer should bring to the ROSA and RHOAI table names the AI worker pool capacity in nodes and in cores, the RHOAI entitlement envelope, the GPU instance family and reservation commit, the workbench profile constraints, and the hackathon decommission cadence. The narrative is a one page document that the AI platform team, the procurement function, and the cloud finance function all sign before the marketplace order is placed.
For the broader cross product reading, see the OpenShift practice hub, the OpenShift AI LLM workload read for the workload sizing detail, the GPU entitlement mechanics read for the worker pool counting, the bare metal versus virtualized cost read for the on premise comparison, and the RHEL on AWS marketplace economics read for the broader marketplace concession dynamics. For the engagement protocol, see renewal negotiation and contact.
OpenShift AI on ROSA pricing reads as three layers stacked through one marketplace bill. The buyer who sizes all three at signature, reserves the GPU capacity that fits the training cadence, and constrains the workbench profile pays the line once. The buyer who lets the bill compose itself across the three layers pays the marketplace invoice at the worst possible rate.
Notes & references
- 1. Red Hat OpenShift Service on AWS product documentation and AWS marketplace billing notes accessed across 2025 and 2026. ROSA is sold through the AWS marketplace under classic and hosted control plane variants priced at vCPU per hour against the worker footprint.
- 2. Red Hat OpenShift AI product documentation accessed across 2025 and 2026. RHOAI is the renamed Red Hat OpenShift Data Science product and bundles the Jupyter notebook, KServe, OpenVINO, TGIS, and Kubeflow Pipelines stacks under a node level entitlement.
- 3. AWS EC2 GPU instance family documentation and on demand and reserved pricing tables accessed across 2025 and 2026. The on demand rate is materially higher than the reserved rate for the same instance family.
- 4. Practice observation across RHOAI on ROSA engagements: the most common procurement miss is the lack of a hackathon decommission playbook and the lack of a workbench profile constraint that bounds the prototype instance family.
- 5. The AWS Enterprise Discount Programme and the Red Hat marketplace concession overlap in characteristic ways on multi year ROSA and RHOAI commits; the concession band is materially better than the marketplace shelf price.
Preparing a response? The practice keeps a one-page Red Hat audit response checklist — what to acknowledge, what to preserve, and what not to volunteer in the first fourteen days after the letter arrives.