Loading...


Updated 21 Sep 2026 • 5 mins read

Allocating GPU costs in a shared Kubernetes cluster is harder than CPU allocation because GPUs are assigned whole, sit idle between jobs, and cost far more per node. This guide covers why standard allocation fails on GPUs, which metrics to charge on, sharing methods, and a practical charge model.
A single eight-GPU node can cost more per month than forty CPU nodes running your entire web platform. Put three of those in a shared cluster, invite four machine learning teams to use them, and you have created the most expensive attribution problem in your infrastructure, one where being wrong by ten percent means being wrong by tens of thousands of dollars.
The awkward part is that the techniques that work for CPU and memory allocation do not transfer cleanly. GPUs are handed out whole rather than sliced, they sit reserved and idle between jobs far more than CPUs do, and utilization inside the device is invisible to Kubernetes by default. So the usual approach, charge each workload for the share of the node it requested, produces numbers that are technically defensible and practically infuriating, because a team that reserved a GPU for a notebook they left open all weekend pays the same as a team that trained a model on it.
The short answer GPU costs in a shared Kubernetes cluster are allocated by charging each team for GPU-time reserved rather than GPU-utilization achieved, because Kubernetes assigns GPUs as whole, exclusive devices that no other workload can use while they are held. The practical model is: take the true hourly cost of each GPU node from your cloud bill, divide it across the GPUs on that node, charge each workload for the GPU-hours it held, and allocate the node's CPU, memory, and idle overhead separately. Utilization metrics from DCGM should be reported alongside the charge as an efficiency signal rather than used as the billing basis, since charging on utilization removes any incentive to release idle GPUs.
Standard container cost allocation, which we cover in our guide to container cost visibility, works because CPU and memory are divisible and elastic. A pod can request 250 millicores, several pods share a core, and the scheduler packs them densely. GPUs work nothing like this by default, and four properties explain the difference.
This is the decision that determines whether your model changes behavior, and there are three candidates.
Charging on GPU utilization, the percentage of device capacity actually used, feels fairest and is the wrong answer. A team that holds a GPU at 5 percent utilization for a week has denied that device to everyone else for a week, and billing them for 5 percent of it tells them that hoarding is free. Utilization-based charging systematically rewards the exact behavior you most need to discourage.
Charging on GPU-hours reserved, the time a workload held the device regardless of what it did with it, reflects the real economics: in an exclusive-assignment model, holding is consuming. It also creates the right pressure, because the cheapest thing a team can do is release GPUs it is not using, which is precisely the outcome you want.
Charging on job completion or output, such as cost per training run, is a useful reporting view for the teams themselves but a poor allocation basis, because it cannot account for the idle time between jobs that your invoice is very much paying for.
So charge on reserved GPU-hours, and report utilization beside it. The charge drives release behavior; the utilization figure tells teams whether the GPUs they legitimately hold are being used well, and tells the platform team where to invest in sharing.
Here is a model that holds up in practice. It separates the GPU from everything else on the node, because mixing them produces numbers nobody can verify.
| Component | How to allocate it | Notes |
|---|---|---|
| GPU devices | Node GPU cost divided by GPU count, charged per GPU-hour held | The dominant line; charge on reservation |
| CPU and memory on GPU nodes | By requests, as with any other workload | Usually small relative to the GPU |
| Storage and data transfer | By actual usage per namespace | Training data movement can be significant |
| Idle GPU capacity | To the platform team, or proportionally by a published rule | Needs an explicit owner, not an unallocated bucket |
| Cluster overhead | Proportionally across allocated cost | Control plane, monitoring, networking |
Deriving the per-GPU hourly rate is straightforward but worth doing carefully. Take the node's real hourly cost from your billing data, not the list price, since Spot, reservations, and committed-use discounts change it substantially. Subtract a reasonable allowance for the CPU, memory, and storage portion of that node, then divide the remainder by the number of GPUs. That gives a defensible per-GPU-hour figure, and because it derives from actual billed cost, it stays correct when your discount mix changes.
In most shared AI clusters, idle GPU time is larger than any single team's usage. Notebooks left running overnight, jobs that finished but whose pods were not cleaned up, development environments provisioned for a project that ended, and capacity held in reserve for a training run scheduled next week all consume devices that bill continuously.
Two things fix it, and both require the allocation model to be in place first. Technical controls do the heavy lifting: idle timeouts that reclaim notebook GPUs after a period of inactivity, queueing systems that hand out devices per job rather than per person, and automatic cleanup of completed pods. Ownership does the rest: idle GPU cost should appear on someone's report, normally the platform team that controls provisioning and can right-size the fleet, rather than disappearing into an unallocated line that grows quietly.
The general principle is the one we apply to idle Kubernetes resources everywhere: waste with no owner is waste nobody removes. GPUs simply make the consequence arrive faster.
A GPU charge model gets scrutinized far harder than a CPU one, because the numbers are large enough to affect team budgets. Three practices earn the trust that makes it stick.
GPU cost allocation in a shared cluster is a variant of the container allocation problem with the volume turned up: the same principles, applied to a device that is assigned whole, costs an order of magnitude more, and spends much of its life reserved but unused. The model that works charges on GPU-hours held rather than utilization achieved, derives the per-GPU rate from real billed node cost, allocates CPU, storage, and overhead separately, and gives idle capacity a named owner.
Put utilization beside the charge and the practice becomes self-correcting. Teams that can see they are paying for 400 GPU-hours and using 12 percent of them do not need a policy to change their behavior; they need the number. That is the whole return on the work, because in an AI cluster the difference between allocated and utilized is not a reporting detail, it is most of the bill.
Charge each workload for the GPU-hours it held, using a per-GPU hourly rate derived from the real billed cost of the node divided by its GPU count. Allocate the node's CPU, memory, storage, and cluster overhead separately, and give idle GPU capacity an explicit owner, usually the platform team that controls provisioning.
On reservation, meaning GPU-hours held. Kubernetes assigns GPUs exclusively, so a workload holding a device blocks every other workload whether it uses it or not. Charging on utilization would make hoarding free and remove any incentive to release idle GPUs. Report utilization alongside the charge as an efficiency metric instead.
Four reasons: GPUs are assigned whole rather than divisible into fractions like CPU millicores; their internal utilization is invisible to Kubernetes without DCGM telemetry; a GPU node can cost ten to fifty times a general-purpose node, so errors are expensive; and GPUs sit idle between jobs far more often, especially under interactive notebooks.
Through NVIDIA's DCGM exporter, which publishes per-GPU metrics such as utilization, memory usage, temperature, and power draw to Prometheus. Kubernetes itself only knows that a GPU has been assigned to a container, not what is happening inside the device, so this telemetry has to be deployed separately.
Time-slicing lets several containers take turns on one GPU without partitioning memory, so cost is usually divided equally among the containers sharing it. Multi-Instance GPU physically partitions a GPU into isolated instances with dedicated memory, so each workload can be charged for the fraction of the device its instance represents.
Give it an explicit owner rather than leaving it unallocated, normally the platform team that controls provisioning and can right-size the fleet. Pair that with technical controls: idle timeouts that reclaim notebook GPUs, job queueing that assigns devices per job rather than per person, and automatic cleanup of completed pods.
Start with showback. Let teams see and challenge their GPU numbers for a quarter so the model can be corrected before budgets depend on it. Because GPU costs are large, an allocation model that arrives as a surprise invoice tends to generate disputes about the data rather than decisions about usage.