CoreWeave Delivers Higher Throughput and Lower Cost per Token on NVIDIA Blackwell: A Three-Year TCO and Tokenomics Analysis
-
Mitch Lewis
Enterprise AI programs increasingly run on rented infrastructure. Rather than absorbing the capital cost, lead times, and ongoing maintenance of standing up hardware on-premises, organizations can reserve dedicated GPU capacity on demand. They can also skip infrastructure management altogether and call hosted models through serverless APIs priced per token. Both paths are widely available across cloud providers. What differs substantially is what they cost.
Pricing for AI cloud services varies significantly between traditional hyperscaler clouds and specialized, AI-optimized clouds such as CoreWeave, and at production scale those differences compound into material sums. This study quantifies them. It models a 5,000-GPU NVIDIA Blackwell deployment over a three-year term against three anonymized hyperscaler clouds, using public on-demand pricing, MLPerf Inference v6.0 throughput results, and published serverless token rates. For inference workloads it accounts for performance as well as price, since how efficiently a provider runs a given model shapes the real cost per token as much as the published rate does.
The results show CoreWeave consistently delivers the lowest total cost of ownership for NVIDIA Blackwell GPU deployments, doing so in six of the seven configurations compared, and that advantage widens once inference performance is taken into account. CoreWeave’s cost advantage additionally extends past GPU rental to storage, Kubernetes, and other supporting services, and again to Serverless Inference.
CoreWeave delivers lower total cost than hyperscaler clouds for 3-year AI deployments on NVIDIA Blackwell GPUs:
- Up to 41% lower cost on NVIDIA HGX B200
- Up to 65% lower cost on NVIDIA GB200 NVL72
- Up to 52% lower cost on NVIDIA HGX B300
CoreWeave also delivers lower cost per token for inference, driven by a combination of lower list pricing and industry-leading performance:
- CoreWeave achieves between 4% and 124% greater throughput than comparable references when running DeepSeek-R1 on NVIDIA GB200 NVL72.
- CoreWeave’s performance and pricing advantages compound to increase per-token cost advantages. Compared to the most competitively priced hyperscaler competitors, CoreWeave’s per-token cost advantage widens from 22% based on infrastructure pricing alone to 65% once performance is factored in.
CoreWeave leads on Serverless Inference pricing.
- Across 17 models with a direct hyperscaler comparison, CoreWeave Serverless Inference is lower priced on 11 and identically priced on 4.
- CoreWeave provides 26% lower median input token costs and 15% lower median output token costs across 17 model comparisons.
Research commissioned by:


