Scaling AI with Dell PowerEdge XE9680
-
Brian Martin
Scaling AI operations across organizations is no longer a purely technical discussion, it is a strategic requirement. Enterprises adopting large-scale AI infrastructure face critical choices around performance, cost, and time-todeployment. This paper analyzes scaling results from Dell PowerEdge XE9680 clusters configured with NVIDIA H200 or AMD MI300X accelerators, each interconnected with Broadcom BCM57608 Thor 2 network controllers and Dell PowerSwitch Z9864F fabric. The findings illustrate not only technical performance, but also how Dell Technologies enables organizations to achieve sustainable AI scaling with flexibility and confidence.
Key Insights
Strategic TCO Advantages: Dell PowerEdge XE9680 deployments with Broadcom Ethernet achieved 19–28% lower TCO versus alternative approaches, driven by operational simplicity, Ethernet standardization, and optimized power/cooling.
Choice Without Compromise: AMD configurations deliver superior memory capacity and cost efficiency, while NVIDIA configurations provide unmatched inference throughput and software maturity. Both run seamlessly on the Dell PowerEdge XE9680, with Dell providing unified management and enterprise support.
Future-Proof Infrastructure: Ethernet’s ubiquity ensures continuity across compute, storage, and networking. Dell and Broadcom’s roadmap to 800GbE and 1.6TbE guarantees that investments made today align with future AI technology cycles.
Sustained Network Utilization: Broadcom NICs with Dell PowerSwitch maintained lossless, near-peak efficiency across collective workloads, ensuring that GPUs remain fully utilized as clusters scale to 64+ accelerators.
AMD Instinct: AMD MI300X GPUs stand out for cost optimization and memory capacity, making this an ideal choice for organizations focused on very large model training and high scale inference.
NVIDIA Hopper: NVIDIA H100 and H200 GPUs excel in computational throughput, especially for mixedprecision training and inference with its advanced Transformer Engine.
Software Ecosystem: NVIDIA CUDA remains the dominant first mover with its long-established and mature ecosystem, providing developers a rich set of tools and libraries, resulting in faster time-to-deployment and reduced development friction.
Research commissioned by:


