Local Vector Search at 10-Billion Scale
-
Brian Martin
Enterprise AI has moved beyond model experimentation and into operational deployment. At this stage, success is no longer defined only by model quality, but by whether the infrastructure can deliver data to those models reliably, consistently, and cost-effectively at production scale. For organizations deploying retrieval-augmented generation (RAG), vector search, and real-time inference, the storage layer has a direct impact on user experience, service levels, infrastructure efficiency, and ultimately business value. When storage becomes the bottleneck, AI investments underperform, latency rises, and expensive compute resources sit underutilized.
In prior work, AI Storage Pipeline Acceleration with Dell PERC13 established that the Dell PERC H975i (PERC13) controller can deliver industry-leading synthetic performance, reaching up to 56 GB/s of sequential throughput and 13 million IOPS. Those results demonstrated the platform’s technical ceiling. The more important question for enterprise decision makers, however, is whether that performance holds up in real deployments, where access patterns are irregular, concurrency is high, and data protection cannot be compromised. That is the difference between an impressive benchmark and a platform that can be trusted in production.
This paper answers that question using a 10-billion-vector FAISS index, sized to exceed system memory and validate storage-based queries, on a Dell PowerEdge R770 equipped with dual PERC H975i controllers, 32 NVMe drives in RAID5 configuration, and 512 GiB of DDR5 memory. Under this realistic workload, the platform delivered 51.6 GB/s while serving concurrent queries at scale and maintaining RAID5 protection, demonstrating high-performance AI infrastructure does not have to come at the expense of resilience. With the right storage architecture and proper tuning, organizations can accelerate retrieval-intensive AI workloads, protect critical data, and improve infrastructure efficiency while reducing deployment risk and increasing confidence in production-scale AI.
Key Highlights:
Research commissioned by:


