Evaluating Agentic AI: Dell AI Solutions with Open Models
-
Mitch Lewis
AI has rapidly moved from a emerging technology to a top priority for enterprise organizations. Continual advances in model development produce larger, more capable, and more reliable LLMs. For enterprise adoption, agentic applications provide tangible business value, enabling AI to solve challenges and complete meaningful business tasks with minimal human involvement.
As agentic AI becomes the center of enterprise strategy, organizations increasingly recognize the need to run AI on premises to maintain data sovereignty, preserve operational control, and achieve predictable, cost-effective performance. Yet, evaluating LLM usefulness and accuracy on agentic tasks remains challenging, as most existing benchmarks focus on reasoning abstractions rather than a model’s ability to execute real enterprise workflows. Static benchmarks can also leak into training data, turning them into exercises of memorization rather than true capability tests.
To address this gap, Signal65 and Kamiwaza established a new AI benchmark to measure model performance on enterprise-focused agentic tasks, validated on modern on-premises infrastructure including Dell AI servers and Broadcom high-performance Ethernet fabrics. These open, standards-based environments reflect how organizations deploy AI today: securely under their control and optimized for both performance and cost. Testing covered 31 models and over 170,000 test conversations, resulting in over 5.5 billion tokens processed.
Key Highlights:
Key findings:
Model Size: In general, accuracy was seen to improve with model size, with the highest scores attributed to very large models with over 100B parameters. Small models (<10B parameters) showed a clear deficiency across most agentic tasks. Some models in the 30 to 100B parameter range, however, such as Llama-3.1-70B-Instruct and Qwen3-30B-A3B (thinking mode) outperformed much larger models, demonstrating compelling options for organizations with limited infrastructure.
Quantization: FP8 quantization does not appear to have adverse effects on agentic capabilities. Across FP8 quantized and full weight model pairs tested, the FP8 variations consistently achieved similar or even slightly greater accuracy.
Thinking: Models with thinking capabilities were generally found to be more accurate in achieving agentic tasks than similar non-thinking models. Non-thinking models, however, became highly competitive when provided basic hints and context clues, offering a possible alternative to the high token usage and cost associated with thinking models.
Research commissioned by:


