Trusted Answers at Enterprise Scale

Dell AI Solutions with Open Models

Authors:

Mitch Lewis
Mitch Lewis
Brian Martin

Brian Martin

September 15, 2026

27 of 91 models achieved at least 95% accuracy with 32K-token document inputs.
3 of 54 models exceeded 95% accuracy with 200K-token document inputs.
At 32K, reasoning raised one model’s accuracy from 58.8% to 96.6%.

Enterprise AI has moved from pilots to production workflows, where its most common job is answering questions from the organization’s own documents: contracts, policies, operational reports, and records. The business value of these systems, from faster decisions and lower operating costs to institutional knowledge available on demand, depends on a deceptively simple question: can the answers be trusted? An AI system that retrieves the wrong clause from a contract, mis-totals figures across reports, or fabricates an answer when none exists creates operational, financial, and compliance risk, and undermines the confidence needed for adoption to spread.

The stakes are rising as enterprises move toward agentic AI, systems that do not simply answer questions but carry out multi-step business processes. Agents retrieve information at nearly every step of their work: checking a policy, querying records, reconciling figures before acting. Retrieval reliability is therefore not one AI capability among many; it is the foundation on which agentic initiatives stand or fail, because small per-step error rates compound across a workflow, and a fabricated answer becomes a fabricated action.

Signal65 and Kamiwaza developed RIKER (Retrieval Intelligence and Knowledge Extraction Rating) to compare how accurately models find facts, combine information across documents, and recognize missing information within supplied document collections. This report helps enterprise teams shortlist models and configurations for further testing. It evaluates model behavior within provided context windows, not complete retrieval-augmented generation (RAG) pipelines or autonomous agent workflows. Testing used synthetic commercial leases, HR records, and facility field reports, with document inputs up to 200K tokens—roughly 150,000 words.

Open models make their model weights available for deployment under their respective licenses, allowing organizations to run them on infrastructure they control. Signal65 conducted the open-model evaluation in its AI Lab using RIKER, developed with Kamiwaza. Dell Technologies provided access to PowerEdge hardware and technical expertise; Broadcom contributed network optimization collaboration and the Ethernet technology used in the tested stack. Proprietary models were evaluated through their native APIs.

What the testing found, and why it matters to the business:

High accuracy over large document inputs is achievable, but uncommon: Only three of 54 models exceeded 95% overall accuracy at 200K tokens. These controlled results support model selection; they do not establish the accuracy of a complete enterprise retrieval system.

Advertised context windows are not a procurement metric: Models with similar advertised limits diverged sharply as workloads grew. Evaluate retrieval stability at the document scale the organization actually operates, not from token limits alone.

Cross-document work is the key bottleneck: At 200K context, multi-document aggregation accuracy declined by an average of 25.6% relative to 32K performance, more than twice the rate of single-document tasks. Pilots built on simple lookups can overstate production readiness.

Retrieval reliability determines agentic readiness: Because agents retrieve information throughout multi-step workflows, per-step accuracy compounds. Assuming independent failures and five retrieval-dependent steps, 90% reliability at each step yields about a 59% probability of completing the workflow without a retrieval error. Validate retrieval stability before delegating consequential actions.

Deployment configuration is a business lever, not an IT detail: At 32K context, enabling reasoning (“thinking”) improved one model from 58.8% to 96.6% accuracy, a gain of 37.8 percentage points, or 64.3% relative. Test reasoning modes independently; the effect varied by model and context size.

Model selection can materially reduce hallucination risk: The best models reliably acknowledged when information was absent, and that behavior remained comparatively stable as context grew. “Can the model find the answer?” and “does it know when the answer is absent?” should be evaluated separately.

Evaluate model, configuration, and infrastructure together: In the tested environment, Dell PowerEdge compute and a Broadcom-based 400GbE RoCEv2 fabric formed the on-premises stack used for open-model evaluation. The benchmark demonstrates that the workloads ran on this configuration; it does not isolate the performance contribution of individual infrastructure components.

Research commissioned by:

Dell Technologies logo