Resources · Benchmarks

Independent benchmarks.

Methodology, datasets, and head-to-head results. Reproducible, open where possible.

Benchmark

Indic-Eval v2: 22-language LLM benchmark

An open multilingual benchmark across 22 Indic languages.

May 20, 2026
Benchmark

Voice-AgentContainment Benchmark (VACB)

Standardised tasks + scoring for voice-agent containment.

Apr 26, 2026
Benchmark

Hallucination-in-Code: 12-language eval suite

Code generation hallucination eval across 12 languages.

Mar 30, 2026
Benchmark

Enterprise LLM Performance Benchmark

Comprehensive evaluation of enterprise LLMs across reasoning, accuracy, latency, and response consistency.

Mar 18, 2026

Ready to build
AI you can trust?

Talk to a solutions architect — get a pilot scoped in 48 hours.