
The judges your
search ranker deserves.
Calibrated human-rated relevance across 60+ languages. NDCG/DCG reports, query intent tagging, side-by-side ranker comparisons — the ground truth your team needs to ship ranking changes confidently.
Built for production, not just demos.
- 4-point and 7-point relevance scales (Perfect → Off-topic)
- Query intent + recall-class tagging
- Side-by-side ranker A/B (NDCG / DCG / MRR)
- Multilingual judges across 60+ languages
- Native-locale judges for region-specific queries
- Per-query calibration with house guidelines
- Long-tail + adversarial query coverage
- Catalog quality + canonicalisation feedback
How a typical engagement runs.
Guidelines
Co-design relevance guidelines + intent taxonomy with your search team.
Calibrate
Judges complete 200 gold judgments; tune anchors and edge cases.
Pilot
5k-query batch with per-judge IAA + per-intent confusion matrix.
Production
Weekly batches with live NDCG dashboards + ranker A/B reports.
Evolve
Quarterly guidelines refresh based on emerging queries + drift signals.
What you get in your bucket.
Questions, answered.
4-point or 7-point relevance?
Default to 4-point (Perfect/Excellent/Fair/Bad) for speed + IAA. Graduate to 7-point for fine-grained ranker A/B where small wins matter.
Can you judge our long-tail queries?
Yes — we curate a representative + long-tail mix and stratify reporting. Long-tail typically reveals the biggest ranker wins.
Code / docs search?
We have PhD-CS and engineering judges for code/doc search, including code-context understanding and API-doc query intent.
Multilingual?
60+ languages with native locale judges. Cross-lingual relevance (e.g. English query, Hindi result) is supported with appropriate guidelines.
Ready to build
AI you can trust?
Talk to a solutions architect — get a pilot scoped in 48 hours.