Managed AI Operations · Search Evaluation

The judges your
search ranker deserves.

Calibrated human-rated relevance across 60+ languages. NDCG/DCG reports, query intent tagging, side-by-side ranker comparisons — the ground truth your team needs to ship ranking changes confidently.

Capabilities

Built for production, not just demos.

  • 4-point and 7-point relevance scales (Perfect → Off-topic)
  • Query intent + recall-class tagging
  • Side-by-side ranker A/B (NDCG / DCG / MRR)
  • Multilingual judges across 60+ languages
  • Native-locale judges for region-specific queries
  • Per-query calibration with house guidelines
  • Long-tail + adversarial query coverage
  • Catalog quality + canonicalisation feedback
Specs at a glance
Scales4 / 7-point ESCS-style
Throughput200k judgments/wk
IAA>0.85
Languages60+
VerticalsRetail · News · Docs · Code · Medical
DeliveryJSONL · Parquet · API · dashboard
Workflow

How a typical engagement runs.

Step 1

Guidelines

Co-design relevance guidelines + intent taxonomy with your search team.

Step 2

Calibrate

Judges complete 200 gold judgments; tune anchors and edge cases.

Step 3

Pilot

5k-query batch with per-judge IAA + per-intent confusion matrix.

Step 4

Production

Weekly batches with live NDCG dashboards + ranker A/B reports.

Step 5

Evolve

Quarterly guidelines refresh based on emerging queries + drift signals.

Deliverables

What you get in your bucket.

Judgment dataset (JSONL / Parquet)
NDCG / DCG / MRR per ranker
Per-intent + per-language slices
Ranker A/B win-rate matrix
Per-judge calibration metrics
Catalog quality feedback log
FAQ

Questions, answered.

4-point or 7-point relevance?

Default to 4-point (Perfect/Excellent/Fair/Bad) for speed + IAA. Graduate to 7-point for fine-grained ranker A/B where small wins matter.

Can you judge our long-tail queries?

Yes — we curate a representative + long-tail mix and stratify reporting. Long-tail typically reveals the biggest ranker wins.

Code / docs search?

We have PhD-CS and engineering judges for code/doc search, including code-context understanding and API-doc query intent.

Multilingual?

60+ languages with native locale judges. Cross-lingual relevance (e.g. English query, Hindi result) is supported with appropriate guidelines.

More within Managed AI Operations

Ready to build
AI you can trust?

Talk to a solutions architect — get a pilot scoped in 48 hours.