All posts
rlhf
indic
case-study

RLHF at scale across 11 Indic languages

Our pipeline for collecting, calibrating, and shipping preference data in Hindi, Tamil, Telugu, Marathi, Bengali, and more.

P. NairJuly 29, 2026

Building reliable preference data in a dominant language like English is hard. Doing it across 11 Indic languages — with dialect, code-switching, and cultural nuance — is a different problem.

In this post we walk through:

  1. Recruitment and calibration of native-speaker raters
  2. Dialect-aware rubric design
  3. Inter-rater agreement tracking
  4. How we feed signal back to the reward model

The full pipeline is in production with two frontier labs as of Q2.

Ready to build
AI you can trust?

Talk to a solutions architect — get a pilot scoped in 48 hours.