Your model speaks Hindi.
Your evals don't.
vindex marks Indic LLM output the way a teacher who reads the language would: wrong script, misread words, answers that only look right. Then it fails your build.
Hindi, Hinglish and 8 more Indian scripts · two checks need no LLM and no API key
Mistakes an English eval marks correct.
Whoever reviews your outputs probably reads Hindi, so these look fine. They aren't — for the user, or for the judge doing the grading.
Start free. Add a judge when you need one.
Every check returns the same result: a score, pass or fail, a stable label to branch on, and one sentence saying why.
| Check | What it marks wrong | Needs |
|---|---|---|
script_adherence |
Replies in a different script from the question: Devanagari to a Hinglish user, English to a Tamil one, emoji instead of an answer. script_adherence(prompt, response) |
free no LLM, no gold |
check_trace |
An LLM judge's reasoning that misread an ambiguous Hindi term. 69 built in — समुद्र तल, कल, उत्तर — and you can add your own. check_trace(question, judge_reasoning) |
free deterministic |
calibrated_similarity |
Answers that drift from the gold answer, with thresholds calibrated per encoder and language. Refuses to grade where the encoder can't discriminate. calibrated_similarity(gold, response, language="hi") |
local encoder gold answer |
indic_judge |
Wrong answers, judged with a Hindi rubric that accepts Romanized Hindi. Temperature 0, conservative, and the judge model is recorded in every result. indic_judge(question, answer, judge=OpenAIJudge()) |
LLM API key Groq · OpenAI · Anthropic · LiteLLM |
Write a question. Write an answer. Get it marked.
This is script_adherence running in your browser, with the same rules as the Python package. Nothing leaves the page.
Mark a whole dataset. Fail the PR if it slips.
$ vindex run examples/cases.jsonl vindex 0.6.0 cases.jsonl · 6 cases ✗ script_adherence 83.3% 5/6 passed ✗ check_trace 0.0% 0/1 passed, 5 skipped Failures ● mumbai-devanagari-reply · script_adherence · script_mismatch prompt Mumbai kahan hai? response मुंबई महाराष्ट्र में है। reason prompt is code-mixed; response in devanagari, not Roman script. ● boiling-hindi-trace · check_trace · misread_detected prompt समुद्र तल पर पानी किस तापमान पर उबलता है? reason trace appears to misread 1 known ambiguous term(s): समुद्र तल. FAILED script_adherence, check_trace below 100%
# .github/workflows/evals.yml on: [pull_request] jobs: indic-evals: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: actions/setup-python@v5 with: { python-version: "3.12" } - run: pip install vindex - run: vindex run evals/cases.jsonl --fail-under 0.95
import pytest from vindex import script_adherence @pytest.mark.parametrize("prompt", [ "Mera refund kab aayega?", "मेरा रिफंड कब आएगा?", "என் பணம் எப்போது திரும்ப வரும்?", ]) def test_replies_in_users_script(bot, prompt): result = script_adherence(prompt, bot(prompt)) assert result.passed, result.reason
from vindex import indic_judge from vindex.judge_model import OpenAIJudge, AnthropicJudge, LiteLLMJudge indic_judge(q, a) # Groq, openai/gpt-oss-120b indic_judge(q, a, judge=OpenAIJudge("gpt-4o")) indic_judge(q, a, judge=AnthropicJudge()) indic_judge(q, a, judge=LiteLLMJudge("gemini/gemini-1.5-pro")) # or from the CLI $ vindex run cases.jsonl -m judge --judge anthropic
Reads the exports you already have
JSONL, JSON or CSV. Columns like input, actual_output and expected_output from DeepEval, promptfoo and Ragas work without renaming.
Fails the build, not silently
Exits 1 when any check's pass rate is under --fail-under. Skipped rows are counted, never passed.
Reports where you look
Terminal, JSON or markdown — and a summary in the GitHub Actions run, automatically.
Plugs into your test suite
Ready adapters for DeepEval, promptfoo and Ragas in examples/, or call it straight from pytest.