Quickstart
vindex scores LLM output in Hindi, Hinglish and eight other Indian scripts. The first two checks are free: no model, no key, no network.
$ pip install vindex $ vindex check "Mumbai kahan hai?" "मुंबई महाराष्ट्र में है।" ✗ FAIL script_mismatch score 0.00 prompt is code-mixed; response in devanagari, not Roman script.
The same check from Python:
from vindex import script_adherence result = script_adherence("Mumbai kahan hai?", "Mumbai Maharashtra mein hai.") result.passed # True result.label # "matched" result.reason # "prompt is code-mixed; response is Roman script, as expected."
Install
| Command | Gives you |
|---|---|
pip install vindex | script_adherence, check_trace, the CLI. One dependency. |
pip install "vindex[similarity]" | calibrated_similarity (sentence-transformers, torch). |
pip install "vindex[judge]" | indic_judge with the default Groq judge. |
pip install openai / anthropic / litellm | Other judge providers. Install only the SDK you use. |
Python 3.10 or newer.
Results
Every check returns a frozen MetricResult:
| Field | Type | Meaning |
|---|---|---|
score | float | 0 to 1, higher is better. |
passed | bool | The check's verdict at its default threshold. |
label | str | Stable outcome name. Branch on this. |
reason | str | One readable sentence explaining the verdict. |
detail | dict | Raw evidence: counts, thresholds, judge reasoning, model id. |
The CLI
vindex run evaluates a dataset file and prints a pass-rate summary plus every failing case with its reason. It exits with status 1 if any metric's pass rate is below --fail-under (default 1.0).
$ vindex run cases.jsonl $ vindex run cases.jsonl -m script,judge --judge openai:gpt-4o $ vindex run cases.csv -m similarity -l hi --fail-under 0.9 $ vindex run cases.json -f json -o report.json
| Option | Default | Description |
|---|---|---|
-m, --metrics | script | Comma-separated: script, similarity, judge, trace. |
-l, --language | — | Default language for similarity: en, hi, hinglish. A row's own language wins. |
--judge | groq | groq, openai[:model], anthropic[:model], litellm:<model>. |
--encoder | mpnet-v2 | sentence-transformers model for similarity. |
--strict-language | off | Also fail Romanized-Hindi prompts answered in English. |
--fail-under | 1.0 | Minimum pass rate per metric before exiting 1. |
-f, --format | table | table, json or markdown. |
-o, --output | — | Also write the full JSON report to a file. |
--no-trace-auto | — | Don't run check_trace automatically on rows that have a trace. |
Two more commands for quick checks: vindex check PROMPT RESPONSE scores one pair, and vindex detect TEXT shows which scripts a string is written in. Both take --json.
Dataset format
JSONL (one object per line), a JSON list, or CSV with a header row. Column names are matched case-insensitively, and the usual names from other eval tools work without renaming:
| Field | Accepted names | Used by |
|---|---|---|
| prompt | prompt question input user_input query | all |
| response | response answer output actual_output completion | all |
| gold | gold expected expected_output reference ground_truth | similarity, judge |
| trace | trace reasoning judge_reasoning | trace |
| language | language lang | similarity |
| id | id case_id name | reports |
{"id": "refund-017", "prompt": "Mera refund kab tak aayega?", "response": "आपका रिफंड 5-7 दिनों में आएगा।"}
{"id": "boil-01", "prompt": "समुद्र तल पर पानी किस तापमान पर उबलता है?", "response": "100°C", "trace": "...sea floor..."}
Rows without a gold reference are skipped by similarity; rows without a trace are skipped by check_trace. Skips are counted in the report, never treated as passes.
GitHub Actions
When $GITHUB_STEP_SUMMARY is set, vindex run appends a markdown report to the job summary.
- run: pip install vindex - run: vindex run evals/cases.jsonl --fail-under 0.95
script_adherence
Checks that the response is written in the script the prompt used. No reference answer and no model needed. Native-script prompts (Hindi in Devanagari, Tamil in Tamil script, …) expect the same script back; Roman-script and code-mixed prompts expect Roman script back.
| Label | Passed | When |
|---|---|---|
matched | yes | Response is in the expected script. |
mixed | yes | Response is code-mixed in a compatible way. |
script_mismatch | no | Response is in a different script. |
language_mismatch | no | Only with strict_language_check=True: Romanized Hindi prompt, English reply. |
no_script_signal | no | No letters in any recognized script (emoji, digits, punctuation). |
empty | no | Prompt or response is blank. |
strict_language_check is off by default because Hinglish detection is a 14-word heuristic and can misfire on English prompts (“Who directed Se7en?” contains “se”). Turn it on when your prompts are reliably Hinglish.
check_trace
Reads a judge's reasoning (from any judge, not only vindex's) and flags it when it took the wrong reading of a known ambiguous Hindi term in the source. Deterministic and free.
from vindex import check_trace r = check_trace( source="समुद्र तल पर पानी किस तापमान पर उबलता है?", trace="The question asks for the boiling point at the sea floor...", ) r.label # "misread_detected"
Labels: no_misread_detected, misread_detected, empty. The bundled dictionary covers 69 terms: polysemous nouns, misleading compounds, tense-flipping time words (कल), fractions and Indian number words. It detects misreadings of those terms, not mistranslation in general. Pass your own trap_words list to extend it.
check_trace_llm_fallback(source, trace, judge) adds an opt-in LLM pass for misreadings a dictionary can't catch. It only sends the segments that don't align, and flags when the model's answer is ambiguous.
calibrated_similarity
Embedding similarity against a gold answer, using a threshold calibrated for each encoder and language (en, hi, hinglish) instead of a fixed 0.5. Requires vindex[similarity].
from vindex import calibrated_similarity r = calibrated_similarity( gold="भारत की राजधानी नई दिल्ली है।", response="नई दिल्ली भारत की राजधानी है।", language="hi", )
Labels: similar, dissimilar, low_discrimination, empty.
Most encoders separate right from wrong answers poorly in Hindi. When the calibrated AUC for your encoder and language is below min_auc, the result is low_discrimination with passed=False rather than a misleading verdict. Default encoder: paraphrase-multilingual-mpnet-base-v2. MuRIL is blocked unless you pass allow_muril=True, because it scores almost every pair near 0.99.
indic_judge
An LLM judge for correctness, with a rubric written in Hindi that treats Romanized Hindi as valid. Reference-free by default. Pass gold for reference mode: an exact match skips the model call entirely.
from vindex import indic_judge r = indic_judge( question="समुद्र तल पर पानी किस तापमान पर उबलता है?", answer="समुद्र तल पर पानी 100 डिग्री सेल्सियस पर उबलता है।", ) r.passed # True r.detail["judge_model_id"] # "openai/gpt-oss-120b" r.detail["judge_reasoning"] # the judge's own reasoning — feed it to check_trace
- Conservative: passes only on the top rubric score at high judge confidence. Anything less is
flagged. - Reproducible: temperature 0, and the judge's model id is stored in every result.
- Self-judging guard: pass
answering_model_idand vindex warns when the judge comes from the same model family.
Labels: matched, flagged, judge_error, empty.
Choosing a judge
from vindex.judge_model import GroqJudge, OpenAIJudge, AnthropicJudge, LiteLLMJudge GroqJudge() # GROQ_API_KEY, default openai/gpt-oss-120b OpenAIJudge("gpt-4o") # OPENAI_API_KEY AnthropicJudge("claude-sonnet-4-5") # ANTHROPIC_API_KEY LiteLLMJudge("gemini/gemini-1.5-pro") # any LiteLLM provider
Anything with a model_id property and a call(prompt) -> str method works as a judge. Pin a dated model version, never a “latest” alias, and never judge a model with itself. Indic competence doesn't track overall model size, so check your judge against the bundled benchmark before trusting it.
Calibrate on your own data
The labelled data behind the shipped thresholds is included, so you can measure your own encoder or judge instead of trusting ours.
from vindex import calibrate, indic_judge from vindex.datasets import ( load_similarity_benchmark, similarity_benchmark_for_calibration, score_judge_benchmark, ) # Encoder: fit a threshold for your own similarity function cases = load_similarity_benchmark() correct, wrong = similarity_benchmark_for_calibration(cases, my_similarity, language="hi") fit = calibrate(correct, wrong) fit.threshold, fit.warnings # Judge: agreement with two human graders on 62 labelled traces report = score_judge_benchmark(lambda t: "correct" if indic_judge(t.question, t.answer).passed else "wrong") report.agreement_grader1, report.agreement_grader2
Integrations
Ready-made adapters live in examples/integrations:
- DeepEval —
ScriptAdherenceMetric, usable withevaluate()andassert_test(). - promptfoo — a Python assertion file and example config.
- Ragas — a
SingleTurnMetric.
Each adapter is about 50 lines. To wrap a different check, swap the script_adherence call for any other vindex function.
Limitations
- Scripts: Devanagari, Bengali, Gurmukhi, Gujarati, Odia, Tamil, Telugu, Kannada, Malayalam and Latin. Text in any other script has no bucket of its own.
- Hinglish detection is a function-word heuristic. It's only used for the opt-in
language_mismatchlabel. - Similarity languages: thresholds ship for
en,hiandhinglish. For anything else, calibrate on your own data. - check_trace catches misreadings of its dictionary terms only. Grow the list from what you see in production.
- indic_judge is an LLM and can be wrong. Its rubric is Hindi-specific; other languages have not been validated.
Missing a language or found a false positive? Open an issue.