Quickstart

vindex scores LLM output in Hindi, Hinglish and eight other Indian scripts. The first two checks are free: no model, no key, no network.

$ pip install vindex
$ vindex check "Mumbai kahan hai?" "मुंबई महाराष्ट्र में है।"

  ✗ FAIL  script_mismatch  score 0.00
  prompt is code-mixed; response in devanagari, not Roman script.

The same check from Python:

from vindex import script_adherence

result = script_adherence("Mumbai kahan hai?", "Mumbai Maharashtra mein hai.")
result.passed   # True
result.label    # "matched"
result.reason   # "prompt is code-mixed; response is Roman script, as expected."

Install

CommandGives you
pip install vindexscript_adherence, check_trace, the CLI. One dependency.
pip install "vindex[similarity]"calibrated_similarity (sentence-transformers, torch).
pip install "vindex[judge]"indic_judge with the default Groq judge.
pip install openai / anthropic / litellmOther judge providers. Install only the SDK you use.

Python 3.10 or newer.

Results

Every check returns a frozen MetricResult:

FieldTypeMeaning
scorefloat0 to 1, higher is better.
passedboolThe check's verdict at its default threshold.
labelstrStable outcome name. Branch on this.
reasonstrOne readable sentence explaining the verdict.
detaildictRaw evidence: counts, thresholds, judge reasoning, model id.

The CLI

vindex run evaluates a dataset file and prints a pass-rate summary plus every failing case with its reason. It exits with status 1 if any metric's pass rate is below --fail-under (default 1.0).

$ vindex run cases.jsonl
$ vindex run cases.jsonl -m script,judge --judge openai:gpt-4o
$ vindex run cases.csv -m similarity -l hi --fail-under 0.9
$ vindex run cases.json -f json -o report.json
OptionDefaultDescription
-m, --metricsscriptComma-separated: script, similarity, judge, trace.
-l, --language—Default language for similarity: en, hi, hinglish. A row's own language wins.
--judgegroqgroq, openai[:model], anthropic[:model], litellm:<model>.
--encodermpnet-v2sentence-transformers model for similarity.
--strict-languageoffAlso fail Romanized-Hindi prompts answered in English.
--fail-under1.0Minimum pass rate per metric before exiting 1.
-f, --formattabletable, json or markdown.
-o, --output—Also write the full JSON report to a file.
--no-trace-auto—Don't run check_trace automatically on rows that have a trace.

Two more commands for quick checks: vindex check PROMPT RESPONSE scores one pair, and vindex detect TEXT shows which scripts a string is written in. Both take --json.

Dataset format

JSONL (one object per line), a JSON list, or CSV with a header row. Column names are matched case-insensitively, and the usual names from other eval tools work without renaming:

FieldAccepted namesUsed by
promptprompt question input user_input queryall
responseresponse answer output actual_output completionall
goldgold expected expected_output reference ground_truthsimilarity, judge
tracetrace reasoning judge_reasoningtrace
languagelanguage langsimilarity
idid case_id namereports
{"id": "refund-017", "prompt": "Mera refund kab tak aayega?", "response": "आपका रिफंड 5-7 दिनों में आएगा।"}
{"id": "boil-01", "prompt": "समुद्र तल पर पानी किस तापमान पर उबलता है?", "response": "100°C", "trace": "...sea floor..."}

Rows without a gold reference are skipped by similarity; rows without a trace are skipped by check_trace. Skips are counted in the report, never treated as passes.

GitHub Actions

When $GITHUB_STEP_SUMMARY is set, vindex run appends a markdown report to the job summary.

- run: pip install vindex
- run: vindex run evals/cases.jsonl --fail-under 0.95

script_adherence

script_adherence(prompt, response, strict_language_check=False) -> MetricResult

Checks that the response is written in the script the prompt used. No reference answer and no model needed. Native-script prompts (Hindi in Devanagari, Tamil in Tamil script, …) expect the same script back; Roman-script and code-mixed prompts expect Roman script back.

LabelPassedWhen
matchedyesResponse is in the expected script.
mixedyesResponse is code-mixed in a compatible way.
script_mismatchnoResponse is in a different script.
language_mismatchnoOnly with strict_language_check=True: Romanized Hindi prompt, English reply.
no_script_signalnoNo letters in any recognized script (emoji, digits, punctuation).
emptynoPrompt or response is blank.

strict_language_check is off by default because Hinglish detection is a 14-word heuristic and can misfire on English prompts (“Who directed Se7en?” contains “se”). Turn it on when your prompts are reliably Hinglish.

check_trace

check_trace(source, trace, trap_words=None) -> MetricResult

Reads a judge's reasoning (from any judge, not only vindex's) and flags it when it took the wrong reading of a known ambiguous Hindi term in the source. Deterministic and free.

from vindex import check_trace

r = check_trace(
    source="समुद्र तल पर पानी किस तापमान पर उबलता है?",
    trace="The question asks for the boiling point at the sea floor...",
)
r.label  # "misread_detected"

Labels: no_misread_detected, misread_detected, empty. The bundled dictionary covers 69 terms: polysemous nouns, misleading compounds, tense-flipping time words (कल), fractions and Indian number words. It detects misreadings of those terms, not mistranslation in general. Pass your own trap_words list to extend it.

check_trace_llm_fallback(source, trace, judge) adds an opt-in LLM pass for misreadings a dictionary can't catch. It only sends the segments that don't align, and flags when the model's answer is ambiguous.

calibrated_similarity

calibrated_similarity(gold, response, language, encoder_name="", allow_muril=False, min_auc=0.7) -> MetricResult

Embedding similarity against a gold answer, using a threshold calibrated for each encoder and language (en, hi, hinglish) instead of a fixed 0.5. Requires vindex[similarity].

from vindex import calibrated_similarity

r = calibrated_similarity(
    gold="भारत की राजधानी नई दिल्ली है।",
    response="नई दिल्ली भारत की राजधानी है।",
    language="hi",
)

Labels: similar, dissimilar, low_discrimination, empty.

Most encoders separate right from wrong answers poorly in Hindi. When the calibrated AUC for your encoder and language is below min_auc, the result is low_discrimination with passed=False rather than a misleading verdict. Default encoder: paraphrase-multilingual-mpnet-base-v2. MuRIL is blocked unless you pass allow_muril=True, because it scores almost every pair near 0.99.

indic_judge

indic_judge(question, answer, gold=None, judge=None, answering_model_id=None) -> MetricResult

An LLM judge for correctness, with a rubric written in Hindi that treats Romanized Hindi as valid. Reference-free by default. Pass gold for reference mode: an exact match skips the model call entirely.

from vindex import indic_judge

r = indic_judge(
    question="समुद्र तल पर पानी किस तापमान पर उबलता है?",
    answer="समुद्र तल पर पानी 100 डिग्री सेल्सियस पर उबलता है।",
)
r.passed                      # True
r.detail["judge_model_id"]     # "openai/gpt-oss-120b"
r.detail["judge_reasoning"]    # the judge's own reasoning — feed it to check_trace

Labels: matched, flagged, judge_error, empty.

Choosing a judge

from vindex.judge_model import GroqJudge, OpenAIJudge, AnthropicJudge, LiteLLMJudge

GroqJudge()                              # GROQ_API_KEY, default openai/gpt-oss-120b
OpenAIJudge("gpt-4o")                     # OPENAI_API_KEY
AnthropicJudge("claude-sonnet-4-5")       # ANTHROPIC_API_KEY
LiteLLMJudge("gemini/gemini-1.5-pro")     # any LiteLLM provider

Anything with a model_id property and a call(prompt) -> str method works as a judge. Pin a dated model version, never a “latest” alias, and never judge a model with itself. Indic competence doesn't track overall model size, so check your judge against the bundled benchmark before trusting it.

Calibrate on your own data

The labelled data behind the shipped thresholds is included, so you can measure your own encoder or judge instead of trusting ours.

from vindex import calibrate, indic_judge
from vindex.datasets import (
    load_similarity_benchmark, similarity_benchmark_for_calibration, score_judge_benchmark,
)

# Encoder: fit a threshold for your own similarity function
cases = load_similarity_benchmark()
correct, wrong = similarity_benchmark_for_calibration(cases, my_similarity, language="hi")
fit = calibrate(correct, wrong)
fit.threshold, fit.warnings

# Judge: agreement with two human graders on 62 labelled traces
report = score_judge_benchmark(lambda t: "correct" if indic_judge(t.question, t.answer).passed else "wrong")
report.agreement_grader1, report.agreement_grader2

Integrations

Ready-made adapters live in examples/integrations:

Each adapter is about 50 lines. To wrap a different check, swap the script_adherence call for any other vindex function.

Limitations

Missing a language or found a false positive? Open an issue.