Bilingual Retrieval Workbench
synthetic corpusselftest pending
Retrieval quality, measured per language

An English-first RAG pipeline quietly fails Arabic learners.

This page builds a bilingual retrieval stack in your browser and scores it: chunking strategies compared on outcome, hybrid dense and sparse retrieval with a reranking pass, Arabic normalisation shown as a metrics delta on a fixture eval set, cross-tenant probes run adversarially, and a cost projection whose Arabic multiplier comes from a tokenizer trained here rather than from a guess.

Every number below is recomputed on load. Nothing is hardcoded next to a label. The corpus, the courses, the learners and the organisations are all invented for this demo.

training tokenizer, chunking corpus, building three indexes