Model evaluation demonstrates superior retrieval accuracy for a compact trilingual embedding architecture trained from scratch, indicating the viability of small distilled models.
We present mentee-embed-v1, a 41M-parameter trilingual text embedding model for Arabic, English, and Urdu trained entirely from scratch — no pre trained backbone, no fine-tuning of existing models. Training uses a custom 50K BPE tokenizer and two stages: masked language modeling followed by relational knowledge distillation from multilingual-e5-base. On in-batch retrieval benchmarks, mentee-embed-v1 achieves avg MRR@10 of 0.585, beating all-MiniLM-L6-v2 (0.396) while remaining functional on all three languages. Code, weights, and tokenizer are publicly released.
No takes yet. Share an insight, caveat, or question.
Shah et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: