Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

Tag: Multilingual

Multilingual evaluation papers that call for domain expertise or specialist review (73 papers).

Papers in tag: 73

Running a Multilingual study?

Post a Job →

Research Utility Snapshot

Evaluation Modes

  • Automatic Metrics (12)
  • Llm As Judge (4)
  • Human Eval (3)

Human Feedback Types

  • Pairwise Preference (9)
  • Critique Edit (2)
  • Red Team (2)

Required Expertise

  • Multilingual (20)
  • Coding (6)
  • Medicine (3)
Source-Free MT Evaluation Is Not MT Evaluation

Baban Gain, Ramakrishna Appicharla, Asif Ekbal · Aug 21, 2026 · Citations: 0

Pairwise Preference Automatic Metrics Multilingual
  • Reference-based metrics remain the standard choice in machine translation evaluation, partly because quality estimation methods often correlate less well with human judgments.
  • As a result, source-free, reference-based evaluation has become the practical norm, even though it is unfaithful to the definition of translation adequacy and unfair to systems whose outputs preserve the source meaning while differing from…
HealMed: Multilingual Evaluation of Large Language Models in Medicine

Yingjian Chen, Fan Gao, Sherry T. Tong, Haoyu Zhang, Aosong Feng, Kevin W. Jin · Aug 20, 2026 · Citations: 0

Critique Edit MedicineMultilingual
  • We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine.
  • The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions.
When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use

Siddharth Chauhan, Thomas Butler, Abhishek Singhania, Pankaj Porwal, Honey Gupta · Aug 12, 2026 · Citations: 0

Automatic Metrics Multilingual
  • We revisit post-training strategies for mitigating ALM and find that, in our benchmark, supervised fine-tuning (SFT) provides a strong baseline, substantially improving argument language consistency and end-to-end function call accuracy.
HULAT2 at MER-TRANS 2026: Governed Multi-Agent Simplification for Spanish Easy-to-Read Generation

Lourdes Moreno, Paloma Martínez, Marco Antonio Sanchez-Escudero, Miguel Domínguez-Gómez · Jul 2, 2026 · Citations: 0

Automatic Metrics Multilingual
  • RUN1 and RUN2 used a LangGraph-based multi-agent workflow combining Gemini 2.5 Flash and RigoChat-7B-v2, parallel generation strategies, internal quality signals, Event-Condition-Action routing, controlled editing and traceable decisions.
  • These results indicate that, in this task setting, signal-guided multi-agent routing outperformed the linear regeneration baseline.
HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety

Navaneeth Sangameswaran, Preetham S, Ashmiya Lenin · Jul 2, 2026 · Citations: 0

Red Team Automatic Metrics Multilingual
  • We present HaloGuard 1.0, an open-weights implementation of the constitutional-classifier paradigm for input safety.
  • Across seven prompt-safety benchmarks, HaloGuard 1.0-0.8B attains the best average F1 (90.9) of any open guard we evaluate, outperforming baselines up to 27B parameters (over 30 times larger) while holding false-positive rate (FPR) to 4.3…
Safety Targeted Embedding Exploit via Refinement

Joshua Adrian Cahyono · Jul 2, 2026 · Citations: 0

Red Team Automatic Metrics CodingMultilingual
  • We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training.
  • To study this phenomenon, we introduce STEER (Safety Targeted Embedding Exploit via Refinement), a gradient-guided attack that identifies words contributing most strongly to the model's refusal behavior and iteratively translates them into…
Disentangling Speaker and Language Effects in Cross-Lingual Speaker Verification for Iberian Languages

Pol Buitrago, Javier Hernando · Jul 1, 2026 · Citations: 0

Pairwise Preference Multilingual
  • However, standard evaluation protocols confound language mismatch with inter-speaker variability, as evaluation is generally performed with different speakers across languages.
  • In this work, we introduce a bilingual same-speaker evaluation set for five Iberian languages, enabling analysis of cross-lingual SV under constant speaker identity.
Efficient Multilingual Reasoning Transfer via Progressive Code-Switching

Zhijun Wang, Junxiao Liu, Hao Zhou, Hao-Ran Wei, Baosong Yang, Shujian Huang · Jul 1, 2026 · Citations: 0

Llm As JudgeAutomatic Metrics CodingMultilingual
  • However, existing transfer approaches typically rely on distilled target-language reasoning traces from stronger LRMs or online supervision from external judge models, which are costly and difficult to scale.
  • Experiments on multiple benchmarks and five typologically diverse languages show that PCS substantially narrows the performance gap between target-language and English reasoning, yielding more language-consistent reasoning while maintaining…
AI translation of literary texts is "fine", but readers still prefer human translations

Yves Ferstler, Adam Podoxin, Ty Brassington, Roman Grundkiewicz, Maite Taboada, Marzena Karpinska · Jun 24, 2026 · Citations: 0

Pairwise Preference Human EvalLlm As Judge Multilingual
  • While the content may be rendered adequately, we do not know enough about how readers experience it in terms of immersiveness and literary effect, aspects poorly captured by automatic machine translation metrics or human evaluation…
  • We ask 15 avid readers to compare recently published human translations (HT) to machine translations (MT) generated with an agentic large language model (LLM)-based pipeline, for 15 recent novels in French, Polish, and Japanese and…
A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models

Soham Dan, Himanshu Beniwal, Thomas Hartvigsen · Jun 24, 2026 · Citations: 0

Pairwise Preference Automatic Metrics CodingMultilingual
  • Large language models (LLMs) are increasingly deployed across languages, but their safety behavior remains uneven across linguistic and cultural contexts.
  • We first catalogue threat models that exploit language choice, translation pivots, code-switching, orthographic variation, multi-turn interaction, and post-deployment fine-tuning to weaken safety alignment.
REDACT: A Systematically Controlled Multilingual Benchmark for Personal Information Detection

Guneesh Vats, Anubha Agrawal, Shikha Singhal, Ajita Dash, Praison Selvaraj, Vidhan Jhawar · Jun 18, 2026 · Citations: 0

Llm As JudgeAutomatic Metrics CodingMultilingual
  • We present REDACT, a systematically controlled multilingual PII benchmark with 13,427 records, 324,078 entity annotations, 51 entity types, 4,127 surface-form patterns, and 25 languages across 9 scripts.
  • From the full benchmark, we evaluate five detectors (Presidio, GLiNER, the OpenAI Privacy Filter, GPT-4.1, and Claude Sonnet 4.6) on a locked, language-stratified sample of 1,000 records.
Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selection

Yutong Wang, Xuebo Liu, Derek F. Wong, Zhilin Li, Rongqing Jiang, Min Zhang · May 28, 2026 · Citations: 0

Pairwise Preference CodingMultilingual
  • To address this, we propose a human-like long document translation agent called Loong, which leverages a 3E memory module (Essence-Exemplar-Entity) to store summaries, sentence pairs, and entity records as historical context.
  • Loong optimizes its context policy through reinforcement learning, utilizing preference data derived from its own sampled observe-and-act reasoning trajectories.
Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization

Yilong Wang, Qianli Wang, Bohao Chu, Yihong Liu, Jing Yang, Simon Ostermann · May 12, 2026 · Citations: 0

Pairwise Preference Multilingual
  • We introduce Macro, a preference alignment framework that applies Direct Preference Optimization (DPO) to multilingual SCE generation, using a composite scoring function to construct preference pairs that effectively translate the trade-off…
  • Compared to supervised fine-tuning, Macro achieves superior performance on both metrics, confirming that explicit preference optimization is essential for balancing this trade-off.
Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages

Srija Anand, Ashwin Sankar, Ishvinder Sethi, Aaditya Pareek, Kartik Rajput, Gaurav Yadav · Apr 23, 2026 · Citations: 0

Pairwise Preference CodingMultilingual
  • We present a controlled multidimensional pairwise evaluation framework for multilingual TTS that combines linguistic control with perceptually grounded annotation.
  • Using 5K+ native and code-mixed sentences across 10 Indic languages, we evaluate 7 state-of-the-art TTS systems and collect over 120K pairwise comparisons from over 1900 native raters.
A Multi-Stage Validation Framework for Trustworthy Large-scale Clinical Information Extraction using Large Language Models

Maria Mahbub, Gregory M. Dams, Josh Arnold, Caitlin Rizy, Sudarshan Srinivasan, Elliot M. Fielstein · Apr 7, 2026 · Citations: 0

Expert Verification Automatic Metrics MedicineMultilingual
  • Conventional evaluation methods rely heavily on annotation-intensive reference standards or incomplete structured data, limiting feasibility at population scale.
  • Using judge-evaluated outputs as references, the primary LLM achieved an F1 score of 0.80 under relaxed matching criteria.
Plausibility as Commonsense Reasoning: Humans Succeed, Large Language Models Do not

Sercan Karakaş · Apr 6, 2026 · Citations: 0

Pairwise Preference Multilingual
  • Large language models achieve strong performance on many language tasks, yet it remains unclear whether they integrate world knowledge with syntactic structure in a human-like, structure-sensitive way during ambiguity resolution.
  • In a speeded forced-choice comprehension experiment, humans show a large, correctly directed plausibility effect.
Blinded Radiologist and LLM-Based Evaluation of LLM-Generated Japanese Translations of Chest CT Reports: Comparative Study

Yosuke Yamagishi, Atsushi Takamatsu, Yasunori Hamaguchi, Tomohiro Kikuchi, Shouhei Hanaoka, Takeharu Yoshikawa · Apr 2, 2026 · Citations: 0

Pairwise Preference Llm As JudgeAutomatic Metrics MedicineMultilingual
  • A board-certified radiologist and a radiology resident independently performed blinded pairwise evaluations across 4 criteria: terminology accuracy, readability, overall quality, and radiologist-style authenticity.
  • Radiologist 2 rated readability as equivalent in 75% of cases and favored the human-edited translation for overall quality (40% vs 21%).
Voxtral TTS

Mistral-AI, :, Alexander H. Liu, Alexis Tacnet, Andy Ehrenberg, Andy Lo · Mar 26, 2026 · Citations: 0

Human EvalAutomatic Metrics Multilingual
  • In human evaluations conducted by native speakers, Voxtral TTS is preferred for multilingual voice cloning due to its naturalness and expressivity, achieving a 68.4\% win rate over ElevenLabs Flash v2.5.