KMITL

Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1

Browse

Search Results

Now showing 1 - 1 of 1
  • Some of the metrics are blocked by your 
    Item type:Item,
    Benchmarking Next-Generation Frontier Models for Document-Level Arabic-Thai Medical Translation: A Reliability Study of LLM-as-a-Judge
    (2026-01-01)
    Lertsuksakda, Rathawut
    ;
    Netisopakul, Ponrudee
    ;
    Boonsang, Siridech
    This research targets Arabic-to-Thai medical translation, a low-resource and linguistically distant pair critical for public health. We utilize the OpenWHO dataset, a resource comprising 293 documents. This dataset is shielded from webcrawling to minimize training data contamination and ensure rigorous zero-shot evaluation. We conduct a comparative analysis between 2026-era frontier models including GPT-5.2, Gemini 3 Pro, and Claude Opus 4.5 and industry-standard NMT represented by Google Translate. Using an LLM-as-a-Judge framework with an AI jury of efficient reasoning models such as GPT-5.1, Gemini 3 Flash, and Claude Sonnet 4.5, we performed 3,516 evaluations of fidelity, fluency, and cultural appropriateness. Results demonstrate that frontier models outperform traditional NMT across all dimensions. While high exact agreement suggests LLM-as-a-Judge frameworks are promising for scalable evaluation, human validation reveals persistent opportunities for improving AI-human alignment, necessitating targeted oversight for safety-critical applications.