Now showing 1 - 7 of 7
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Preface
    (2020-01-01)
    Yang, Haiqin
    ;
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Preface
    (2020-01-01)
    Yang, Haiqin
    ;
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Preface
    (2020-01-01)
    Yang, Haiqin
    ;
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Preface
    (2024-01-01) ;
    Yang, Haiqin
    ;
    Mahmud, Mufti
    ;
    Cho, Sung Bae
    ;
    Pan, Min
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Preface
    (2020-01-01)
    Yang, Haiqin
    ;
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Preface
    (2020-01-01)
    Yang, Haiqin
    ;
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
    (2025-01-01)
    Singh, Punit Kumar
    ;
    Kumar, Nishant
    ;
    Ghosh, Akash
    ;
    Pasad, Kunal
    ;
    Soni, Khushi
    Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional and indigenous sporting traditions. To address this gap, we introduce CultSportQA, a benchmark designed to assess LMs' understanding of traditional sports across 60 countries and 6 continents, encompassing four distinct cultural categories. The dataset features 33,000 multiple-choice questions (MCQs) across text and image modalities, each of which is categorized into three key types: history-based, rule-based, and scenario-based. To evaluate model performance, we employ zero-shot, few-shot, and chain-of-thought (CoT) prompting across a diverse set of Large Language Models (LLMs), Small Language Models (SLMs), and Multimodal Large Language Models (MLMs). By providing a comprehensive multilingual and multicultural sports benchmark, CultSportQA establishes a new standard for assessing AI's ability to understand and reason about traditional sports.