KMITL

Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1

Browse

Search Results

Now showing 1 - 9 of 9
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Fin-Ally: Pioneering the Development of an Advanced, Commonsense-Embedded Conversational AI for Money Matters
    (2025-10-21)
    Das, Sarmistha
    ;
    Mathur, Priya
    ;
    Sharma, Ishani
    ;
    Saha, Sriparna
    ;
    Pasupa, Kitsuchart
    The exponential technological breakthrough of the FinTech industry has significantly enhanced user engagement through sophisticated advisory chatbots. However, large-scale fine-tuning of LLMs can occasionally yield unprofessional or flippant remarks, such as 'With that money, you're going to change the world,' which, though factually correct, can be contextually inappropriate and erode user trust. The scarcity of domain-specific datasets has led previous studies to focus on isolated components, such as reasoning-aware frameworks or the enhancement of human-like response generation. To address this research gap, we present Fin-Solution 2.O, an advanced solution that 1) introduces the multi-turn financial conversational dataset, Fin-Vault, and 2) incorporates a unified model, Fin-Ally, which integrates commonsense reasoning, politeness, and human-like conversational dynamics. Fin-Ally is powered by COMET-BART-embedded commonsense context and optimized with a Direct Preference Optimization (DPO) mechanism to generate human-aligned responses. The novel Fin-Vault dataset, consisting of 1,417 annotated multi-turn dialogues, enables Fin-Ally to extend beyond basic account management to provide personalized budgeting, real-time expense tracking, and automated financial planning. Our comprehensive results demonstrate that incorporating commonsense context enables language models to generate more refined, textually precise, and professionally grounded financial guidance, positioning this approach as a next-generation AI solution for the FinTech sector.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
    (2025-01-01)
    Singh, Punit Kumar
    ;
    Kumar, Nishant
    ;
    Ghosh, Akash
    ;
    Pasad, Kunal
    ;
    Soni, Khushi
    Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional and indigenous sporting traditions. To address this gap, we introduce CultSportQA, a benchmark designed to assess LMs' understanding of traditional sports across 60 countries and 6 continents, encompassing four distinct cultural categories. The dataset features 33,000 multiple-choice questions (MCQs) across text and image modalities, each of which is categorized into three key types: history-based, rule-based, and scenario-based. To evaluate model performance, we employ zero-shot, few-shot, and chain-of-thought (CoT) prompting across a diverse set of Large Language Models (LLMs), Small Language Models (SLMs), and Multimodal Large Language Models (MLMs). By providing a comprehensive multilingual and multicultural sports benchmark, CultSportQA establishes a new standard for assessing AI's ability to understand and reason about traditional sports.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    M3Retrieve: Benchmarking Multimodal Retrieval for Medicine
    (2025-01-01)
    Acharya, Arkadeep
    ;
    Ghosh, Akash
    ;
    Verma, Pradeepika
    ;
    Pasupa, Kitsuchart
    ;
    Saha, Sriparna
    With the increasing use of Retrieval-Augmented Generation (RAG), strong retrieval models have become more important than ever. In healthcare, multimodal retrieval models that combine information from both text and images offer major advantages for many downstream tasks such as question answering, cross-modal retrieval, and multimodal summarization, since medical data often includes both formats. However, there is currently no standard benchmark to evaluate how well these models perform in medical settings. To address this gap, we introduce M3Retrieve, a Multimodal Medical Retrieval Benchmark. M3Retrieve, spans 5 domains,16 medical fields, and 4 distinct tasks, with over 1.2 Million text documents and 164K multimodal queries, all collected under approved licenses. We evaluate leading multimodal retrieval models on this benchmark to explore the challenges specific to different medical specialities and to understand their impact on retrieval performance. By releasing M3Retrieve, we aim to enable systematic evaluation, foster model innovation, and accelerate research toward building more capable and reliable multimodal retrieval systems for medical applications. The dataset and the baselines code are available in this github page https://github.com/AkashGhosh/M3Retrieve.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    ToxVI: a Multimodal LLM-based Framework for Generating Intervention in Toxic Code-Mixed Videos
    (2024-10-21)
    Maity, Krishanu
    ;
    Poornash, A. S.
    ;
    Saha, Sriparna
    ;
    Pasupa, Kitsuchart
    While considerable research has delved into detecting toxic content in text-based data, the realm of video content, particularly in languages other than English, has received less attention. Prior studies have primarily focused on creating automated tools to identify online toxic speech but have often overlooked the crucial next steps of mitigating its impact and discouraging future use. We can discourage social media users from sharing such material by automatically generating interventions that explain why certain content is inappropriate. To bridge this research gap, we propose an innovative task: generating interventions for toxic videos in code-mixed languages which go beyond existing methods focusing on text and images to combat online toxicity. We are introducing a Toxic Code-Mixed Intervention Video benchmark dataset (ToxCMI), comprising 1697 code-mixed toxic video utterances sourced from YouTube. Each utterance in this dataset has been meticulously annotated for toxicity and severity, accompanied by interventions provided in Hindi-English code-mixed languages. We have developed an advanced multimodal framework ToxVI, specifically designed for the task of generating Toxic Video appropriate Interventions, leveraging Large Language Models (LLMs), which comprises three modules - Modality module, Cross-Modal Synchronization module and Generation module. Our experiments demonstrate that integrating multiple modalities from the videos significantly enhances the performance of the proposed task and outperforms all the baselines by a significant margin.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Breast cancer prognosis through the use of multi-modal classifiers: current state of the art and the way forward
    (2024-09-01)
    Mathur, Archana
    ;
    Arya, Nikhilanand
    ;
    Pasupa, Kitsuchart
    ;
    Saha, Sriparna
    ;
    Roy Dey, Sudeepa
    We present a survey of the current state-of-the-art in breast cancer detection and prognosis. We analyze the evolution of Artificial Intelligence-based approaches from using just uni-modal information to multi-modality for detection and how such paradigm shift facilitates the efficacy of detection, consistent with clinical observations. We conclude that interpretable AI-based predictions and ability to handle class imbalance should be considered priority.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    HANCaps: A Two-Channel Deep Learning Framework for Fake News Detection in Thai
    (2024-01-01)
    Maity, Krishanu
    ;
    Bhattacharya, Shaubhik
    ;
    Phosit, Salisa
    ;
    Kongsamlit, Sawarod
    ;
    Saha, Sriparna
    The rapid advancement of internet technology, widespread smartphone usage, and the rise of social media platforms have drastically transformed the global communication landscape. These developments have resulted in both positive and negative consequences. On the one hand, they have facilitated the dissemination of information, connecting individuals across vast distances and fostering diverse perspectives. On the other hand, the ease of access to online platforms has led to the proliferation of misinformation, often in the form of fake news. Detecting and combatting fake news has become crucial to mitigate its adverse effects on society. This paper presents an investigation into fake news detection in the Thai language. It addresses current limitations in this domain by proposing a novel two-channel deep learning model named HANCaps, which integrates BERT and FastText embeddings with a hierarchical attention network and capsule network. The HANCaps model utilizes the BERT language model as one channel input, while the other channel incorporates pre-trained FastText embeddings. The proposed model undergoes evaluation using a benchmark Thai fake news dataset, and extensive experimentation demonstrates that HANCaps outperforms state-of-the-art methods by up to 3.28% in terms of F1 score, showcasing its superior performance.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    HateThaiSent: Sentiment-Aided Hate Speech Detection in Thai Language
    (2024-01-01)
    Maity, Krishanu
    ;
    Poornash, A. S.
    ;
    Bhattacharya, Shaubhik
    ;
    Phosit, Salisa
    ;
    Kongsamlit, Sawarod
    Social media platforms are a double-edged sword: on the one hand, they enable the dissemination of information; but on the other hand, they also provide an avenue for spreading online abuse and harassment, such as hate speech. While significant research efforts are being devoted to detecting online hate speech in the English language, little attention has been paid to the Thai language. In this study, we created a benchmark dataset, called HateThaiSent, which labels each post with both hate speech and sentiment information. To detect hate speech, we created a multitask model that uses a dual-channel deep learning approach based on FastText and BERT embeddings, with an added capsule network. One channel utilizes pretrained FastText embeddings while the other uses embeddings from the BERT language model. We aimed to answer two research questions: (Q1) Does incorporating sentiment information improves the performance of hate speech detection (HD) in the Thai language? (Q2) What is the comparative effectiveness of two different approaches for sentiment-Aware HD in the Thai language: feature engineering versus multitasking? Our proposed approach outperformed other baselines and state-of-The-Art models on the HateThaiSent dataset, with overall accuracy/macro-F1 values of 89.67%/89.79%, and 80.92%/80.97% for hate speech and sentiment detection tasks, respectively. We concluded that multitasking is more effective than feature engineering in enhancing the performance of the main task (HD).
  • Some of the metrics are blocked by your 
    Item type:Publication,
    FastThaiCaps: A Transformer Based Capsule Network for Hate Speech Detection in Thai Language
    (2023-01-01)
    Maity, Krishanu
    ;
    Bhattacharya, Shaubhik
    ;
    Saha, Sriparna
    ;
    Janoai, Suwika
    ;
    Pasupa, Kitsuchart
    The advent of technology has led to people sharing their views openly like never before. Parallelly, cyberbullying and hate speech content have also increased as a side effect that is potentially hazardous to society. While plenty of research is going on to detect online hate speech in English, there is very little research on the Thai language. To investigate how noisy Thai posts can be handled effectively, in this work, we have developed a two-channel deep learning model FastThaiCaps based on BERT and FastText embedding along with a capsule network. The input to one channel is the BERT language model, and that to the other is the pre-trained FastText embedding. Our model has been evaluated on a benchmark Thai dataset categorized into four categories, i.e., peace speech, neutral speech, level-1 hate speech, and level-2 hate speech. Experiments show that FastThaiCaps outperforms state-of-the-art methods by up to 3.11% in terms F1 score.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Ex-ThaiHate: A Generative Multi-task Framework for Sentiment and Emotion Aware Hate Speech Detection with Explanation in Thai
    (2023-01-01)
    Maity, Krishanu
    ;
    Bhattacharya, Shaubhik
    ;
    Phosit, Salisa
    ;
    Kongsamlit, Sawarod
    ;
    Saha, Sriparna
    Social media platforms have both positive and negative impacts on users in diverse societies. One of the adverse effects of social media platforms is the usage of hate and offensive language, which not only fosters prejudice but also harms the vulnerable. Additionally, a person’s sentiment and emotional state heavily influence the intended content of any social media post. Despite extensive research being conducted to detect online hate speech in English, there is a lack of similar studies on low-resource languages such as Thai. The recent enactment of laws like the “right to explanations” in the General Data Protection Regulation has stimulated the development of interpretable models rather than solely focusing on performance. Motivated by this, we created the first benchmark hate speech corpus, called Ex-ThaiHate, in the Thai language. Each post is annotated with four labels, namely hate, sentiment, emotion, and rationales (explainability), which specify the phrases that are responsible for annotating the post as hate. In order to investigate the effect of sentiment and emotional information on detecting hate speech posts, we propose a unified generative framework called GenX, which redefines this multi-task problem as a text-to-text generation task to simultaneously solve four tasks: hate-speech identification, rationale detection, sentiment, and emotion detection. Our extensive experiments demonstrate that GenX significantly outperforms all baselines and state-of-the-art models, thereby highlighting its effectiveness in detecting hate speech and identifying the rationales in low-resource languages. The code and dataset are available at https://github.com/dsmlr/Ex-ThaiHate. Disclaimer: The article contains offensive text and profanity. This is due to the nature of the work and does not reflect any opinion or stance of the authors.