COMPARING THE PERFORMANCE OF QUESTION ANSWERING BY LLMS USING QUANTIZATION AND RETRIEVAL AUGMENTED GENERATION TECHNIQUES

dc.contributor.authorNetisopakul, Ponrudee
dc.contributor.authorHaruethaipree, Sira
dc.date.accessioned2026-08-06T10:50:36Z
dc.date.available2026-08-06T10:50:36Z
dc.date.issued2025-03-01
dc.description.abstractThe development of large language models (LLMs) like ChatGPT and Google Bard has led to the creation of intelligent chatbots and question-answering systems that are gaining widespread popularity. However, there are still limitations in using LLMs to develop applications, including the substantial computational resources required for fine-tuning and deployment. This paper studies and experiments with two techniques to reduce the computing resources required for developing a question-answering system using LLMs. A quantization technique is employed to compress the model’s size, and the application of Retrieval Augmented Generation (RAG) techniques is utilized for information retrieval. The study compares the performance of compressed-size models using Quantization and RAG against the original-sized models. The results show that quantizing the model can compress the VRAM resources used in the GPU between 38% to 57% while still achieving 68.9% accuracies compared to 70% in the non-compress model.
dc.identifier.citationIcic Express Letters, 19(3), 261-269, 2025
dc.identifier.doi10.24507/icicel.19.03.261
dc.identifier.issn1881803X
dc.identifier.other2-s2.0-85216980667
dc.identifier.urihttps://dspace.kmitl.ac.th/handle/123456789/16835
dc.sourceIcic Express Letters
dc.subjectDocument Question Answering (DQA)
dc.subjectLarge language models
dc.subjectQuantization
dc.subjectResource-constrained machine
dc.subjectRetrieval Augmented Generation (RAG)
dc.titleCOMPARING THE PERFORMANCE OF QUESTION ANSWERING BY LLMS USING QUANTIZATION AND RETRIEVAL AUGMENTED GENERATION TECHNIQUES
dc.typeArticle

Files

Collections