Improving OpenAI's Whisper Model for Transcribing Homophones in Legal News

dc.contributor.authorSiriket, Lattapon
dc.contributor.authorJitkajornwanich, Kulsawasd
dc.contributor.authorJaiyen, Saichon
dc.contributor.authorIntakosum, Sarun
dc.date.accessioned2026-08-06T10:44:24Z
dc.date.available2026-08-06T10:44:24Z
dc.date.issued2024-01-01
dc.description.abstractThe 'Whisper' model provides a tool for those who require transcription of human voice. It equips with opensource features and diverse functionalities. The model is capable of effectively deciphering messages in multiple languages, including support for the Thai language. This paper focuses on improving the transcription process of Thai homophones using the Whisper model in reducing the word error rate (WER). We focus on words in the legal news category and identify factors that lead to Whisper's incorrect sound predictions. We examined homophones using snippets of legal news video clips and compiled them into a homophone dictionary. We compare words extracted from the Whisper model by determining the word error rate and spelling of words. Based on the initial results obtained from the original Whisper model and the created homophone dictionary, 48 % of the words were incorrectly transcribed out of a total of 94 words. Then, we propose a methodology by which the performance of the Whisper is improved. That way, the automatic speech recognition of Thai language using the Whisper model can fully be utilized and used in other applications.
dc.identifier.citation2024 10th International Conference on Engineering Applied Sciences and Technology Iceast 2024, 108-113, 2024
dc.identifier.doi10.1109/ICEAST61342.2024.10554018
dc.identifier.other2-s2.0-85197245952
dc.identifier.urihttps://dspace.kmitl.ac.th/handle/123456789/15183
dc.source2024 10th International Conference on Engineering Applied Sciences and Technology Iceast 2024
dc.subjectAutomatic Speech Recognition
dc.subjectNatural Language Processing
dc.subjectOpenAI
dc.subjectVoice to Text
dc.subjectWhisper Model
dc.titleImproving OpenAI's Whisper Model for Transcribing Homophones in Legal News
dc.typeConference Paper

Files

Collections