Liang, Xiao (2026) Bridging the divide in low-resource speech translation through semantic corpus alignment and accessible model fine-tuning. PhD thesis, UTAR.
| PDF Download (8Mb) |
Abstract
The rapid expansion of economic and cultural exchanges between China and Malaysia has created unprecedented demand for efficient cross-lingual communication technologies. However, real-time Chinese–Malay speech translation remains severely constrained by three fundamental challenges: extreme data scarcity, profound typological divergence between the languages, and computational accessibility barriers to state-of-the-art models. This dissertation presents a comprehensive framework that addresses these challenges through three synergistic innovations, achieving state-of-the-art quality while remaining trainable on consumer-grade hardware. First, we introduce GPT-Aligner, a novel semantic alignment technique that recasts corpus creation as a reasoning task for large language models. This approach dramatically improves alignment accuracy over traditional methods for this typologically distant pair, enabling the cost-effective construction of a high-quality parallel corpus where none previously existed. Second, we develop Progressive Cross-Layer Low-Rank Adaptation (P-CLRA), a principled parameter-efficient fine-tuning method that uses a quantitative benefit-to-cost ratio to identify optimal components for adaptation. This achieves significant quality improvements that exceed full fine-tuning while adding a negligible number of trainable parameters. Third, we propose Progressive Cross-Layer Freezing (P-CLF), an adaptive strategy for speech translation that uses a novel gradient-based sensitivity analysis to identify and freeze non-critical model layers. This technique not only makes fine-tuning feasible on consumer-grade hardware but also surpasses the performance of a full fine-tuning approach. The integrated pipeline achieves a +0.98 BLEU improvement over full fine-tuning (8.32 → 9.30), delivering high-quality, real-time speech-to-speech translation on consumer hardware. This research establishes a new framework for resource-constrained multilingual speech translation, providing both immediate practical solutions for Chinese–Malay communication and generalizable methodologies for addressing language technology inequality in other underserved linguistic communities.
| Item Type: | Final Year Project / Dissertation / Thesis (PhD thesis) |
|---|---|
| Subjects: | L Education > LA History of education P Language and Literature > PE English T Technology > T Technology (General) T Technology > TD Environmental technology. Sanitary engineering |
| Divisions: | Institute of Postgraduate Studies & Research > Faculty of Information and Communication Technology (FICT) - Kampar Campus > Doctor of Philosophy (Computer Science) |
| Depositing User: | ML Main Library |
| Date Deposited: | 08 Aug 2026 00:28 |
| Last Modified: | 08 Aug 2026 00:28 |
| URI: | http://eprints.utar.edu.my/id/eprint/7836 |
Actions (login required)
| View Item |

