UTAR Institutional Repository

Bridging the divide in low-resource speech translation through semantic corpus alignment and accessible model fine-tuning

Liang, Xiao (2026) Bridging the divide in low-resource speech translation through semantic corpus alignment and accessible model fine-tuning. PhD thesis, UTAR.

[img] PDF
Download (8Mb)

    Abstract

    The rapid expansion of economic and cultural exchanges between China and Malaysia has created unprecedented demand for efficient cross-lingual communication technologies. However, real-time Chinese–Malay speech translation remains severely constrained by three fundamental challenges: extreme data scarcity, profound typological divergence between the languages, and computational accessibility barriers to state-of-the-art models. This dissertation presents a comprehensive framework that addresses these challenges through three synergistic innovations, achieving state-of-the-art quality while remaining trainable on consumer-grade hardware. First, we introduce GPT-Aligner, a novel semantic alignment technique that recasts corpus creation as a reasoning task for large language models. This approach dramatically improves alignment accuracy over traditional methods for this typologically distant pair, enabling the cost-effective construction of a high-quality parallel corpus where none previously existed. Second, we develop Progressive Cross-Layer Low-Rank Adaptation (P-CLRA), a principled parameter-efficient fine-tuning method that uses a quantitative benefit-to-cost ratio to identify optimal components for adaptation. This achieves significant quality improvements that exceed full fine-tuning while adding a negligible number of trainable parameters. Third, we propose Progressive Cross-Layer Freezing (P-CLF), an adaptive strategy for speech translation that uses a novel gradient-based sensitivity analysis to identify and freeze non-critical model layers. This technique not only makes fine-tuning feasible on consumer-grade hardware but also surpasses the performance of a full fine-tuning approach. The integrated pipeline achieves a +0.98 BLEU improvement over full fine-tuning (8.32 → 9.30), delivering high-quality, real-time speech-to-speech translation on consumer hardware. This research establishes a new framework for resource-constrained multilingual speech translation, providing both immediate practical solutions for Chinese–Malay communication and generalizable methodologies for addressing language technology inequality in other underserved linguistic communities.

    Item Type: Final Year Project / Dissertation / Thesis (PhD thesis)
    Subjects: L Education > LA History of education
    P Language and Literature > PE English
    T Technology > T Technology (General)
    T Technology > TD Environmental technology. Sanitary engineering
    Divisions: Institute of Postgraduate Studies & Research > Faculty of Information and Communication Technology (FICT) - Kampar Campus > Doctor of Philosophy (Computer Science)
    Depositing User: ML Main Library
    Date Deposited: 08 Aug 2026 00:28
    Last Modified: 08 Aug 2026 00:28
    URI: http://eprints.utar.edu.my/id/eprint/7836

    Actions (login required)

    View Item