Advanced Search
Tashi Nyima, Mabao Ban, Tsering Thupten, et al. Yangguang Qingyan: Construction of a Tibetan General-purpose Large Language Model Based on a Strong Foundation ModelJ. Computing Magazine of the CCF, 2026, 2(7): 32−41. DOI: 10.11991/cccf.202607007
Citation: Tashi Nyima, Mabao Ban, Tsering Thupten, et al. Yangguang Qingyan: Construction of a Tibetan General-purpose Large Language Model Based on a Strong Foundation ModelJ. Computing Magazine of the CCF, 2026, 2(7): 32−41. DOI: 10.11991/cccf.202607007

Yangguang Qingyan: Construction of a Tibetan General-purpose Large Language Model Based on a Strong Foundation Model

  • “Yangguang Qingyan” (SunshineGLM-V1.0) is a Tibetan general-purpose large language model developed based on GLM-4.5, a strong foundation model developed by Zhipu AI. This article presents a relatively complete technical pathway formed during the development of “Yangguang Qingyan” to address several key technical challenges, including the insufficient accumulation of large-scale high-quality Tibetan data resources, the limited Tibetan adaptation capability of the foundation model, inadequate transformation of task capabilities, and insufficient validation in application scenarios. The pathway covers key stages such as data system construction, foundation model adaptation, continual pretraining, general instruction tuning, evaluation-driven iteration, and application demonstration. Specifically, “Yangguang Qingyan” first constructs a series of data resources, including Tibetan monolingual corpora, document-level corpora, Chinese-Tibetan parallel corpora, Tibetan-English parallel corpora, Chinese-Tibetan bilingual lexicon resources, task-oriented instruction data, and Tibetan evaluation datasets. It then follows a staged and progressive technical route of “strong foundation model + Tibetan continual pretraining + general instruction tuning + evaluation-driven iteration,” through which the model gradually develops capabilities in Tibetan language understanding, text generation, knowledge question answering, multi-turn interaction, and machine translation. Application tests show that “Yangguang Qingyan” has achieved basic usability in typical scenarios such as Tibetan knowledge question answering, definition generation, multi-turn dialogue, and Chinese-Tibetan and Tibetan-English bidirectional translation. Overall, “Yangguang Qingyan” integrates resource construction, capability injection, task alignment, evaluation governance, and application validation into a unified framework, forming a complete research and development chain for low-resource language large language models from data construction and model training to scenario-based deployment. It provides a referential technical paradigm for the development of large language models for other low-resource languages.
  • loading

Catalog

    Turn off MathJax
    Article Contents

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return