高级检索

阳光清言:依托强基座模型的藏语通用大模型构建

Yangguang Qingyan: Construction of a Tibetan General-purpose Large Language Model Based on a Strong Foundation Model

  • 摘要: “阳光清言”是依托智谱AI的强基座模型GLM-4.5研发的藏语通用大模型。文章介绍了“阳光清言”大模型研发过程中,针对大规模高质量藏语数据资源积累不足、基座模型藏语适配能力有限、任务能力转化不充分和场景应用验证不足等关键技术问题,围绕数据体系构建、基座适配、继续预训练、通用指令微调、评测迭代和应用示范等关键环节形成的较为完整的技术路径。首先,“阳光清言”构建了藏语单语语料、篇章级语料、汉藏平行语料、藏英平行语料、汉藏双语词典资源、任务型指令数据和藏语评测数据集等一系列数据资源,再通过“强基座模型+藏语继续预训练+通用指令微调+评测迭代”的路径分阶段进行技术递进,最后逐步使模型形成藏语理解、文本生成、知识问答、多轮对话和机器翻译等能力。应用实测表明,“阳光清言”在藏语知识问答、释义生成、多轮对话以及汉藏、藏英互译等典型场景中已具备基础可用性。总体来看,“阳光清言”将资源建设、能力注入、任务对齐、评测治理与应用验证纳入统一框架,形成了低资源语言大模型从数据建设、模型训练到场景落地的完整研发链条,为其他低资源语言大模型研发提供了可借鉴的技术范式。

     

    Abstract: “Yangguang Qingyan” (SunshineGLM-V1.0) is a Tibetan general-purpose large language model developed based on GLM-4.5, a strong foundation model developed by Zhipu AI. This article presents a relatively complete technical pathway formed during the development of “Yangguang Qingyan” to address several key technical challenges, including the insufficient accumulation of large-scale high-quality Tibetan data resources, the limited Tibetan adaptation capability of the foundation model, inadequate transformation of task capabilities, and insufficient validation in application scenarios. The pathway covers key stages such as data system construction, foundation model adaptation, continual pretraining, general instruction tuning, evaluation-driven iteration, and application demonstration. Specifically, “Yangguang Qingyan” first constructs a series of data resources, including Tibetan monolingual corpora, document-level corpora, Chinese-Tibetan parallel corpora, Tibetan-English parallel corpora, Chinese-Tibetan bilingual lexicon resources, task-oriented instruction data, and Tibetan evaluation datasets. It then follows a staged and progressive technical route of “strong foundation model + Tibetan continual pretraining + general instruction tuning + evaluation-driven iteration,” through which the model gradually develops capabilities in Tibetan language understanding, text generation, knowledge question answering, multi-turn interaction, and machine translation. Application tests show that “Yangguang Qingyan” has achieved basic usability in typical scenarios such as Tibetan knowledge question answering, definition generation, multi-turn dialogue, and Chinese-Tibetan and Tibetan-English bidirectional translation. Overall, “Yangguang Qingyan” integrates resource construction, capability injection, task alignment, evaluation governance, and application validation into a unified framework, forming a complete research and development chain for low-resource language large language models from data construction and model training to scenario-based deployment. It provides a referential technical paradigm for the development of large language models for other low-resource languages.

     

/

返回文章
返回