高级检索

谁是编程语言家族的世界语?

Who is the Universal Language in the Programming Language Family?

  • 摘要: 本文从大语言模型训练的视角重新解读编程语言的彼此关联性,提出一种数据驱动的编程语言关系分析方法,将多种编程语言的特征实例化为程序片段,并借助大语言模型进行向量嵌入,通过聚类分析得到编程语言的远近亲疏关系。文章得到众多合理且新颖的发现,例如编程语言之间展示出清晰的家族关系,传统视角上的编程语言衍生在大模型新方法中得到印证。另外,我们发现Go语言是最兼具其他语言特征的编程语言,具有“世界语”的特点。基于这种编程语言“族谱”的发现,笔者设计了多个实验验证它的实际价值,包括中介语言翻译、课程学习预训练,以及面向近亲的迁移学习。在三类任务中,我们实验发现均能助力任务的表现。该研究启发我们优化现有的多语言预训练策略,同时探索新的AI友好的编程语言。

     

    Abstract: The rapid advancement of large code models has blurred the boundaries between traditional programming languages, enabling models to comprehend, generate, and seamlessly translate code across diverse language paradigms. This evolution raises a profound question: much like human languages, do programming languages possess “family” relationships based on structural and semantic proximity? Furthermore, is there a universal “lingua franca” among them? If we could map the genealogical spectrum of programming languages—akin to linguists charting the Indo-European language family—could we leverage this topology to optimize the training of large language models for greater intelligence and efficiency? Our recent research provides a definitive answer. For the first time, we employ a data-driven approach to automatically mine the underlying “family tree” of 19 mainstream programming languages. By integrating these genealogical relationships to guide LLM training, our proposed method demonstrates significant performance improvements across four critical downstream tasks: code summarization, code search, code generation, and code translation.

     

/

返回文章
返回