高级检索

软件工程数据构建的挑战、探索与新标准

Challenges, Explorations, and New Standards of Data Construction in Software Engineering

  • 摘要: 数据是驱动现代软件工程(software engineering, SWE)技术创新与验证的核心资源,也是提升大语言模型(LLMs)能力的基石。然而,如何系统地构建高质量的评测数据和训练数据,是当前SWE与人工智能(AI)交叉领域面临的核心挑战之一。本文从实践视角出发,梳理了当前SWE数据构建中的三大核心矛盾:质量悖论、场景割裂和基准老化,以SWE系列基准为案例深入分析了环境构建的技术瓶颈与自动化突破,探讨了合成数据中幻觉污染的现象与机制,并提出了弥合基准测试(benchmark)与业务场景鸿沟的方法论思考。在此基础上,本文进一步指出,大模型时代需要建立以可溯源性、系统化验证、人类偏好对齐为核心的数据构建新标准,以应对合成数据规模化带来的质量危机。

     

    Abstract: Data is the core resource for driving technological innovation and validation in modern software engineering(SWE), and it also underpins the advancement of large language models (LLMs) capabilities. However, systematically constructing high-quality evaluation data and training data remains one of the central challenges at the intersection of SWE and artificial intelligence(AI). From a practical perspective, this article examines three major contradictions in current SWE data construction: the quality paradox, scenario fragmentation, and benchmark aging. Taking the SWE-series benchmarks as case studies, it conducts an in-depth analysis of the technical bottlenecks and automation breakthroughs in environment construction, explores the phenomena and mechanisms of hallucination contamination in synthetic data, and offers methodological reflections on bridging the gap between benchmarks and real-world business scenarios. On this basis, the article further argues that, in the era of large models, it is essential to establish new standards for data construction centered on traceability, systematic validation, and alignment with human preferences, so as to address the quality crisis brought by the large-scale production of synthetic data.

     

/

返回文章
返回