Challenges, Explorations, and New Standards of Data Construction in Software Engineering
-
Abstract
Data is the core resource for driving technological innovation and validation in modern software engineering(SWE), and it also underpins the advancement of large language models (LLMs) capabilities. However, systematically constructing high-quality evaluation data and training data remains one of the central challenges at the intersection of SWE and artificial intelligence(AI). From a practical perspective, this article examines three major contradictions in current SWE data construction: the quality paradox, scenario fragmentation, and benchmark aging. Taking the SWE-series benchmarks as case studies, it conducts an in-depth analysis of the technical bottlenecks and automation breakthroughs in environment construction, explores the phenomena and mechanisms of hallucination contamination in synthetic data, and offers methodological reflections on bridging the gap between benchmarks and real-world business scenarios. On this basis, the article further argues that, in the era of large models, it is essential to establish new standards for data construction centered on traceability, systematic validation, and alignment with human preferences, so as to address the quality crisis brought by the large-scale production of synthetic data.
-
-