高级检索

人文智能高质量数据集的建设与价值释放

High-Quality Datasets in AI for Humanities: Construction and Value Creation

  • 摘要: 让沉睡的文化资源转化为可计算、可关联的高质量数据集,是文化与科技深度融合中的一项重要课题。本文以首届CCF人文智能大会汇聚的一批典型人文智能数据集为对象,沿建设与价值释放2条线索展开观察。研究指出,人文数据在内涵、语境、关联、敏感性与评判等方面区别于通用数据,其高质量首先体现为标注的专业质量,并有赖于关联的保全与先进计算技术的支撑;在价值层面,人文智能数据集既支撑前沿人文研究,也赋能文化产业与社会,推动文化资源向文化新质生产力转化。文章进而讨论了标注专业性与规模化、通用方法适用性、数据与算力储备、跨学科协作等方面的共性挑战,以期为人文智能数据集的建设与应用提供来自一线的参照。

     

    Abstract: Transforming dormant cultural resources into high-quality, computable, and interconnected datasets is a key task in advancing the integration of culture and technology. Drawing on a collection of representative humanities-AI datasets presented at the 1st CCF AI for Humanities Conference, this article examines the field from two perspectives: dataset construction and value creation. Humanities data differ from general-purpose data in semantic richness, contextual dependence, relational structure, cultural sensitivity, and evaluation criteria. Their quality depends fundamentally on expert-driven annotation, supported by the preservation of relational structure and by advanced computational techniques. These datasets support both cutting-edge humanities research and applications in cultural industries and broader society, transforming cultural resources into new quality cultural productive forces. The article further discusses shared challenges concerning annotation expertise and scalability, the applicability of general-purpose methods, the availability of data and computing resources, interdisciplinary collaboration, and the governance and long-term stewardship, providing practical insights for the construction and application of humanities-AI datasets.

     

/

返回文章
返回