高级检索

大模型机理研究的方法论谱系

Continuous Methodological Spectrum for Mechanistic Research on Large Models

  • 摘要: 近年来,大模型研究迅速成为人工智能领域的中心议题。围绕这一技术体系,学术界与产业界大体沿着4个方向发力:能力更强、应用更广、效率更高、基础更牢。前2个方向直接推动模型能力进化和场景落地,第3个方向支撑大规模训练与推理的成本优化,而第4个方向——基础更牢——则关乎我们能否真正理解大模型、解释其行为、刻画其边界,并最终建立可信、可控、可持续发展的人工智能科学。本文认为,大模型机理研究不能被理解为单一方法的工作,而应被视为一条由“数学—物理学—生物学—心理学”构成的连续方法论谱系:数学方法论强调理想化假设下的定理证明与边界刻画;物理学方法论强调受控实验、统计规律与唯象定律;生物学方法论关注结构解剖、内部回路与功能定位;心理学方法论则通过黑盒测试与行为诱导理解模型的认知边界。4种路径既相互区分,又彼此联通,共同构成理解大模型的完整框架。本文进一步提出,大模型在数学上不是一个单一对象,而是函数逼近器、概率采样机、信息压缩器与通用图灵机的“四位一体”。因此,未来的大模型基础研究应超越单一视角,建立跨方法、跨尺度、跨层次的研究范式。

     

    Abstract: While the rapid advancement of Large Models has driven unprecedented successes in capabilities, efficiency, and real-world applications, a fundamental question remains: do we truly understand how they work? Establishing a solid theoretical foundation for these highly black-boxed systems is no longer just an academic curiosity, but a crucial prerequisite for developing safe, interpretable, and controllable AI. This article argues that mathematically, large models are multifaceted entities, acting simultaneously as function approximators, probability samplers, information compressors, and universal Turing machines. To fully capture this complexity, the study of AI mechanisms must transcend isolated viewpoints. The author proposes a continuous methodological spectrum that draws analogies from four fundamental scientific disciplines:•  Mathematics delineates the theoretical boundaries and absolute capacities of models under idealized assumptions through rigorous proofs.•  Physics extracts empirical laws (e.g., scaling laws) and statistical patterns via controlled experiments and synthetic data.•  Biology conducts mechanistic “anatomy” to decode internal representations, neural circuits, and functional modules within real-world models.•  Psychology explores cognitive boundaries and emergent behaviors through black-box testing, prompting, and input-output interactions. These four paradigms are not mutually exclusive but deeply interconnected and mutually reinforcing. By bridging abstract theoretical proofs with empirical behavioral observations, this continuous spectrum provides a holistic framework for AI mechanism research. Ultimately, this paradigm shift will transition the field of large models from empirical engineering into a rigorous, cumulative scientific discipline.

     

/

返回文章
返回