高级检索

大语言模型的信息、计算和统计边界

Informational, Computational, and Statistical Limits of Large Language Models

  • 摘要: 大语言模型已经在应用上取得了巨大的成果,但其能力边界、适用范围与失败机制仍缺乏充分的基础理论解释。系统刻画其理论边界对于理解大语言模型有效性、适用性,以及构建真正安全可控的人工智能模型具有重要意义。本文从计算复杂性理论、信息与编码理论、统计机器学习理论3个视角分析并总结当前语言大语言模型范式的理论限制:在计算层面,大语言模型是基于 Transformer 的计算电路,其表达能力受到电路复杂性约束;在信息层面,大语言模型是对世界语料进行编码的压缩器,其知识容量受到可压缩结构与模型复杂度约束;在统计层面,大语言模型是执行下一词元预测的高维分类模型,受到样本复杂度、泛化与校准权衡等约束。因此,大语言模型能力边界应被理解为计算可表达性、信息可压缩性与统计可学习性共同作用的结果。

     

    Abstract: Large language models have achieved substantial empirical success, but their capability boundaries, applicable scope, and failure mechanisms still lack sufficient theoretical explanation. This article reviews the theoretical limitations of the current large language models paradigm from three perspectives: computational complexity theory, information and coding theory, and statistical learning theory. From the computational perspective, Transformer-based language models can be viewed as computational circuits whose expressiveness is constrained by circuit depth, precision, context length, and generation steps. From the information-theoretic perspective, pretraining can be interpreted as learning a compressor of world corpora, whose capacity is limited by compressibility, Kolmogorov complexity, and scaling-law constraints. From the statistical perspective, next-token prediction forms a high-dimensional learning problem that involves tradeoffs among generalization, memorization, calibration, and hallucination. These analyses suggest that improving model capability should not rely only on scale, but should also consider architecture, data distribution, verification mechanisms, and system-level constraints.

     

/

返回文章
返回