高级检索

电脑新传(1):大模型

New Legend of eBrain (1): Large Model

  • 摘要: 大模型经历了20多年的发明和优化历程。2000年约书亚·本吉奥率先采用词向量和自回归路线训练语言模型,2014年提出自注意机制,2017年底谷歌发明只需要注意机制的Transformer神经网络架构,2018年OpenAI训练出大语言模型GPT,实现了智能涌现,引爆智能革命。北京智源人工智能研究院2021年首次提出“大模型”概念,2025年训练出多模态大模型EMU,表明GPT自回归路线同样适用于视觉等模态,有望成为生成式人工智能的统一路线。

     

    Abstract: Large model has undergone more than 20 years of invention and optimization. In 2000, Yoshua Bengio pioneered the use of word embeddings and the autoregressive approach for training language models, and proposed self-attention mechanism in 2014. In late 2017, Google invented the Transformer neural network architecture, which relies solely on the attention mechanism. In 2018, OpenAI trained the large language model GPT, which realized intelligence emergence and ignited the intelligence revolution. In 2021, the Beijing Academy of Artificial Intelligence (BAAI) first put forward the concept of “large model” and in 2025, developed the multimodal large model EMU, demonstrating that the autoregressive paradigm practiced by GPT is also applicable to vision and other modalities, and is expected to become a unified path toward generative artificial intelligence.

     

/

返回文章
返回