Abstract:
Since the Transformer emerged as the foundational architecture for large language models (LLMs), their capability evolution has followed a remarkably consistent technical trajectory−systematically scaling a core dimension to drive performance gains. This principle was later formalized as scaling laws, which establish power-law relationships between model performance and parameter size, data volume, and computational resources. Over time, the scaling paradigm has expanded beyond the two traditional axes of model size and training data to include post-training reasoning enhancement and runtime environmental adaptation. The former encompasses outcome-supervised reinforcement learning and learning within interactive settings, while the latter spans inference-time scaling, integration of external tools and experiential knowledge, and multi-agent collaboration—reflecting, respectively, the trade-off between computation and reasoning fidelity, sustained domain-specific progress, and the shift from individual to collective intelligence. Nevertheless, the continued advancement of LLMs faces significant challenges, including diminishing returns from conventional pre-training scaling, the delicate balance between simplicity and openness in post-training environment design, and the limited understanding of internal working mechanisms. To move forward, future research needs to pursue innovations in model architectures, training efficiency, and environmental design, thereby steering LLMs capability growth toward a more principled and sustainable trajectory.