Abstract:
With the advancement of computing power and energy efficiency of AI chips, the system-level bottleneck has shifted from individual chips to the interconnection and system coordination issues across chips, nodes, and domains. This article classifies computing power interconnections into Scale-Up, Scale-Out, and Wide-Area Interconnection based on their scope of operation, summarizes the implementation paths of various solutions, and identifies that interconnection hardware, congestion control, load balancing, communication libraries, scheduling and operations are the core challenges for large-scale deployment. Based on this, a roadmap centered on hardware-software co-design is proposed: advancing switch-RNIC collaborative telemetry and control, establishing a closed loop of topology and link between communication libraries and schedulers, balancing performance and maintainability through standardization and incremental deployment, and enhancing operator libraries and toolchains to narrow the ecosystem gap. Finally, future research directions are discussed, including tighter end-network collaboration, improved maintainability of interconnection solutions, and industry-level standardization.