Abstract As artificial intelligence (AI) models continue to expand rapidly and evolve at an accelerating pace, the sustainability of scaling law faces challenges as computational resources cannot expand indefinitely. A more promising strategy is to use high-performance interconnections to compose multiple modest-capacity computing chips into a high-efficiency system, enabling comparable capability to that of much larger resource-intensive systems. However, an effective on-chip communication and network that provides high bandwidth, low latency, and large-scale parallelism to accelerate distributed computation is still lacking. Here, we address this challenge by proposing an on-chip, all-optical-interconnect-based hybrid optoelectronic distributed computing system, achieving two orders of magnitude lower inference latency than a single graphics processing unit (GPU) while using only one-ninth of its computational resources. The 400-gigabits per second (Gbps) silicon photonic transceiver chips can provide high-speed and error-free optical input/output (I/O) for processing chips, and an ultra-low-loss (≤ 5 dB at 1300 nm), non-blocking 16×16 optical switch chip can further scale out the all-optical network for massive and flexible interconnection. As a validation, we configure the system to realize optical pipeline parallelism (OPP) and execute a five-layer convolutional neural network (CNN) denoising model, which processes a total of 1,000 images of 32,768 bits each in only 105.16 µs. These results convincingly illustrate a transformative pathway toward future high-efficiency computing architectures.
Tao et al. (Mon,) studied this question.