Key points are not available for this paper at this time.
Increasing demands for distributed machine learning (DML) have posed significant pressure on data-center networks (DCNs). This promotes the adoption of reconfigurable all-optical interconnects (AOI) in DCNs leveraging optical circuit switching (OCS) for better performance on throughput, energy efficiency, and data transfer latency. Despite their benefits, these OCS-based DCNs (ODCNs) still bear limited flexibility due to the larger switching granularity and longer reconfiguration latency of OCS. To address this issue, this work introduces in-network computing (INC) in an ODCN to realize P4INC-AOI, which can orchestrate INC and AOI to explore their mutual benefits for accelerating the training of DML jobs with less AOI reconfigurations. In the control plane of P4INC-AOI, we address the scheduling of concurrent DML jobs by formulating a mixed integer linear programming (MILP) model and proposing a time-efficient heuristic, to allocate multi-dimensional resources and configure AOI for minimizing the longest job completion time (JCT) across workloads. For the data plane, we extend existing in-network gradient aggregation schemes to accelerate DML jobs more efficiently. We first implement P4INC-AOI and verify its performance in a small-scale ODCN testbed, and further justify its effectiveness with large-scale simulations. Our experimental results demonstrate that compared with an ODCN without INC, P4INC-AOI not only cuts down AOI reconfigurations effectively but also reduces the average JCT of DML jobs in ResNet50 and VGG16 by 46. 66\% and 56. 34\%, respectively.
Xie et al. (Fri,) studied this question.