Objective: Traditional Full Waveform Inversion (FWI) methods for Ultrasound Computed Tomography (UCT) are computationally expensive and can be sensitive to strong acoustic contrasts. In this work, we propose the Multi-Channel Transducer Network (CUCT-Net), a deep learning framework that directly maps received ultrasound signals to image-space outputs for quantized speed-of-sound (SoS) estimation and for direct tissue-level segmentation over both low- and high-contrast regions, enabling end-to-end recovery of both contrast-driven and anatomically meaningful structures from raw measurements. Method: CUCT-Net uses a multi-input encoder–decoder architecture that maps raw multi-static UCT measurements to quantized SoS (or tissue-class) maps without requiring an initial guess or iterative optimization. Parallel per-transducer encoders extract view-specific features that are fused and refined by a decoder, with Shift Units (SU) used to enhance fine-scale feature modeling under sparse sensing. Experiments are performed on k-Wave simulations using (i) Shepp–Logan-inspired disc phantoms with Original/Distorted/Mixed variants and (ii) DBB-derived anatomical brain phantoms, under clean and noisy measurement conditions. Results: The proposed network achieves accurate quantized SoS estimation and direct tissue-level segmentation across synthetic and anatomically derived phantom experiments. Strong robustness to noise is demonstrated through transfer learning. Compared with FWI, CUCT-Net significantly reduces computational cost while maintaining stable performance under reduced-sensor conditions for quantized SoS estimation and complex tissue heterogeneity for segmentation. Conclusions: CUCT-Net formulates UCT as a direct signal-to-image learning problem that supports both quantized SoS estimation and tissue-level segmentation. By learning an end-to-end mapping from raw ultrasound measurements to quantized SoS or tissue representations, the proposed framework bypasses iterative inversion and achieves efficient and robust performance under reduced-sensor and strong-contrast conditions. The multi-input architecture enables effective integration of information from multiple transducers, demonstrating the feasibility and potential of data-driven end-to-end quantized SoS estimation and tissue segmentation for UCT.
Gao et al. (Thu,) studied this question.