Key points are not available for this paper at this time.
Cone-beam computed tomography (CBCT) is widely used in image-guided radiotherapy because of its low radiation dose and on-board acquisition capability. However, CBCT images often suffer from scatter artifacts, increased noise, reduced soft-tissue contrast, and inaccurate Hounsfield Unit (HU) values, which limit their direct use for accurate dose calculation and quantitative analysis. To address this limitation, we propose a CBCT-to-CT synthesis framework based on 2.5D context encoding (concatenating five adjacent slices along the channel dimension) and latent-space variational diffusion. The proposed method combines a Vector Quantized Variational Autoencoder (VQ-VAE) and a U-shaped Vision Transformer (U-ViT)-based latent-space Variational Diffusion Model (VDM) to translate CBCT images into synthetic CT (sCT) images in a compressed latent space. To incorporate inter-slice anatomical context while preserving the computational efficiency of 2D processing, five adjacent CBCT slices are concatenated along the channel dimension and used as input. We evaluated the proposed method on the SynthRAD2025 paired CBCT-CT dataset covering head-and-neck, thoracic, and abdominal regions. Under the provided benchmark setting, quantitative evaluation on the validation set showed that the proposed 2.5D model improved peak signal-to-noise ratio (PSNR) from 25.39 dB to 27.44 dB (averaged across regions), structural similarity index measure (SSIM) from 0.813 to 0.846, reduced mean squared error (MSE) from 0.00313 to 0.00200, and lowered Fréchet inception distance (FID) from 1009.33 to 869.53 compared with the 2D baseline. Qualitative results also showed improved anatomical consistency and reduced artifact-related distortions. These findings suggest that neighboring-slice context can enhance HU fidelity and overall image quality in a computationally practical synthesis framework, supporting the usefulness of efficient AI-based cross-modality reconstruction for radiotherapy-related imaging workflows.
Park et al. (Fri,) studied this question.