Modernizing Convolutional Segmentation: A ConvNeXt-UNet Architecture for Multi-Modal Medical Image Analysis

avoin
Julkaisu on tekijänoikeussäännösten alainen. Teosta voi lukea ja tulostaa henkilökohtaista käyttöä varten. Käyttö kaupallisiin tarkoituksiin on kielletty.
Lataukset1

Verkkojulkaisu

DOI

Tiivistelmä

Manual tumour delineation in radiation oncology is time-intensive and prone to inter-observer variability. While deep learning offers a solution, the widely used U-Net is often limited by 3x3 kernels, which provide an insufficient effective receptive field for diffuse targets. This thesis evaluates an improved ConvNeXt-UNet architecture featuring 7x7 depthwise convolutions, an inverted bottleneck structure, and Layer Normalization to maximize spatial context. Two tasks were studied: Task A (Brain tumour segmentation, Magnetic Resonance Imaging [MRI]) and Task B (Head and neck cancer segmentation, Computed Tomography [CT]). Evaluation was conducted via patient-wise 5-fold cross-validation. The ConvNeXt-UNet significantly outperformed the U-Net in Task A, achieving a mean Dice Similarity Coefficient (DSC) of 0.9328 compared to 0.5496. However, in Task B, the U-Net proved superior (DSC: 0.3000 vs. 0.2558). These findings suggest that architectural advantages may depend on the imaging modality: while large kernels benefit MRI, they may introduce noise in lower-contrast CT scans. The results emphasize the necessity of multi-modal benchmarking to establish the generalizability of medical imaging architectures.

item.page.okmtext