TransNorm: Transformer Provides a Strong Spatial Normalization Mechanism for a Deep Segmentation Model

Azad, Reza and Al-Antary, Mohammad T. and Heidari, Moein and Merhof, Dorit (2022) TransNorm: Transformer Provides a Strong Spatial Normalization Mechanism for a Deep Segmentation Model. IEEE ACCESS, 10. pp. 108205-108215. ISSN 2169-3536

Full text not available from this repository. (Request a copy)

Abstract

In the past few years, convolutional neural networks (CNNs), particularly U-Net, have been the prevailing technique in the medical image processing era. Specifically, the U-Net model, as well as its alternatives, have successfully managed to address a wide variety of medical image segmentation tasks. However, these architectures are intrinsically imperfect as they fail to exhibit long-range interactions and spatial dependencies leading to a severe performance drop in the segmentation of medical images with variable shapes and structures. Transformers, preliminary proposed for sequence-to-sequence prediction, have arisen as surrogate architectures to precisely model global information assisted by the self-attention mechanism. Despite being feasibly designed, utilizing a pure Transformer for image segmentation purposes can result in limited localization capacity stemming from inadequate low-level features. Thus, a line of research strives to design robust variants of Transformer-based U-Net. In this paper, we propose Trans-Norm, a novel deep segmentation framework which concomitantly consolidates a Transformer module into both encoder and skip-connections of the standard U-Net. We argue that the expedient design of skip-connections can be crucial for accurate segmentation as it can assist feature fusion between the expanding and contracting paths. In this respect, we derive a Spatial Normalization mechanism from the Transformer module to adaptively recalibrate the skip connection path. Extensive experiments across three typical tasks for medical image segmentation demonstrate the effectiveness of TransNorm. The codes and trained models are publicly available at github.

Item Type:	Article
Uncontrolled Keywords:	Transformers; Image segmentation; Semantics; Medical diagnostic imaging; Computer architecture; Decoding; Convolutional neural networks; Biomedical image processing; Transformer; semantic segmentation; attention; medical image analysis
Subjects:	000 Computer science, information & general works > 004 Computer science
Divisions:	Informatics and Data Science
Depositing User:	Dr. Gernot Deinzer
Date Deposited:	15 Feb 2024 11:40
Last Modified:	15 Feb 2024 11:40
URI:	https://pred.uni-regensburg.de/id/eprint/57610

Actions (login required)

View Item