Abstract
The aim of stereo image super-resolution is to improve resolution of stereo images by using complementary information obtained from binocular systems. Recently, Transformer has been widely applied in stereo image super-resolution to address the limitation of convolutional networks that can not capture long-range information. However, Transformer also brings problem of high computational complexity. To better balance reconstruction performance and computational efficiency, we introduced and developed Mamba into stereo image super-resolution firstly. In order to efficiently establish long-range dependencies within/between views, a Stereo Interactive Mamba Network (SIMNet) is proposed to reconstruct high-quality stereo image pairs based on proposed residual visual state space block (RVSSB) and stereo Mamba fusion block (SMFB). Specifically, our SIMNet performs efficient feature extraction by developed RVSSB and fuses stereo features at different levels via a multi-level residual mechanism. In a meanwhile, a novel SMFB is proposed to explore complementary information between left view and right view. It designs a spatial channel interaction attention module (SCIAM) to capture and interact spatial features and channel features of the left view and right view, and then designs a stereo interaction mamba module (SIMM) to fuse feature information extracted by developed RVSSB. Extensive experiments illustrate that our SIMNet can achieve significant improvements in quantitative and qualitative evaluations with fewer parameters while maintaining global receptive field and linear complexity.