Skip to main navigation Skip to search Skip to main content

Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model

  • Jizhen LI
  • , Weiping TU*
  • , Yuhong YANG
  • , Xinmeng XU
  • , Yiqun ZHANG
  • , Yanzhen REN
  • *Corresponding author for this work

Research output: Book Chapters | Papers in Conference ProceedingsConference paper (refereed)Researchpeer-review

Abstract

Recently, the state space model (SSM) represented by Mamba has shown remarkable performance in long-term sequence modeling tasks, including speech enhancement. However, due to substantial differences in sub-band features, applying the same SSM to all sub-bands limits its inference capability. Additionally, when processing each time frame of the time-frequency representation, the SSM may forget certain high-frequency information of low energy, making the restoration of structure in the high-frequency bands challenging. For this reason, we propose Cross- and Sub-band Mamba (CSMamba). To assist the SSM in handling different sub-band features flexibly, we propose a band split block that splits the full-band into four sub-bands with different widths based on their information similarity. We then allocate independent weights to each subband, thereby reducing the inference burden on the SSM. Furthermore, to mitigate the forgetting of low-energy information in the high-frequency bands by the SSM, we introduce a spectrum restoration block that enhances the representation of the cross-band features from multiple perspectives. Experimental results on the DNS Challenge 2021 dataset demonstrate that CSMamba outperforms several state-of-the-art (SOTA) speech enhancement methods in three objective evaluation metrics with fewer parameters.
Original languageEnglish
Title of host publication2025 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025: Proceedings
EditorsBhaskar D RAO, Isabel TRANCOSO, Gaurav SHARMA, Neelesh B. MEHTA
PublisherIEEE
Number of pages5
ISBN (Electronic)9798350368741
DOIs
Publication statusPublished - 2025
Externally publishedYes
Event2025 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025 - Hyderabad, India
Duration: 6 Apr 202511 Apr 2025

Conference

Conference2025 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025
Country/TerritoryIndia
CityHyderabad
Period6/04/2511/04/25

Bibliographical note

Publisher Copyright:
© 2025 IEEE.

Funding

This work is supported by the National Nature Science Foundation of China (No. 62471343, No. 62071342, No.62171326), the Special Fund of Hubei Luojia Laboratory (No. 220100019).

Keywords

  • band split
  • spectrum restoration
  • speech enhancement
  • state space model

Fingerprint

Dive into the research topics of 'Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model'. Together they form a unique fingerprint.

Cite this