Abstract
Speech enhancement is an essential task of improving speech quality in noise scenario. Several state-of-the-art approaches have introduced visual information for speech enhancement, since the visual aspect of speech is essentially unaffected by acoustic environment. This paper proposes a novel framework that involves visual information for speech enhancement, by incorporating a Generative Adversarial Network (GAN). In particular, the proposed visual speech enhancement GAN consists of two networks trained in adversarial manner, i) a generator that adopts multi-layer feature fusion convolution network to enhance input noisy speech, and ii) a discriminator that attempts to minimize the discrepancy between the distributions of the clean speech signal and enhanced speech signal. Experiment results demonstrated superior performance of the proposed model against several state-of-the-art models.
| Original language | English |
|---|---|
| Title of host publication | 2022 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022: Proceedings |
| Publisher | IEEE |
| Pages | 7307-7311 |
| Number of pages | 5 |
| ISBN (Electronic) | 9781665405409 |
| DOIs | |
| Publication status | Published - 2022 |
| Externally published | Yes |
| Event | 47th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 - , Singapore Duration: 23 May 2022 → 27 May 2022 |
Publication series
| Name | ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings |
|---|---|
| Volume | 2022-May |
| ISSN (Print) | 1520-6149 |
Conference
| Conference | 47th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 |
|---|---|
| Country/Territory | Singapore |
| Period | 23/05/22 → 27/05/22 |
Bibliographical note
Publisher Copyright:© 2022 IEEE
Keywords
- generative adversarial network
- multi-layer feature fusion convolution network
- speech enhancement
- visual information
Fingerprint
Dive into the research topics of 'VSEGAN: Visual Speech Enhancement Generative Adversarial Network'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver