Skip to main navigation Skip to search Skip to main content

VSEGAN: Visual Speech Enhancement Generative Adversarial Network

  • Xinmeng XU
  • , Yang WANG
  • , Dongxiang XU
  • , Yiyuan PENG
  • , Cong ZHANG
  • , Jie JIA
  • , Binbin CHEN

Research output: Book Chapters | Papers in Conference ProceedingsConference paper (refereed)Researchpeer-review

Abstract

Speech enhancement is an essential task of improving speech quality in noise scenario. Several state-of-the-art approaches have introduced visual information for speech enhancement, since the visual aspect of speech is essentially unaffected by acoustic environment. This paper proposes a novel framework that involves visual information for speech enhancement, by incorporating a Generative Adversarial Network (GAN). In particular, the proposed visual speech enhancement GAN consists of two networks trained in adversarial manner, i) a generator that adopts multi-layer feature fusion convolution network to enhance input noisy speech, and ii) a discriminator that attempts to minimize the discrepancy between the distributions of the clean speech signal and enhanced speech signal. Experiment results demonstrated superior performance of the proposed model against several state-of-the-art models.
Original languageEnglish
Title of host publication2022 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022: Proceedings
PublisherIEEE
Pages7307-7311
Number of pages5
ISBN (Electronic)9781665405409
DOIs
Publication statusPublished - 2022
Externally publishedYes
Event47th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022 - , Singapore
Duration: 23 May 202227 May 2022

Publication series

NameICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
Volume2022-May
ISSN (Print)1520-6149

Conference

Conference47th IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2022
Country/TerritorySingapore
Period23/05/2227/05/22

Bibliographical note

Publisher Copyright:
© 2022 IEEE

Keywords

  • generative adversarial network
  • multi-layer feature fusion convolution network
  • speech enhancement
  • visual information

Fingerprint

Dive into the research topics of 'VSEGAN: Visual Speech Enhancement Generative Adversarial Network'. Together they form a unique fingerprint.

Cite this