Skip to main navigation Skip to search Skip to main content

Dynamic Weighted Combiner for Mixed-Modal Image Retrieval

  • Fuxiang HUANG
  • , Lei ZHANG*
  • , Xiaowei FU
  • , Suqi SONG
  • *Corresponding author for this work

Research output: Book Chapters | Papers in Conference ProceedingsConference paper (refereed)Referred Conference Paperpeer-review

Abstract

Mixed-Modal Image Retrieval (MMIR) as a flexible search paradigm has attracted wide attention. However, previous approaches always achieve limited performance, due to two critical factors are seriously overlooked. 1) The contribution of image and text modalities is different, but incorrectly treated equally. 2) There exist inherent labeling noises in describing users' intentions with text in web datasets from diverse real-world scenarios, giving rise to overfitting. We propose a Dynamic Weighted Combiner (DWC) to tackle the above challenges, which includes three merits. First, we propose an Editable Modality De-equalizer (EMD) by taking into account the contribution disparity between modalities, containing two modality feature editors and an adaptive weighted combiner. Second, to alleviate labeling noises and data bias, we propose a dynamic soft-similarity label generator (SSG) to implicitly improve noisy supervision. Finally, to bridge modality gaps and facilitate similarity learning, we propose a CLIP-based mutual enhancement module alternately trained by a mixed-modality contrastive loss. Extensive experiments verify that our proposed model significantly outperforms state-of-the-art methods on real-world datasets. The source code is available at https://github.com/fuxianghuang1/DWC.
Original languageEnglish
Title of host publicationProceedings of the 38th AAAI Conference on Artificial Intelligence
EditorsMichael Wooldridge, Jennifer Dy, Sriraam Natarajan
PublisherAssociation for the Advancement of Artificial Intelligence
Pages2303-2311
Number of pages9
Edition3
ISBN (Print)9781577358879
DOIs
Publication statusPublished - 25 Mar 2024
Externally publishedYes
EventThe 38th Annual AAAI Conference on Artificial Intelligence - Vancouver, Canada
Duration: 20 Feb 202427 Feb 2024

Publication series

NameProceedings of the AAAI Conference on Artificial Intelligence
PublisherAssociation for the Advancement of Artificial Intelligence
Number3
Volume38
ISSN (Print)2159-5399
ISSN (Electronic)2374-3468

Conference

ConferenceThe 38th Annual AAAI Conference on Artificial Intelligence
Abbreviated titleAAAI-24
Country/TerritoryCanada
CityVancouver
Period20/02/2427/02/24

Bibliographical note

Publisher Copyright:
Copyright © 2024, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved.

Funding

This work was partially supported by National Key R&D Program of China (2021YFB3100800), National Natural Science Fund of China (62271090, 61771079), Chongqing Natural Science Fund (cstc2021jcyj-jqX0023) and National Youth Talent Project. This work is also supported by Huawei computational power of Chongqing Artificial Intelligence Innovation Center.

Keywords

  • CV: Image and Video Retrieval
  • ML: Multimodal Learning
  • ML: Representation Learning

Fingerprint

Dive into the research topics of 'Dynamic Weighted Combiner for Mixed-Modal Image Retrieval'. Together they form a unique fingerprint.

Cite this