Skip to main navigation Skip to search Skip to main content

Multimodal Medical Image Binding via Shared Text Embeddings

  • Yunhao LIU
  • , Suyang XI
  • , Shiqi LIU
  • , Hong DING
  • , Chicheng JIN
  • , Chong ZHONG
  • , Junjun HE
  • , Catherine C. LIU
  • , Yiqing SHEN

Research output: Book Chapters | Papers in Conference ProceedingsConference paper (refereed)Referred Conference Paperpeer-review

Abstract

Medical image analysis increasingly relies on the integration of multiple imaging modalities to capture complementary anatomical and functional information, enabling more accurate diagnosis and treatment planning. Achieving aligned feature representations across these diverse modalities is therefore important for effective multimodal analysis. While contrastive language-image pre-training (CLIP) and its variant have enabled image-text alignments, they require explicitly paired data between arbitrary two modalities, which is difficult to acquire in medical contexts. To address the gap, we present Multimodal Medical Image Binding with Text (M3 Bind), a novel pre-training framework that enables seamless alignment of multiple medical imaging modalities through a shared text representation space without requiring explicit paired data between any two medical image modalities. Specifically, based on the insight that different images can naturally bind with text, M3Bind first fine-tunes pre-trained CLIP-like image-text models, which are derived from different medical modalities, to align their modality-specific text embedding space while preserving their original image-text alignments. Subsequently, we distill these modality-specific text encoders into a unified model, creating a shared text embedding space. Notably, M3Bind is a flexible framework in which the selection of CLIP-like models is not fixed and can be adapted according to the requirements of the task. Experiments on X-ray, CT, retina, ECG, and pathological images on multiple downstream tasks demonstrate that M3 Bind achieves competitive or even superior performance in zero-shot, few-shot classification and cross-modal retrieval tasks compared to its CLIP-like counterparts. These results validate M3Bind's effectiveness in achieving cross-image-modal alignment for medical analysis.
Original languageEnglish
Title of host publication2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV): Proceedings
PublisherIEEE
Pages1610-1620
Number of pages11
ISBN (Electronic)9798331555115
DOIs
Publication statusPublished - Mar 2026
Externally publishedYes
Event2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026 - Tucson, United States
Duration: 6 Mar 202610 Mar 2026

Conference

Conference2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
Abbreviated titleWACV 2026
Country/TerritoryUnited States
CityTucson
Period6/03/2610/03/26

Bibliographical note

Publisher Copyright:
© 2026 IEEE.

Funding

This work was funded by the Innovation and Technology Commission of the Hong Kong SAR Government [grant number K-BBY1], and the Otto Poon Research Institute for Climate-Resilient Infrastructure (RICRI), The Hong Kong Polytechnic University.

Fingerprint

Dive into the research topics of 'Multimodal Medical Image Binding via Shared Text Embeddings'. Together they form a unique fingerprint.

Cite this