Skip to main navigation Skip to search Skip to main content

How Multi-Modal LLMs Reshape Visual Deep Learning Testing? A Comprehensive Study Through the Lens of Image Mutation

  • Liwen WANG
  • , Yuanyuan YUAN*
  • , Ao SUN
  • , Zongjie LI
  • , Pingchuan MA
  • , Daoyuan WU
  • , Shuai WANG
  • *Corresponding author for this work

Research output: Journal PublicationsJournal Article (refereed)peer-review

Abstract

Visual deep learning (VDL) systems power safety-critical applications like autonomous driving and demand comprehensive testing, where generating diverse image mutations is crucial. Recently, multi-modal large language models (MLLMs) introduced instruction-driven methods, offering a unified paradigm of creating diverse and complex image mutations by describing them in text. Despite the promise, their mutation quality and applicability remain largely unexplored. We present the first study assessing MLLM's adequacy in VDL testing, focusing on the semantic validity of mutated images, their alignment with text instructions (prompts), and the faithfulness in preserving unchanged semantics. Through large-scale human studies and quantitative analysis, we identify MLLMs’ potential. Notably, while SoTA MLLMs (e.g., GPT-4V) struggle with editing already-appeared semantics in images (e.g., rotating objects), they enable novel “semantic-replacement” mutations (e.g., “dress a dog with clothes”) that were infeasible previously. These mutations bring new semantics, effectively triggering unique VDL faults. Thus, MLLM-based mutations are a vital complement to traditional methods, not a replacement. Building on our findings, we develop new testing metrics tailored to MLLM-enabled mutations and summarize lessons for employing and designing MLLM mutation pipelines in different testing scenarios. Our evaluation framework and generated dataset form a reusable benchmark for future research.
Original languageEnglish
JournalACM Transactions on Software Engineering and Methodology
DOIs
Publication statusPublished - 13 May 2026

Fingerprint

Dive into the research topics of 'How Multi-Modal LLMs Reshape Visual Deep Learning Testing? A Comprehensive Study Through the Lens of Image Mutation'. Together they form a unique fingerprint.

Cite this