Abstract
Visual deep learning (VDL) systems power safety-critical applications like autonomous driving and demand comprehensive testing, where generating diverse image mutations is crucial. Recently, multi-modal large language models (MLLMs) introduced instruction-driven methods, offering a unified paradigm of creating diverse and complex image mutations by describing them in text. Despite the promise, their mutation quality and applicability remain largely unexplored. We present the first study assessing MLLM's adequacy in VDL testing, focusing on the semantic validity of mutated images, their alignment with text instructions (prompts), and the faithfulness in preserving unchanged semantics. Through large-scale human studies and quantitative analysis, we identify MLLMs’ potential. Notably, while SoTA MLLMs (e.g., GPT-4V) struggle with editing already-appeared semantics in images (e.g., rotating objects), they enable novel “semantic-replacement” mutations (e.g., “dress a dog with clothes”) that were infeasible previously. These mutations bring new semantics, effectively triggering unique VDL faults. Thus, MLLM-based mutations are a vital complement to traditional methods, not a replacement. Building on our findings, we develop new testing metrics tailored to MLLM-enabled mutations and summarize lessons for employing and designing MLLM mutation pipelines in different testing scenarios. Our evaluation framework and generated dataset form a reusable benchmark for future research.
| Original language | English |
|---|---|
| Journal | ACM Transactions on Software Engineering and Methodology |
| DOIs | |
| Publication status | Published - 13 May 2026 |
Fingerprint
Dive into the research topics of 'How Multi-Modal LLMs Reshape Visual Deep Learning Testing? A Comprehensive Study Through the Lens of Image Mutation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver