Abstract
Driving fatigue detection is a critical component of intelligent driving and road safety, playing a significant role in accident prevention. Methods based on facial features hold an important position in this field due to their non-contact nature, ease of implementation, and low cost. However, existing approaches typically rely on large amounts of annotated data and use coarse category labels (such as “alert” and “drowsy”), which struggle to fully capture the rich semantic information in facial behaviors. Moreover, most of them analyze static features from single frames (e.g., eye/mouth aspect ratios) and fail to model the temporal dynamics of fatigue behaviors (such as frequent eye-closing and intermittent yawning), resulting in limited feature representation and generalization. To address these issues, we propose DFD-CLIP, a novel unified vision-language framework based on CLIP for driving fatigue detection, which supports both static image and video inputs while demonstrating strong cross-dataset generalization. Specifically, we restructure CLIP into a temporally enhanced framework by integrating a temporal Transformer module and leveraging learnable class tokens to extract video-level fatigue features. We utilize an LLM to generate fine-grained facial behavior descriptions instead of traditional labels and combine them with adaptive prompt learning to enhance semantic representation. Moreover, a composite loss function incorporating a dynamic boundary loss and a fine-grained weighting loss is designed to mitigate classification instability caused by label ambiguity and class imbalance. Experimental results demonstrate that DFD-CLIP achieves state-of-the-art performance on public datasets.
| Original language | English |
|---|---|
| Number of pages | 14 |
| Journal | IEEE Transactions on Multimedia |
| DOIs | |
| Publication status | E-pub ahead of print - 11 Jun 2026 |
Bibliographical note
Publisher Copyright:© 1999-2012 IEEE.
Keywords
- Driving fatigue detection
- vision-language model
- temporal modeling
- learnable prompt
Fingerprint
Dive into the research topics of 'DFD-CLIP: A Unified Vision-Language Framework for Driving Fatigue Detection with Temporal Modeling and Adaptive Prompt Learning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver