Skip to main navigation Skip to search Skip to main content

DFD-CLIP: A Unified Vision-Language Framework for Driving Fatigue Detection with Temporal Modeling and Adaptive Prompt Learning

  • Jing BAI
  • , Xiwen FU
  • , Guopu ZHU
  • , Ligang WU
  • , Jinkun YOU
  • , Yang GAO
  • , Sam KWONG

Research output: Journal PublicationsJournal Article (refereed)peer-review

Abstract

Driving fatigue detection is a critical component of intelligent driving and road safety, playing a significant role in accident prevention. Methods based on facial features hold an important position in this field due to their non-contact nature, ease of implementation, and low cost. However, existing approaches typically rely on large amounts of annotated data and use coarse category labels (such as “alert” and “drowsy”), which struggle to fully capture the rich semantic information in facial behaviors. Moreover, most of them analyze static features from single frames (e.g., eye/mouth aspect ratios) and fail to model the temporal dynamics of fatigue behaviors (such as frequent eye-closing and intermittent yawning), resulting in limited feature representation and generalization. To address these issues, we propose DFD-CLIP, a novel unified vision-language framework based on CLIP for driving fatigue detection, which supports both static image and video inputs while demonstrating strong cross-dataset generalization. Specifically, we restructure CLIP into a temporally enhanced framework by integrating a temporal Transformer module and leveraging learnable class tokens to extract video-level fatigue features. We utilize an LLM to generate fine-grained facial behavior descriptions instead of traditional labels and combine them with adaptive prompt learning to enhance semantic representation. Moreover, a composite loss function incorporating a dynamic boundary loss and a fine-grained weighting loss is designed to mitigate classification instability caused by label ambiguity and class imbalance. Experimental results demonstrate that DFD-CLIP achieves state-of-the-art performance on public datasets.
Original languageEnglish
Number of pages14
JournalIEEE Transactions on Multimedia
DOIs
Publication statusE-pub ahead of print - 11 Jun 2026

Bibliographical note

Publisher Copyright:
© 1999-2012 IEEE.

Keywords

  • Driving fatigue detection
  • vision-language model
  • temporal modeling
  • learnable prompt

Fingerprint

Dive into the research topics of 'DFD-CLIP: A Unified Vision-Language Framework for Driving Fatigue Detection with Temporal Modeling and Adaptive Prompt Learning'. Together they form a unique fingerprint.

Cite this