Skip to main navigation Skip to search Skip to main content

Multimodal Reasoning with LLM for Encrypted Traffic Interpretation: A Benchmark

  • Longgang ZHANG
  • , Chenxu ZHANG
  • , Xiaowei FU
  • , Fuxiang HUANG
  • , Lei ZHANG*
  • *Corresponding author for this work

Research output: Journal PublicationsJournal Article (refereed)peer-review

Abstract

Network traffic, as a key media format, is crucial for ensuring security and communications in modern internet infrastructure. While existing methods offer excellent performance, they face two key bottlenecks: (1) They fail to capture multidimensional semantics beyond unimodal sequence patterns. (2) Their “black box” property, i.e., providing only category labels, lacks an auditable reasoning process. We identify a key factor that existing network traffic datasets are primarily designed for classification and inherently lack rich semantic annotations, failing to generate human-readable evidence report. To address data scarcity, this paper proposes a Byte-Grounded Traffic Description (BGTD) benchmark for the first time, combining raw bytes with structured expert annotations. BGTD provides necessary behavioral features and verifiable chains of evidence for multimodal reasoning towards explainable encrypted traffic interpretation. Built upon BGTD, this paper proposes an end-to-end traffic-language representation framework (mmTraffic), a multimodal reasoning architecture bridging physical traffic encoding and semantic interpretation. In order to alleviate modality interference and generative hallucinations, mmTraffic adopts a jointly-optimized perception-cognition architecture. By incorporating a perception-centered traffic encoder and a cognition-centered LLM generator, mmTraffic achieves refined traffic interpretation with guaranteed category prediction. Extensive experiments demonstrate that mmTraffic generates high-fidelity, human-readable and evidence-grounded traffic interpretation reports, while maintaining highly competitive classification accuracy comparing to specialized unimodal model (e.g., NetMamba). The project is available at Traffic-Reasoning-Benchmark.
Original languageEnglish
Number of pages11
JournalIEEE Transactions on Multimedia
DOIs
Publication statusE-pub ahead of print - 2 Sept 2026

Funding

This work was partially supported by National Natural Science Fund of China under Grants 92570110 and 62271090, Chongqing Natural Science Fund under Grant CSTB2024NSCQ-JQX0038, and National Youth Talent Project.

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 9 - Industry, Innovation, and Infrastructure
    SDG 9 Industry, Innovation, and Infrastructure

Keywords

  • Encrypted traffic classification
  • network traffic interpretation
  • large language model
  • multimodal reasoning

Fingerprint

Dive into the research topics of 'Multimodal Reasoning with LLM for Encrypted Traffic Interpretation: A Benchmark'. Together they form a unique fingerprint.

Cite this