Skip to main navigation Skip to search Skip to main content

A Mathematical Explanation of Transformers

  • Xue-Cheng TAI
  • , Hao LIU*
  • , Lingfeng LI
  • , Raymond H. CHAN
  • *Corresponding author for this work

Research output: Journal PublicationsJournal Article (refereed)peer-review

Abstract

The Transformer architecture has revolutionized the field of sequence modeling and underpins the recent breakthroughs in large language models (LLMs). However, a comprehensive mathematical theory that explains its structure and operations remains elusive. In this work, we propose a novel continuous framework that rigorously interprets the Transformer as a discretization of a structured integro-differential equation. Within this formulation, the self-attention mechanism emerges naturally as a nonlocal integral operator, and layer normalization is characterized as a projection to a time-dependent constraint. This operator-theoretic and variational perspective offers a unified and interpretable foundation for understanding the architecture’s core components, including attention, feedforward layers, and normalization. Our approach extends beyond previous theoretical analyses by embedding the entire Transformer operation in continuous domains for both token indices and feature dimensions. This leads to a principled and flexible framework that not only deepens on theoretical insight but also offers new directions for architecture design, analysis, and control-based interpretations. This new interpretation provides a step toward bridging the gap between deep learning architectures and continuous mathematical modeling, and contributes a foundational perspective to the ongoing development of interpretable and theoretically grounded neural network models.
Original languageEnglish
Pages (from-to)1542-1568
Number of pages27
JournalSIAM Journal on Imaging Sciences
Volume19
Issue number3
Early online date17 Jul 2026
DOIs
Publication statusE-pub ahead of print - 17 Jul 2026

Bibliographical note

Publisher Copyright:
© by SIAM and 2026 Society for Industrial and Applied Mathematics.

Funding

The work of the first author was partially supported by NORCE Kompetanseoppbygging program. The work of the second author was partially supported by HKRGC ECS 22302123, HKRGC GRF 12301925, and Guangdong and Hong Kong Universities “1+1+1” Joint Research Collaboration Scheme 2025A0505000007. The work of the third author was supported by the start-up fund of Hetao Institute of Mathematics and Interdisciplinary Sciences (Shenzhen) and HKRGC GRF Grant LU13300125. The work of the fourth author was partially supported by HKRGC grants LU13300125 and LU11309922, ITF grant MHP/054/22, and LUBGR105824.

Keywords

  • attention
  • integro-differential equation
  • operator splitting
  • transformer

Fingerprint

Dive into the research topics of 'A Mathematical Explanation of Transformers'. Together they form a unique fingerprint.

Cite this