Projects per year
Abstract
Large Language Models (LLMs) have achieved remarkable generative capabilities but often underperform in sequence- and token-level classification tasks due to the causal masking constraint in decoder-only architectures. This unidirectional attention prevents tokens from accessing bidirectional context, limiting representation learning for discriminative prediction. We propose Label-Supervised Bi-directional Large Language Models (LS-BiLLMs), a lightweight adaptation method that (1) employs direct label supervision to align latent representations with task-specific labels and (2) removes the causal mask to enable bidirectional information flow. Implemented with LoRA-based fine-tuning, LS-BiLLMs efficiently adapt compact open-weight LLMs, such as LLaMA, Qwen, and Mistral, for classification without complex prompt engineering. Experiments across text classification, named-entity recognition, and commonsense reasoning benchmarks show consistent gains over instruction-tuned and encoder-based baselines. While unmasking sacrifices autoregressive generation, it substantially enhances discriminative understanding and efficiency. These findings reveal how causal directionality in attention mechanisms affects representational learning and reasoning in modern LLMs.
| Original language | English |
|---|---|
| Article number | 104568 |
| Journal | Information Processing and Management |
| Volume | 63 |
| Issue number | 4 |
| Early online date | 7 Jan 2026 |
| DOIs | |
| Publication status | Published - Jun 2026 |
Bibliographical note
Publisher Copyright:© 2026 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
Funding
Zongxi Li and Haoran Xie have been supported by Lingnan University through Faculty Research Grants (SDS24A2, SDS24A8, SDS24A12, SDS24A19), Direct Grant (No. DR25E8), and Lam Woo Research Fund (No. LWP20040), and by the Hong Kong Research Grants Council through the Faculty Development Schemes (No. UGC/FDS16/E10/23). Xianming Li and Jing Li are partially supported by a grant from the Hong Kong Research Grants Council (Project No. PolyU/25200821), the Innovation and Technology Fund (Project No. PRP/047/22FX), and a gift fund from Huawei (N-ZGM3). Qing Li has been supported by the Hong Kong Research Grants Council under the General Research Fund (No. 15216225).
Keywords
- Large language models
- Natural language processing
- Named entity recognition
- Sequence classification
- Token classification
Fingerprint
Dive into the research topics of 'LS-BiLLMs: Label supervised bi-directional large language models for token- and sequence-level information extraction'. Together they form a unique fingerprint.-
Hierarchical Evaluation Framework for AI Scientists: Beyond Surface Metrics to Deep Scientific Understanding
LI, Z. (PI)
1/08/25 → 31/07/26
Project: Grant Research
-
An Integrated Fake Financial News Detection Framework: Knowledge Graph, Large Language Models, Uncertainty Modeling, and Contrastive Learning
XIE, H. (PI)
1/07/25 → 30/06/27
Project: Grant Research
-
Scalable Sentence Representation with Mixture-of-Experts and Dy-namic Routing
LI, Z. (PI), CHEN, X. (CoI) & WANG, W. (CoI)
1/07/25 → 30/06/28
Project: Grant Research
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver