Abstract
The advent of large language models (LLMs) has sparked interest among educators across various fields, particularly in medical education. However, general-purpose pretrained LLMs may perform inadequately on medical examination questions without domain adaptation. Full fine-tuning of a pretrained LLM requires substantial computing resources, making it impractical for resource-constrained settings. In this work, we investigate a Composition to Augment Language Models (CALM) framework for medical domain adaptation in answering real-world medical examination questions. Specifically, we fine-tune a small pretrained language model on an open-ended medical question-answering dataset to inject medical domain knowledge, and then develop MedCALM by composing this medical-specialized augmenting model with an anchor LLM. Experiments on real-world medical exam question datasets show that MedCALM outperforms the compared baselines. The results indicate that using a multi-head cross-attention module to connect two LLMs improves performance on multiple-choice question answering. In addition, the selection of connected layers significantly affects the effectiveness of composition between the anchor and augmenting models. Our experiments suggest that cross-attention between the last few layers of the two models achieves promising performance.
| Original language | English |
|---|---|
| Article number | 100638 |
| Journal | Computers and Education: Artificial Intelligence |
| Volume | 11 |
| Early online date | 30 Jun 2026 |
| DOIs | |
| Publication status | E-pub ahead of print - 30 Jun 2026 |
Bibliographical note
Publisher Copyright:© 2026 The Authors.
Funding
The research described in this article has been supported by a grant from the Research Grants Council of the Hong Kong Special Administrative Region, China (UGC/FDS16/E17/23), and the Fund for Innovative Technology-in-Education (FITE) for Inter-institutional Collaborative Activities (IICA) Project (120045) entitled “Advancing Digital Competency for University Teachers and Students in the Era of Generative Artificial Intelligence” and the IICA Project (102745) entitled “GAI-Powered Feedback and Learning: A Domain-Adapted Large Language Model Approach for Bilingual and Data Science Education” funded by the Teaching Development and Language Enhancement Grant (TDLEG) from University Grants Council of the Hong Kong Special Administrative Region, the Katie Shu Sui Pui Charitable Trust — HKMU Transport and Publication Subsidy Scheme (Project Reference No. KSTP/2024/12).
Keywords
- Medical question answering
- Domain adaptation
- LLM composition
- Augmenting model
Fingerprint
Dive into the research topics of 'Enhancing domain adaptation of LLM via model composition in solving medical exam questions'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver