Logo image
MedVisualLLM: A Parameter-Efficient Multimodal Alignment Framework for Medical Visual Question Answering
Conference paper

MedVisualLLM: A Parameter-Efficient Multimodal Alignment Framework for Medical Visual Question Answering

Fatima Zulfiqar, Kok Wai Wong, Hamid Laga, Deval Mehta and Guanjin Wang
Machine Learning and Knowledge Engineering for Decision Making, pp.251-276
Lecture Notes in Computer Science, Springer Nature Singapore
The 17th International FLINS Conference on Fuzzy Logic for Intelligent Systems and The 21st International Conference on Intelligent Systems and Knowledge Engineering (Sydney, NSW, Australia, 15/07/2026–19/07/2026)
2027

Abstract

Low-Rank Adaptation Medical Visual Question Answering Mixture-of-Experts Multimodal Learning Parameter-Efficient Fine-Tuning Visual Language Model
Medical Visual Question Answering (VQA) requires models to reason over medical images and generate clinically meaningful responses. While recent approaches increasingly rely on multimodal foundation models and full-parameter fine-tuning, such methods impose substantial computational and memory costs. In this work, we present MedVisualLLM, a lightweight generative framework for medical VQA that balances efficiency and performance. The proposed architecture integrates a biomedical vision encoder with a question-aware Mixture-of-Experts (MoE) projection module and a parameter-efficient large language model. Visual features are aligned with the language embedding space and incorporated as contextual tokens, enabling multimodal reasoning without requiring specialized cross-modal transformer architectures. The training follows a two-stage strategy: domain alignment on PMC-VQA and task-specific fine-tuning with Low Rank Adaptation. Experiments on both closed-ended and open-ended questions on Path-VQA and SLAKE demonstrate that MedVisualLLM obtained 91.97% 91.97% closed accuracy and 35.81% 35.81% open accuracy on Path-VQA, along with 88.22% 88.22% and 82.79% 82.79% on SLAKE, respectively. These results are obtained while updating only 2–4% of the total model parameters. Although large fully fine-tuned multimodal Vision Language Models achieve higher open-ended accuracy, especially on Path-VQA, the proposed framework maintains competitive closed accuracy while substantially reducing training overhead. These results suggest that efficient domain alignment and parameter-efficient fine-tuning may reduce reliance on full-parameter training in medical VQA under resource-constrained settings.

Details

Metrics

5 Record Views
Logo image