Skip to main navigation Skip to search Skip to main content

Learning to Instruct for Visual Instruction Tuning

  • Zhihan Zhou (Co-first author)
  • , Feng Hong (Co-first author)
  • , Jiaan Luo
  • , Yushi Ye
  • , Jiangchao Yao*
  • , Dongsheng Li
  • , Bo Han
  • , Ya Zhang
  • , Yanfeng Wang
  • *Corresponding author for this work

Research output: Chapter in book/report/conference proceedingConference proceedingpeer-review

Abstract

We propose L2T, an advancement of visual instruction tuning (VIT). While VIT equips Multimodal LLMs (MLLMs) with promising multimodal capabilities, the current design choices for VIT often result in overfitting and shortcut learning, potentially degrading performance. This gap arises from an overemphasis on instruction-following abilities, while neglecting the proactive understanding of visual information. Inspired by this, L2T adopts a simple yet effective approach by incorporating the loss function into both the instruction and response sequences. It seamlessly expands the training data, and regularizes the MLLMs from overly relying on language priors. Based on this merit, L2T achieves a significant relative improvement of up to 9% on comprehensive multimodal benchmarks, requiring no additional training data and incurring negligible computational overhead. Surprisingly, L2T attains exceptional fundamental visual capabilities, yielding up to an 18% improvement in captioning performance, while simultaneously alleviating hallucination in MLLMs. Github code: https://github.com/Feng-Hong/L2T.
Original languageEnglish
Title of host publication39th Conference on Neural Information Processing Systems, NeurIPS 2025
EditorsD. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen
PublisherNeural Information Processing Systems Foundation
Pages1-28
Number of pages28
Publication statusPublished - Dec 2025
Event39th Conference on Neural Information Processing Systems, NeurIPS 2025 - San Diego, United States
Duration: 2 Dec 20257 Dec 2025
https://neurips.cc/Conferences/2025 (Conference website)
https://neurips.cc/virtual/2025/loc/san-diego/papers.html (Conference schedule)
https://proceedings.neurips.cc/paper_files/paper/2025 (Conference proceedings)

Publication series

NameAdvances in Neural Information Processing Systems
Volume38
NameNeurIPS Proceedings

Conference

Conference39th Conference on Neural Information Processing Systems, NeurIPS 2025
Abbreviated titleNeurIPS 2025
Country/TerritoryUnited States
CitySan Diego
Period2/12/257/12/25
Internet address

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 9 - Industry, Innovation, and Infrastructure
    SDG 9 Industry, Innovation, and Infrastructure

Fingerprint

Dive into the research topics of 'Learning to Instruct for Visual Instruction Tuning'. Together they form a unique fingerprint.

Cite this