publications

Selected peer-reviewed publications by Haoran Chen on multimodal language models and visual representation.

paper roulette

Not sure where to begin? Let chance choose a research question.

Preview for Multimodal Language Models See Better When They Look Shallower

EMNLP 2025 Main, Oral

Multimodal Language Models See Better When They Look Shallower

Which visual encoder layer should a multimodal language model actually use?

Later is not always better. Shallower visual layers can preserve information that deeper layers discard.

2025

  1. Multimodal Language Models See Better When They Look Shallower
    Haoran Chen, Junyan Lin, Xinhao Chen, Yue Fan, Xin Jin, Hui Su, Jianfeng Dong, Jinlan Fu, and Xiaoyu Shen
    In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025
    Main Conference, Oral Presentation
  2. Multi-Layer Visual Feature Fusion in Multimodal LLMs: Methods, Analysis, and Best Practices
    Junyan Lin*, Haoran Chen*, Yue Fan, Yingqi Fan, Xin Jin, Hui Su, Jinlan Fu, and Xiaoyu Shen
    In Proceedings of the Computer Vision and Pattern Recognition Conference, 2025

2024

  1. To Preserve or To Compress: An In-Depth Study of Connector Selection in Multimodal Large Language Models
    Junyan Lin, Haoran Chen, Dawei Zhu, and Xiaoyu Shen
    In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024