<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://muskai.github.io/feed.xml" rel="self" type="application/atom+xml"/><link href="https://muskai.github.io/" rel="alternate" type="text/html" hreflang="en"/><updated>2026-08-09T09:28:51+00:00</updated><id>https://muskai.github.io/feed.xml</id><title type="html">blank</title><subtitle>Haoran Chen is a computer science master&apos;s student researching multimodal language models, computer vision, and AI-generated content detection. </subtitle><entry><title type="html">a research thread on visual information in multimodal LLMs</title><link href="https://muskai.github.io/blog/2026/visual-information-in-multimodal-llms/" rel="alternate" type="text/html" title="a research thread on visual information in multimodal LLMs"/><published>2026-08-07T00:00:00+00:00</published><updated>2026-08-07T00:00:00+00:00</updated><id>https://muskai.github.io/blog/2026/visual-information-in-multimodal-llms</id><content type="html" xml:base="https://muskai.github.io/blog/2026/visual-information-in-multimodal-llms/"><![CDATA[<p>My recent work follows one question: how can multimodal language models make better use of the information produced by a visual encoder?</p> <p>The first study examines <strong>visual layer selection</strong> and asks where useful visual signals appear inside the encoder. The second studies <strong>multi-layer feature fusion</strong>, comparing ways to combine complementary information across layers. The third examines <strong>connector selection</strong>, focusing on what happens when visual information is preserved or compressed before it reaches the language model.</p> <p>Together, these projects form a connected research thread across visual representation, feature integration, and vision-language alignment. The corresponding papers and resources are available on the <a href="/publications/">publications page</a>.</p>]]></content><author><name></name></author><category term="research"/><category term="multimodal-llm"/><category term="computer-vision"/><summary type="html"><![CDATA[Three connected studies on visual layer selection, feature fusion, and connector design.]]></summary></entry></feed>