Learning What to See: Efficient Multimodal Reasoning with Vision-Language Models

Author

Sunandan Chakraborty

Published

October 9, 2026

Illegal wildlife trade is being increasingly conducted through online channels, posing a significant risk to global biodiversity and environmental sustainability. Online wildlife marketplaces present a challenging multimodal learning problem: identifying species and product types from noisy, incomplete, and heterogeneous combinations of images and text. This talk presents a modular framework that combines Vision-Language Models (VLMs) and Large Language Models (LLMs) for multimodal classification and structured information extraction. Using online shark-product advertisements as a case study, we evaluate the framework on species and product identification and examine its ability to generalize to an out-of-distribution task: classifying products by anatomical body part. Building on this work, our current research focuses on making multimodal models more efficient and selective in how they process information across modalities. I will present TaMe, a training-free, text-aware visual token merging method that uses textual context to preserve question-relevant visual tokens while compressing redundant ones, improving the accuracy–efficiency trade-off of VLM inference. I will also discuss ongoing work on streaming video understanding, where we are developing methods for more efficient video processing and long-term visual memory, with the goal of allocating limited visual context to the most relevant information over time.

The second part of the talk will shift to my research at the intersection of large language models and causality. Together, these research directions examine how foundation models can be made more efficient and adaptable, while also investigating what these models learn and how they reason.