Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jain, Jitesh, Yang, Zhengyuan, Shi, Humphrey, Gao, Jianfeng, Yang, Jianwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking Object Detectors with COCO: A New Path Forward
von: Singh, Shweta, et al.
Veröffentlicht: (2024)
von: Singh, Shweta, et al.
Veröffentlicht: (2024)
List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
von: Yan, An, et al.
Veröffentlicht: (2024)
von: Yan, An, et al.
Veröffentlicht: (2024)
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
von: Li, Honglin, et al.
Veröffentlicht: (2024)
von: Li, Honglin, et al.
Veröffentlicht: (2024)
DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
von: Meng, Lingchen, et al.
Veröffentlicht: (2024)
von: Meng, Lingchen, et al.
Veröffentlicht: (2024)
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
Towards Flexible Visual Relationship Segmentation
von: Zhu, Fangrui, et al.
Veröffentlicht: (2024)
von: Zhu, Fangrui, et al.
Veröffentlicht: (2024)
Slow-Fast Architecture for Video Multi-Modal Large Language Models
von: Shi, Min, et al.
Veröffentlicht: (2025)
von: Shi, Min, et al.
Veröffentlicht: (2025)
Towards Understanding Graphical Perception in Large Multimodal Models
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
von: Zhang, Kai, et al.
Veröffentlicht: (2025)
Attention Distillation: A Unified Approach to Visual Characteristics Transfer
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
von: Fu, Xingyu, et al.
Veröffentlicht: (2025)
von: Fu, Xingyu, et al.
Veröffentlicht: (2025)
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
von: Li, Qi, et al.
Veröffentlicht: (2026)
von: Li, Qi, et al.
Veröffentlicht: (2026)
SITE: towards Spatial Intelligence Thorough Evaluation
von: Wang, Wenqi, et al.
Veröffentlicht: (2025)
von: Wang, Wenqi, et al.
Veröffentlicht: (2025)
Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs
von: Kanade, Aditya, et al.
Veröffentlicht: (2025)
von: Kanade, Aditya, et al.
Veröffentlicht: (2025)
Pix2Gif: Motion-Guided Diffusion for GIF Generation
von: Kandala, Hitesh, et al.
Veröffentlicht: (2024)
von: Kandala, Hitesh, et al.
Veröffentlicht: (2024)
Matryoshka Multimodal Models
von: Cai, Mu, et al.
Veröffentlicht: (2024)
von: Cai, Mu, et al.
Veröffentlicht: (2024)
Adversarial Robustness for Visual Grounding of Multimodal Large Language Models
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
von: Guan, Tongkun, et al.
Veröffentlicht: (2026)
von: Guan, Tongkun, et al.
Veröffentlicht: (2026)
One-Step Diffusion-based Real-World Image Super-Resolution with Visual Perception Distillation
von: Wu, Xue, et al.
Veröffentlicht: (2025)
von: Wu, Xue, et al.
Veröffentlicht: (2025)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
Visual Bridge: Universal Visual Perception Representations Generating
von: Gao, Yilin, et al.
Veröffentlicht: (2025)
von: Gao, Yilin, et al.
Veröffentlicht: (2025)
Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge
von: Wu, Xiangyu, et al.
Veröffentlicht: (2024)
von: Wu, Xiangyu, et al.
Veröffentlicht: (2024)
Customized Visual Storytelling with Unified Multimodal LLMs
von: Li, Wei-Hua, et al.
Veröffentlicht: (2026)
von: Li, Wei-Hua, et al.
Veröffentlicht: (2026)
FishRoPE: Projective Rotary Position Embeddings for Omnidirectional Visual Perception
von: Ahuja, Rahul, et al.
Veröffentlicht: (2026)
von: Ahuja, Rahul, et al.
Veröffentlicht: (2026)
Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving
von: Zhou, Hao, et al.
Veröffentlicht: (2024)
von: Zhou, Hao, et al.
Veröffentlicht: (2024)
Understanding Depth and Height Perception in Large Visual-Language Models
von: Azad, Shehreen, et al.
Veröffentlicht: (2024)
von: Azad, Shehreen, et al.
Veröffentlicht: (2024)
SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
von: Jain, Jitesh, et al.
Veröffentlicht: (2025)
von: Jain, Jitesh, et al.
Veröffentlicht: (2025)
Adaptive Perception for Unified Visual Multi-modal Object Tracking
von: Hu, Xiantao, et al.
Veröffentlicht: (2025)
von: Hu, Xiantao, et al.
Veröffentlicht: (2025)
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
Learning Brain Representation with Hierarchical Visual Embeddings
von: Zheng, Jiawen, et al.
Veröffentlicht: (2026)
von: Zheng, Jiawen, et al.
Veröffentlicht: (2026)
Multi-Teacher Knowledge Distillation with Reinforcement Learning for Visual Recognition
von: Yang, Chuanguang, et al.
Veröffentlicht: (2025)
von: Yang, Chuanguang, et al.
Veröffentlicht: (2025)
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
von: Zhou, Yikang, et al.
Veröffentlicht: (2025)
von: Zhou, Yikang, et al.
Veröffentlicht: (2025)
Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception
von: He, Junwen, et al.
Veröffentlicht: (2024)
von: He, Junwen, et al.
Veröffentlicht: (2024)
Multimodal Rationales for Explainable Visual Question Answering
von: Li, Kun, et al.
Veröffentlicht: (2024)
von: Li, Kun, et al.
Veröffentlicht: (2024)
Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation
von: Zhai, Yuanhao, et al.
Veröffentlicht: (2024)
von: Zhai, Yuanhao, et al.
Veröffentlicht: (2024)
KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering
von: Li, Zhiyang, et al.
Veröffentlicht: (2026)
von: Li, Zhiyang, et al.
Veröffentlicht: (2026)
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
Instruction Tuning-free Visual Token Complement for Multimodal LLMs
von: Wang, Dongsheng, et al.
Veröffentlicht: (2024)
von: Wang, Dongsheng, et al.
Veröffentlicht: (2024)
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
von: Jiang, Tianxiang, et al.
Veröffentlicht: (2025)
von: Jiang, Tianxiang, et al.
Veröffentlicht: (2025)
Vision Function Layer in Multimodal LLMs
von: Shi, Cheng, et al.
Veröffentlicht: (2025)
von: Shi, Cheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Benchmarking Object Detectors with COCO: A New Path Forward
von: Singh, Shweta, et al.
Veröffentlicht: (2024) -
List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
von: Yan, An, et al.
Veröffentlicht: (2024) -
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts
von: Li, Jiachen, et al.
Veröffentlicht: (2024) -
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
von: Li, Honglin, et al.
Veröffentlicht: (2024) -
DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
von: Meng, Lingchen, et al.
Veröffentlicht: (2024)