Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Gen, Zhou, Yiyi, Zhang, Yuxin, Zheng, Xiawu, Sun, Xiaoshuai, Ji, Rongrong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
$γ-$MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
von: Luo, Yaxin, et al.
Veröffentlicht: (2024)
von: Luo, Yaxin, et al.
Veröffentlicht: (2024)
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression
von: Tong, Bo, et al.
Veröffentlicht: (2024)
von: Tong, Bo, et al.
Veröffentlicht: (2024)
Grounded Chain-of-Thought for Multimodal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
Deep Instruction Tuning for Segment Anything Model
von: Huang, Xiaorui, et al.
Veröffentlicht: (2024)
von: Huang, Xiaorui, et al.
Veröffentlicht: (2024)
INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric Regularization
von: Zhang, Jinlu, et al.
Veröffentlicht: (2024)
von: Zhang, Jinlu, et al.
Veröffentlicht: (2024)
Omni-Referring Image Segmentation
von: Zheng, Qiancheng, et al.
Veröffentlicht: (2025)
von: Zheng, Qiancheng, et al.
Veröffentlicht: (2025)
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
von: Chen, Tao, et al.
Veröffentlicht: (2026)
von: Chen, Tao, et al.
Veröffentlicht: (2026)
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
von: Wu, Qiong, et al.
Veröffentlicht: (2024)
Multi-branch Collaborative Learning Network for 3D Visual Grounding
von: Qian, Zhipeng, et al.
Veröffentlicht: (2024)
von: Qian, Zhipeng, et al.
Veröffentlicht: (2024)
Test-Time Computing for Referring Multimodal Large Language Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
von: Zhou, Ziyin, et al.
Veröffentlicht: (2025)
von: Zhou, Ziyin, et al.
Veröffentlicht: (2025)
MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
von: Li, Jiale, et al.
Veröffentlicht: (2025)
von: Li, Jiale, et al.
Veröffentlicht: (2025)
Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient Adaptation
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
Few-Shot Image Quality Assessment via Adaptation of Vision-Language Models
von: Li, Xudong, et al.
Veröffentlicht: (2024)
von: Li, Xudong, et al.
Veröffentlicht: (2024)
AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection
von: Zhang, Shuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Shuheng, et al.
Veröffentlicht: (2025)
MICON-Bench: Benchmarking and Enhancing Multi-Image Context Image Generation in Unified Multimodal Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
Image Captioning via Dynamic Path Customization
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
Prototype-Based Test-Time Adaptation of Vision-Language Models
von: Huang, Zhaohong, et al.
Veröffentlicht: (2026)
von: Huang, Zhaohong, et al.
Veröffentlicht: (2026)
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
von: Fu, Chaoyou, et al.
Veröffentlicht: (2023)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2023)
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
von: Zhou, Chenyu, et al.
Veröffentlicht: (2024)
Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval
von: Chen, Tao, et al.
Veröffentlicht: (2025)
von: Chen, Tao, et al.
Veröffentlicht: (2025)
GS-Bias: Global-Spatial Bias Learner for Single-Image Test-Time Adaptation of Vision-Language Models
von: Huang, Zhaohong, et al.
Veröffentlicht: (2025)
von: Huang, Zhaohong, et al.
Veröffentlicht: (2025)
HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
von: Gong, Ziyang, et al.
Veröffentlicht: (2025)
von: Gong, Ziyang, et al.
Veröffentlicht: (2025)
3D-GRES: Generalized 3D Referring Expression Segmentation
von: Wu, Changli, et al.
Veröffentlicht: (2024)
von: Wu, Changli, et al.
Veröffentlicht: (2024)
ForestPrune: High-ratio Visual Token Compression for Video Multimodal Large Language Models via Spatial-Temporal Forest Modeling
von: Ju, Shaobo, et al.
Veröffentlicht: (2026)
von: Ju, Shaobo, et al.
Veröffentlicht: (2026)
CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
Any-to-3D Generation via Hybrid Diffusion Supervision
von: Fan, Yijun, et al.
Veröffentlicht: (2024)
von: Fan, Yijun, et al.
Veröffentlicht: (2024)
Speculative Decoding Reimagined for Multimodal Large Language Models
von: Lin, Luxi, et al.
Veröffentlicht: (2025)
von: Lin, Luxi, et al.
Veröffentlicht: (2025)
Earth-Adapter: Bridge the Geospatial Domain Gaps with Mixture of Frequency Adaptation
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
StealthDiffusion: Towards Evading Diffusion Forensic Detection through Diffusion Model
von: Zhou, Ziyin, et al.
Veröffentlicht: (2024)
von: Zhou, Ziyin, et al.
Veröffentlicht: (2024)
Depth-Guided Semi-Supervised Instance Segmentation
von: Chen, Xin, et al.
Veröffentlicht: (2024)
von: Chen, Xin, et al.
Veröffentlicht: (2024)
RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation
von: Wu, Changli, et al.
Veröffentlicht: (2024)
von: Wu, Changli, et al.
Veröffentlicht: (2024)
Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy
von: Shen, Yunhang, et al.
Veröffentlicht: (2025)
von: Shen, Yunhang, et al.
Veröffentlicht: (2025)
An Efficient and Mixed Heterogeneous Model for Image Restoration
von: Gu, Yubin, et al.
Veröffentlicht: (2025)
von: Gu, Yubin, et al.
Veröffentlicht: (2025)
M4-BLIP: Advancing Multi-Modal Media Manipulation Detection through Face-Enhanced Local Analysis
von: Wu, Hang, et al.
Veröffentlicht: (2025)
von: Wu, Hang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
$γ-$MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
von: Luo, Yaxin, et al.
Veröffentlicht: (2024) -
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression
von: Tong, Bo, et al.
Veröffentlicht: (2024) -
Grounded Chain-of-Thought for Multimodal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2025) -
Deep Instruction Tuning for Segment Anything Model
von: Huang, Xiaorui, et al.
Veröffentlicht: (2024) -
INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)