DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Zhenhailong, Purushwalkam, Senthil, Xiong, Caiming, Savarese, Silvio, Ji, Heng, Xu, Ran |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Trust but Verify: Programmatic VLM Evaluation in the Wild
di: Prabhu, Viraj, et al.
Pubblicazione: (2024)
di: Prabhu, Viraj, et al.
Pubblicazione: (2024)
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions
di: Awadalla, Anas, et al.
Pubblicazione: (2024)
di: Awadalla, Anas, et al.
Pubblicazione: (2024)
BootPIG: Bootstrapping Zero-shot Personalized Image Generation Capabilities in Pretrained Diffusion Models
di: Purushwalkam, Senthil, et al.
Pubblicazione: (2024)
di: Purushwalkam, Senthil, et al.
Pubblicazione: (2024)
xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations
di: Qin, Can, et al.
Pubblicazione: (2024)
di: Qin, Can, et al.
Pubblicazione: (2024)
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
di: Ryoo, Michael S., et al.
Pubblicazione: (2024)
di: Ryoo, Michael S., et al.
Pubblicazione: (2024)
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
di: Chen, Jiuhai, et al.
Pubblicazione: (2025)
di: Chen, Jiuhai, et al.
Pubblicazione: (2025)
WALT: Web Agents that Learn Tools
di: Prabhu, Viraj, et al.
Pubblicazione: (2025)
di: Prabhu, Viraj, et al.
Pubblicazione: (2025)
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
di: Zhang, Jieyu, et al.
Pubblicazione: (2024)
di: Zhang, Jieyu, et al.
Pubblicazione: (2024)
What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
di: Kang, Inha, et al.
Pubblicazione: (2025)
di: Kang, Inha, et al.
Pubblicazione: (2025)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
di: Zhou, Honglu, et al.
Pubblicazione: (2025)
di: Zhou, Honglu, et al.
Pubblicazione: (2025)
PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging
di: Shao, Zibo, et al.
Pubblicazione: (2026)
di: Shao, Zibo, et al.
Pubblicazione: (2026)
Visually Descriptive Language Model for Vector Graphics Reasoning
di: Wang, Zhenhailong, et al.
Pubblicazione: (2024)
di: Wang, Zhenhailong, et al.
Pubblicazione: (2024)
DyABD: The Abdominal Muscle Segmentation in Dynamic MRI Benchmark
di: Belton, Niamh, et al.
Pubblicazione: (2026)
di: Belton, Niamh, et al.
Pubblicazione: (2026)
SYNTHIA: Novel Concept Design with Affordance Composition
di: Ha, Hyeonjeong, et al.
Pubblicazione: (2025)
di: Ha, Hyeonjeong, et al.
Pubblicazione: (2025)
DyRA: Portable Dynamic Resolution Adjustment Network for Existing Detectors
di: Seo, Daeun, et al.
Pubblicazione: (2023)
di: Seo, Daeun, et al.
Pubblicazione: (2023)
DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation
di: Ye, Bo, et al.
Pubblicazione: (2026)
di: Ye, Bo, et al.
Pubblicazione: (2026)
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
di: Chen, Yanlong, et al.
Pubblicazione: (2026)
di: Chen, Yanlong, et al.
Pubblicazione: (2026)
Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
di: Shabtay, Nimrod, et al.
Pubblicazione: (2026)
di: Shabtay, Nimrod, et al.
Pubblicazione: (2026)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
di: Berman, Shmuel, et al.
Pubblicazione: (2025)
di: Berman, Shmuel, et al.
Pubblicazione: (2025)
VACoT: Rethinking Visual Data Augmentation with VLMs
di: Xu, Zhengzhuo, et al.
Pubblicazione: (2025)
di: Xu, Zhengzhuo, et al.
Pubblicazione: (2025)
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
di: Xue, Le, et al.
Pubblicazione: (2024)
di: Xue, Le, et al.
Pubblicazione: (2024)
Accurate and Efficient Low-Rank Model Merging in Core Space
di: Panariello, Aniello, et al.
Pubblicazione: (2025)
di: Panariello, Aniello, et al.
Pubblicazione: (2025)
Parameter-Efficient VLMs for Gastrointestinal Endoscopy: Medical Image Generation and Clinical Visual Question Answering
di: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Pubblicazione: (2026)
di: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Pubblicazione: (2026)
Concept-as-Tree: A Controllable Synthetic Data Framework Makes Stronger Personalized VLMs
di: An, Ruichuan, et al.
Pubblicazione: (2025)
di: An, Ruichuan, et al.
Pubblicazione: (2025)
Caption This, Reason That: VLMs Caught in the Middle
di: Weng, Zihan, et al.
Pubblicazione: (2025)
di: Weng, Zihan, et al.
Pubblicazione: (2025)
HIVE: Harnessing Human Feedback for Instructional Visual Editing
di: Zhang, Shu, et al.
Pubblicazione: (2023)
di: Zhang, Shu, et al.
Pubblicazione: (2023)
TAR-TVG: Enhancing VLMs with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding
di: Guo, Chaohong, et al.
Pubblicazione: (2025)
di: Guo, Chaohong, et al.
Pubblicazione: (2025)
Learning to Merge Tokens via Decoupled Embedding for Efficient Vision Transformers
di: Lee, Dong Hoon, et al.
Pubblicazione: (2024)
di: Lee, Dong Hoon, et al.
Pubblicazione: (2024)
ProMerge: Prompt and Merge for Unsupervised Instance Segmentation
di: Li, Dylan, et al.
Pubblicazione: (2024)
di: Li, Dylan, et al.
Pubblicazione: (2024)
D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
di: Huang, Yiyang, et al.
Pubblicazione: (2025)
di: Huang, Yiyang, et al.
Pubblicazione: (2025)
Evaluating Compositional Generalisation in VLMs and Diffusion Models
di: Pearson, Beth, et al.
Pubblicazione: (2025)
di: Pearson, Beth, et al.
Pubblicazione: (2025)
Listener-Rewarded Thinking in VLMs for Image Preferences
di: Gambashidze, Alexander, et al.
Pubblicazione: (2025)
di: Gambashidze, Alexander, et al.
Pubblicazione: (2025)
Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning
di: Tan, Zhangyun, et al.
Pubblicazione: (2026)
di: Tan, Zhangyun, et al.
Pubblicazione: (2026)
ImDy: Human Inverse Dynamics from Imitated Observations
di: Liu, Xinpeng, et al.
Pubblicazione: (2024)
di: Liu, Xinpeng, et al.
Pubblicazione: (2024)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
di: Zhang, Jianrui, et al.
Pubblicazione: (2026)
di: Zhang, Jianrui, et al.
Pubblicazione: (2026)
DeepMerge: Deep-Learning-Based Region-Merging for Image Segmentation
di: Lv, Xianwei, et al.
Pubblicazione: (2023)
di: Lv, Xianwei, et al.
Pubblicazione: (2023)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
Same Answer, Different Representations: Hidden instability in VLMs
di: Wani, Farooq Ahmad, et al.
Pubblicazione: (2026)
di: Wani, Farooq Ahmad, et al.
Pubblicazione: (2026)
Stateful Token Reduction for Long-Video Hybrid VLMs
di: Jiang, Jindong, et al.
Pubblicazione: (2026)
di: Jiang, Jindong, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Trust but Verify: Programmatic VLM Evaluation in the Wild
di: Prabhu, Viraj, et al.
Pubblicazione: (2024) -
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions
di: Awadalla, Anas, et al.
Pubblicazione: (2024) -
BootPIG: Bootstrapping Zero-shot Personalized Image Generation Capabilities in Pretrained Diffusion Models
di: Purushwalkam, Senthil, et al.
Pubblicazione: (2024) -
xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations
di: Qin, Can, et al.
Pubblicazione: (2024) -
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
di: Ryoo, Michael S., et al.
Pubblicazione: (2024)