D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yiyang, Wang, Yizhou, Fu, Yun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoDe-NeRF: Neural Rendering via Dynamic Coefficient Decomposition
by: Xing, Wenpeng, et al.
Published: (2025)
by: Xing, Wenpeng, et al.
Published: (2025)
Distorted or Fabricated? A Survey on Hallucination in Video LLMs
by: Huang, Yiyang, et al.
Published: (2026)
by: Huang, Yiyang, et al.
Published: (2026)
Spatial Transcriptomics as Images for Large-Scale Pretraining
by: Zhu, Yishun, et al.
Published: (2026)
by: Zhu, Yishun, et al.
Published: (2026)
Towards Lossless Ultimate Vision Token Compression for VLMs
by: Zheng, Dehua, et al.
Published: (2025)
by: Zheng, Dehua, et al.
Published: (2025)
SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
by: Huang, Yiyang, et al.
Published: (2025)
by: Huang, Yiyang, et al.
Published: (2025)
Parameter-Efficient VLMs for Gastrointestinal Endoscopy: Medical Image Generation and Clinical Visual Question Answering
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2026)
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2026)
Don't Judge by the Look: Towards Motion Coherent Video Representation
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models
by: Chen, Zhawnen, et al.
Published: (2024)
by: Chen, Zhawnen, et al.
Published: (2024)
3D MRI Image Pretraining via Controllable 2D Slice Navigation Task
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering
by: Yu, Ting, et al.
Published: (2024)
by: Yu, Ting, et al.
Published: (2024)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
by: Wang, Xingrui, et al.
Published: (2024)
by: Wang, Xingrui, et al.
Published: (2024)
Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
by: Duan, Yicheng, et al.
Published: (2025)
by: Duan, Yicheng, et al.
Published: (2025)
BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
by: Wang, Shengao, et al.
Published: (2025)
by: Wang, Shengao, et al.
Published: (2025)
Stateful Token Reduction for Long-Video Hybrid VLMs
by: Jiang, Jindong, et al.
Published: (2026)
by: Jiang, Jindong, et al.
Published: (2026)
Long-Text-to-Image Generation via Compositional Prompt Decomposition
by: Huang, Jen-Yuan, et al.
Published: (2026)
by: Huang, Jen-Yuan, et al.
Published: (2026)
Listener-Rewarded Thinking in VLMs for Image Preferences
by: Gambashidze, Alexander, et al.
Published: (2025)
by: Gambashidze, Alexander, et al.
Published: (2025)
DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs
by: Wang, Zhenhailong, et al.
Published: (2025)
by: Wang, Zhenhailong, et al.
Published: (2025)
CoDe: Blockwise Control for Denoising Diffusion Models
by: Singh, Anuj, et al.
Published: (2025)
by: Singh, Anuj, et al.
Published: (2025)
Location-Aware Pretraining for Medical Difference Visual Question Answering
by: Musinguzi, Denis, et al.
Published: (2026)
by: Musinguzi, Denis, et al.
Published: (2026)
Early Exit and Multi Stage Knowledge Distillation in VLMs for Video Summarization
by: Khan, Anas Anwarul Haq, et al.
Published: (2025)
by: Khan, Anas Anwarul Haq, et al.
Published: (2025)
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining
by: Xu, Boshen, et al.
Published: (2025)
by: Xu, Boshen, et al.
Published: (2025)
ScaleNet: Scaling up Pretrained Neural Networks with Incremental Parameters
by: Hao, Zhiwei, et al.
Published: (2025)
by: Hao, Zhiwei, et al.
Published: (2025)
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
by: Yu, Ting, et al.
Published: (2024)
by: Yu, Ting, et al.
Published: (2024)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
by: Zhou, Guanyu, et al.
Published: (2026)
by: Zhou, Guanyu, et al.
Published: (2026)
Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks
by: Li, Chenjun
Published: (2026)
by: Li, Chenjun
Published: (2026)
Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality
by: Toibazar, Daulet, et al.
Published: (2025)
by: Toibazar, Daulet, et al.
Published: (2025)
3D Primitives are a Spatial Language for VLMs
by: Liu, Junze, et al.
Published: (2026)
by: Liu, Junze, et al.
Published: (2026)
AdaVid: Adaptive Video-Language Pretraining
by: Patel, Chaitanya, et al.
Published: (2025)
by: Patel, Chaitanya, et al.
Published: (2025)
DELST: Dual Entailment Learning for Hyperbolic Image-Gene Pretraining in Spatial Transcriptomics
by: Chen, Xulin, et al.
Published: (2025)
by: Chen, Xulin, et al.
Published: (2025)
Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
by: Xu, Zhaoqi, et al.
Published: (2025)
by: Xu, Zhaoqi, et al.
Published: (2025)
LCV2: An Efficient Pretraining-Free Framework for Grounded Visual Question Answering
by: Chen, Yuhan, et al.
Published: (2024)
by: Chen, Yuhan, et al.
Published: (2024)
CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
by: Han, Hongyong, et al.
Published: (2025)
by: Han, Hongyong, et al.
Published: (2025)
Skeleton-to-Image Encoding: Enabling Skeleton Representation Learning via Vision-Pretrained Models
by: Yang, Siyuan, et al.
Published: (2026)
by: Yang, Siyuan, et al.
Published: (2026)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
by: Ji, Longbin, et al.
Published: (2026)
by: Ji, Longbin, et al.
Published: (2026)
GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs
by: Eltahir, Mohamed, et al.
Published: (2026)
by: Eltahir, Mohamed, et al.
Published: (2026)
TAR-TVG: Enhancing VLMs with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding
by: Guo, Chaohong, et al.
Published: (2025)
by: Guo, Chaohong, et al.
Published: (2025)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
Edge Reliability Gap in Vision-Language Models: Quantifying Failure Modes of Compressed VLMs Under Visual Corruption
by: Erol, Mehmet Kaan
Published: (2026)
by: Erol, Mehmet Kaan
Published: (2026)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
by: Oh, Ju-Young, et al.
Published: (2025)
by: Oh, Ju-Young, et al.
Published: (2025)
Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge
by: Lu, Shuai, et al.
Published: (2026)
by: Lu, Shuai, et al.
Published: (2026)
Similar Items
-
CoDe-NeRF: Neural Rendering via Dynamic Coefficient Decomposition
by: Xing, Wenpeng, et al.
Published: (2025) -
Distorted or Fabricated? A Survey on Hallucination in Video LLMs
by: Huang, Yiyang, et al.
Published: (2026) -
Spatial Transcriptomics as Images for Large-Scale Pretraining
by: Zhu, Yishun, et al.
Published: (2026) -
Towards Lossless Ultimate Vision Token Compression for VLMs
by: Zheng, Dehua, et al.
Published: (2025) -
SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
by: Huang, Yiyang, et al.
Published: (2025)