VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | He, Zhihao, Chen, Tieyuan, Wang, Kangyu, Qin, Ziran, Shao, Yang, Gan, Chaofan, Li, Shijie, Wu, Zuxuan, Lin, Weiyao |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
par: He, Zhihao, et autres
Publié: (2025)
par: He, Zhihao, et autres
Publié: (2025)
Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
par: Gan, Chaofan, et autres
Publié: (2025)
par: Gan, Chaofan, et autres
Publié: (2025)
Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
par: Chen, Tieyuan, et autres
Publié: (2025)
par: Chen, Tieyuan, et autres
Publié: (2025)
From Priors to Perception: Grounding Video-LLMs in Physical Reality
par: Zhao, Zicheng, et autres
Publié: (2026)
par: Zhao, Zicheng, et autres
Publié: (2026)
Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
par: Qin, Ziran, et autres
Publié: (2025)
par: Qin, Ziran, et autres
Publié: (2025)
MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning
par: Chen, Tieyuan, et autres
Publié: (2025)
par: Chen, Tieyuan, et autres
Publié: (2025)
Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations
par: Gan, Chaofan, et autres
Publié: (2025)
par: Gan, Chaofan, et autres
Publié: (2025)
OmniVid: A Generative Framework for Universal Video Understanding
par: Wang, Junke, et autres
Publié: (2024)
par: Wang, Junke, et autres
Publié: (2024)
MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning
par: Chen, Tieyuan, et autres
Publié: (2024)
par: Chen, Tieyuan, et autres
Publié: (2024)
CogStream: Context-guided Streaming Video Question Answering
par: Zhao, Zicheng, et autres
Publié: (2025)
par: Zhao, Zicheng, et autres
Publié: (2025)
DND: Boosting Large Language Models with Dynamic Nested Depth
par: Chen, Tieyuan, et autres
Publié: (2025)
par: Chen, Tieyuan, et autres
Publié: (2025)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
par: Shi, Jiapeng, et autres
Publié: (2026)
par: Shi, Jiapeng, et autres
Publié: (2026)
DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and Correction
par: Gan, Chaofan, et autres
Publié: (2024)
par: Gan, Chaofan, et autres
Publié: (2024)
CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit
par: Wang, Kangyu, et autres
Publié: (2025)
par: Wang, Kangyu, et autres
Publié: (2025)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
par: Lan, Xiaohan, et autres
Publié: (2024)
par: Lan, Xiaohan, et autres
Publié: (2024)
LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model
par: AI, Inclusion, et autres
Publié: (2026)
par: AI, Inclusion, et autres
Publié: (2026)
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
par: Wang, Guankun, et autres
Publié: (2025)
par: Wang, Guankun, et autres
Publié: (2025)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
par: Qin, Bosheng, et autres
Publié: (2023)
par: Qin, Bosheng, et autres
Publié: (2023)
Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
par: Qin, Ziran, et autres
Publié: (2025)
par: Qin, Ziran, et autres
Publié: (2025)
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
par: Li, Chaoyu, et autres
Publié: (2024)
par: Li, Chaoyu, et autres
Publié: (2024)
CSTA: Spatial-Temporal Causal Adaptive Learning for Exemplar-Free Video Class-Incremental Learning
par: Chen, Tieyuan, et autres
Publié: (2025)
par: Chen, Tieyuan, et autres
Publié: (2025)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
par: Huang, Xiaohu, et autres
Publié: (2024)
par: Huang, Xiaohu, et autres
Publié: (2024)
MCA: 2D-3D Retrieval with Noisy Labels via Multi-level Adaptive Correction and Alignment
par: Zou, Gui, et autres
Publié: (2025)
par: Zou, Gui, et autres
Publié: (2025)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
par: Xu, Yicheng, et autres
Publié: (2025)
par: Xu, Yicheng, et autres
Publié: (2025)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
par: Wang, Yi, et autres
Publié: (2023)
par: Wang, Yi, et autres
Publié: (2023)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
par: Fang, Ye, et autres
Publié: (2025)
par: Fang, Ye, et autres
Publié: (2025)
HieraVid: Hierarchical Token Pruning for Fast Video Large Language Models
par: Guo, Yansong, et autres
Publié: (2026)
par: Guo, Yansong, et autres
Publié: (2026)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
par: Yang, Zhoufaran, et autres
Publié: (2025)
par: Yang, Zhoufaran, et autres
Publié: (2025)
VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
par: Liang, Baoyu, et autres
Publié: (2025)
par: Liang, Baoyu, et autres
Publié: (2025)
Vid-SME: Membership Inference Attacks against Large Video Understanding Models
par: Li, Qi, et autres
Publié: (2025)
par: Li, Qi, et autres
Publié: (2025)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
par: Qiu, Zongyang, et autres
Publié: (2025)
par: Qiu, Zongyang, et autres
Publié: (2025)
VidDoS: Universal Denial-of-Service Attack on Video-based Large Language Models
par: Tang, Duoxun, et autres
Publié: (2026)
par: Tang, Duoxun, et autres
Publié: (2026)
Rodimus*: Breaking the Accuracy-Efficiency Trade-Off with Efficient Attentions
par: He, Zhihao, et autres
Publié: (2024)
par: He, Zhihao, et autres
Publié: (2024)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
par: Xiao, Xinyu, et autres
Publié: (2026)
par: Xiao, Xinyu, et autres
Publié: (2026)
FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
par: Xie, Yiweng, et autres
Publié: (2026)
par: Xie, Yiweng, et autres
Publié: (2026)
IG-Pruning: Input-Guided Block Pruning for Large Language Models
par: Qiao, Kangyu, et autres
Publié: (2025)
par: Qiao, Kangyu, et autres
Publié: (2025)
VidLA: Video-Language Alignment at Scale
par: Rizve, Mamshad Nayeem, et autres
Publié: (2024)
par: Rizve, Mamshad Nayeem, et autres
Publié: (2024)
AdaVid: Adaptive Video-Language Pretraining
par: Patel, Chaitanya, et autres
Publié: (2025)
par: Patel, Chaitanya, et autres
Publié: (2025)
Understanding Privacy Risks of Embeddings Induced by Large Language Models
par: Zhu, Zhihao, et autres
Publié: (2024)
par: Zhu, Zhihao, et autres
Publié: (2024)
WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models
par: He, Zijian, et autres
Publié: (2024)
par: He, Zijian, et autres
Publié: (2024)
Documents similaires
-
Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
par: He, Zhihao, et autres
Publié: (2025) -
Massive Activations are the Key to Local Detail Synthesis in Diffusion Transformers
par: Gan, Chaofan, et autres
Publié: (2025) -
Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
par: Chen, Tieyuan, et autres
Publié: (2025) -
From Priors to Perception: Grounding Video-LLMs in Physical Reality
par: Zhao, Zicheng, et autres
Publié: (2026) -
Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
par: Qin, Ziran, et autres
Publié: (2025)