Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning Synergy
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Haijier, Xu, Bo, Zhang, Shoujian, Liu, Haoze, Lin, Jiaxuan, Wang, Jingrong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
3DGS_LSR:Large_Scale Relocation for Autonomous Driving Based on 3D Gaussian Splatting
by: Lu, Haitao, et al.
Published: (2025)
by: Lu, Haitao, et al.
Published: (2025)
Sky-GVIO: an enhanced GNSS/INS/Vision navigation with FCN-based sky-segmentation in urban canyon
by: Wang, Jingrong, et al.
Published: (2024)
by: Wang, Jingrong, et al.
Published: (2024)
CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
by: Li, Jingyao, et al.
Published: (2025)
by: Li, Jingyao, et al.
Published: (2025)
SafeVid: Toward Safety Aligned Video Large Multimodal Models
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
HiVid: LLM-Guided Video Saliency For Content-Aware VOD And Live Streaming
by: Chen, Jiahui, et al.
Published: (2026)
by: Chen, Jiahui, et al.
Published: (2026)
Object Navigation with Structure-Semantic Reasoning-Based Multi-level Map and Multimodal Decision-Making LLM
by: Yan, Chongshang, et al.
Published: (2025)
by: Yan, Chongshang, et al.
Published: (2025)
How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM
by: Zha, Jirong, et al.
Published: (2025)
by: Zha, Jirong, et al.
Published: (2025)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
by: Fang, Ye, et al.
Published: (2025)
by: Fang, Ye, et al.
Published: (2025)
VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer
by: Lin, Rui, et al.
Published: (2026)
by: Lin, Rui, et al.
Published: (2026)
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
by: Feng, Weixi, et al.
Published: (2025)
by: Feng, Weixi, et al.
Published: (2025)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
by: Fu, Rao, et al.
Published: (2024)
by: Fu, Rao, et al.
Published: (2024)
AdaVid: Adaptive Video-Language Pretraining
by: Patel, Chaitanya, et al.
Published: (2025)
by: Patel, Chaitanya, et al.
Published: (2025)
SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs
by: Li, Jiawei, et al.
Published: (2026)
by: Li, Jiawei, et al.
Published: (2026)
VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
by: Zheng, Haojie, et al.
Published: (2024)
by: Zheng, Haojie, et al.
Published: (2024)
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
by: Tang, Yolo Y., et al.
Published: (2024)
by: Tang, Yolo Y., et al.
Published: (2024)
DualPrim: Compact 3D Reconstruction with Positive and Negative Primitives
by: Meng, Xiaoxu, et al.
Published: (2026)
by: Meng, Xiaoxu, et al.
Published: (2026)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
by: Li, Chenglin, et al.
Published: (2026)
by: Li, Chenglin, et al.
Published: (2026)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
by: Tzachor, Issar, et al.
Published: (2026)
by: Tzachor, Issar, et al.
Published: (2026)
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
by: He, Zhihao, et al.
Published: (2026)
by: He, Zhihao, et al.
Published: (2026)
VidDoS: Universal Denial-of-Service Attack on Video-based Large Language Models
by: Tang, Duoxun, et al.
Published: (2026)
by: Tang, Duoxun, et al.
Published: (2026)
VidTwin: Video VAE with Decoupled Structure and Dynamics
by: Wang, Yuchi, et al.
Published: (2024)
by: Wang, Yuchi, et al.
Published: (2024)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
by: Xu, Yicheng, et al.
Published: (2025)
by: Xu, Yicheng, et al.
Published: (2025)
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
by: Wang, Ziyang, et al.
Published: (2024)
by: Wang, Ziyang, et al.
Published: (2024)
VidSketch: Hand-drawn Sketch-Driven Video Generation with Diffusion Control
by: Jiang, Lifan, et al.
Published: (2025)
by: Jiang, Lifan, et al.
Published: (2025)
VMID: A Multimodal Fusion LLM Framework for Detecting and Identifying Misinformation of Short Videos
by: Zhong, Weihao, et al.
Published: (2024)
by: Zhong, Weihao, et al.
Published: (2024)
Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation
by: Zhao, Bingrui, et al.
Published: (2025)
by: Zhao, Bingrui, et al.
Published: (2025)
ToxVidLM: A Multimodal Framework for Toxicity Detection in Code-Mixed Videos
by: Maity, Krishanu, et al.
Published: (2024)
by: Maity, Krishanu, et al.
Published: (2024)
VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors
by: Tang, Jimin, et al.
Published: (2026)
by: Tang, Jimin, et al.
Published: (2026)
4D3R: Motion-Aware Neural Reconstruction and Rendering of Dynamic Scenes from Monocular Videos
by: Guo, Mengqi, et al.
Published: (2025)
by: Guo, Mengqi, et al.
Published: (2025)
EdgeVidSum: Real-Time Personalized Video Summarization at the Edge
by: Mujtaba, Ghulam, et al.
Published: (2025)
by: Mujtaba, Ghulam, et al.
Published: (2025)
VidLBEval: Benchmarking and Mitigating Language Bias in Video-Involved LVLMs
by: Yang, Yiming, et al.
Published: (2025)
by: Yang, Yiming, et al.
Published: (2025)
Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free Simulation
by: Chen, Chuhao, et al.
Published: (2025)
by: Chen, Chuhao, et al.
Published: (2025)
VidFormer: A novel end-to-end framework fused by 3DCNN and Transformer for Video-based Remote Physiological Measurement
by: Li, Jiachen, et al.
Published: (2025)
by: Li, Jiachen, et al.
Published: (2025)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
by: Zheng, Sixiao, et al.
Published: (2025)
by: Zheng, Sixiao, et al.
Published: (2025)
SceneLLM: Implicit Language Reasoning in LLM for Dynamic Scene Graph Generation
by: Zhang, Hang, et al.
Published: (2024)
by: Zhang, Hang, et al.
Published: (2024)
CVBench: Benchmarking Cross-Video Synergies for Complex Multimodal Reasoning
by: Zhu, Nannan, et al.
Published: (2025)
by: Zhu, Nannan, et al.
Published: (2025)
VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
by: Liang, Baoyu, et al.
Published: (2025)
by: Liang, Baoyu, et al.
Published: (2025)
Vid-SME: Membership Inference Attacks against Large Video Understanding Models
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
Similar Items
-
3DGS_LSR:Large_Scale Relocation for Autonomous Driving Based on 3D Gaussian Splatting
by: Lu, Haitao, et al.
Published: (2025) -
Sky-GVIO: an enhanced GNSS/INS/Vision navigation with FCN-based sky-segmentation in urban canyon
by: Wang, Jingrong, et al.
Published: (2024) -
CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
by: Li, Jingyao, et al.
Published: (2025) -
SafeVid: Toward Safety Aligned Video Large Multimodal Models
by: Wang, Yixu, et al.
Published: (2025) -
HiVid: LLM-Guided Video Saliency For Content-Aware VOD And Live Streaming
by: Chen, Jiahui, et al.
Published: (2026)