Video Understanding with Large Language Models: A Survey
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Yolo Y., Bi, Jing, Xu, Siting, Song, Luchuan, Liang, Susan, Wang, Teng, Zhang, Daoan, An, Jie, Lin, Jingyang, Zhu, Rongyi, Vosoughi, Ali, Huang, Chao, Zhang, Zeliang, Liu, Pinxin, Feng, Mingqian, Zheng, Feng, Zhang, Jianguo, Luo, Ping, Luo, Jiebo, Xu, Chenliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
GaussianStyle: Gaussian Head Avatar via StyleGAN
von: Liu, Pinxin, et al.
Veröffentlicht: (2024)
von: Liu, Pinxin, et al.
Veröffentlicht: (2024)
Forward Learning for Gradient-based Black-box Saliency Map Generation
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
Generative AI for Cel-Animation: A Survey
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
Will the Inclusion of Generated Data Amplify Bias Across Generations in Future Image Classification Models?
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?
von: Feng, Mingqian, et al.
Veröffentlicht: (2024)
von: Feng, Mingqian, et al.
Veröffentlicht: (2024)
Discover and Mitigate Multiple Biased Subgroups in Image Classifiers
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
von: Tang, Yunlong, et al.
Veröffentlicht: (2025)
von: Tang, Yunlong, et al.
Veröffentlicht: (2025)
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2024)
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2024)
Why Reasoning Matters? A Survey of Advancements in Multimodal Reasoning (v1)
von: Bi, Jing, et al.
Veröffentlicht: (2025)
von: Bi, Jing, et al.
Veröffentlicht: (2025)
TDMM-LM: Bridging Facial Understanding and Animation via Language Models
von: Song, Luchuan, et al.
Veröffentlicht: (2026)
von: Song, Luchuan, et al.
Veröffentlicht: (2026)
Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?
von: Liang, Susan, et al.
Veröffentlicht: (2026)
von: Liang, Susan, et al.
Veröffentlicht: (2026)
V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning
von: Hua, Hang, et al.
Veröffentlicht: (2024)
von: Hua, Hang, et al.
Veröffentlicht: (2024)
Learning to Transform Dynamically for Better Adversarial Transferability
von: Zhu, Rongyi, et al.
Veröffentlicht: (2024)
von: Zhu, Rongyi, et al.
Veröffentlicht: (2024)
Adaptive Super Resolution For One-Shot Talking-Head Generation
von: Song, Luchuan, et al.
Veröffentlicht: (2024)
von: Song, Luchuan, et al.
Veröffentlicht: (2024)
Tri$^{2}$-plane: Thinking Head Avatar via Feature Pyramid
von: Song, Luchuan, et al.
Veröffentlicht: (2024)
von: Song, Luchuan, et al.
Veröffentlicht: (2024)
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
von: Tang, Yolo Y., et al.
Veröffentlicht: (2024)
von: Tang, Yolo Y., et al.
Veröffentlicht: (2024)
GestureLSM: Latent Shortcut based Co-Speech Gesture Generation with Spatial-Temporal Modeling
von: Liu, Pinxin, et al.
Veröffentlicht: (2025)
von: Liu, Pinxin, et al.
Veröffentlicht: (2025)
TextToon: Real-Time Text Toonify Head Avatar from Single Video
von: Song, Luchuan, et al.
Veröffentlicht: (2024)
von: Song, Luchuan, et al.
Veröffentlicht: (2024)
Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning
von: Tan, Zhangyun, et al.
Veröffentlicht: (2026)
von: Tan, Zhangyun, et al.
Veröffentlicht: (2026)
$I^2G$: Generating Instructional Illustrations via Text-Conditioned Diffusion
von: Bi, Jing, et al.
Veröffentlicht: (2025)
von: Bi, Jing, et al.
Veröffentlicht: (2025)
LaunchpadGPT: Language Model as Music Visualization Designer on Launchpad
von: Xu, Siting, et al.
Veröffentlicht: (2023)
von: Xu, Siting, et al.
Veröffentlicht: (2023)
EAGLE: Egocentric AGgregated Language-video Engine
von: Bi, Jing, et al.
Veröffentlicht: (2024)
von: Bi, Jing, et al.
Veröffentlicht: (2024)
Intentional Gesture: Deliver Your Intentions with Gestures for Speech
von: Liu, Pinxin, et al.
Veröffentlicht: (2025)
von: Liu, Pinxin, et al.
Veröffentlicht: (2025)
When to Think and When to Look: Uncertainty-Guided Lookback
von: Bi, Jing, et al.
Veröffentlicht: (2025)
von: Bi, Jing, et al.
Veröffentlicht: (2025)
Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives
von: Zhang, Zeliang, et al.
Veröffentlicht: (2025)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2025)
Learning Brain Tumor Representation in 3D High-Resolution MR Images via Interpretable State Space Models
von: Hu, Qingqiao, et al.
Veröffentlicht: (2024)
von: Hu, Qingqiao, et al.
Veröffentlicht: (2024)
Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
WorldGenBench: A World-Knowledge-Integrated Benchmark for Reasoning-Driven Text-to-Image Generation
von: Zhang, Daoan, et al.
Veröffentlicht: (2025)
von: Zhang, Daoan, et al.
Veröffentlicht: (2025)
Multi-modal Segment Assemblage Network for Ad Video Editing with Importance-Coherence Reward
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2022)
von: Tang, Yolo Yunlong, et al.
Veröffentlicht: (2022)
Sphinx: Benchmarking and Modeling for LLM-Driven Pull Request Review
von: Zhang, Daoan, et al.
Veröffentlicht: (2026)
von: Zhang, Daoan, et al.
Veröffentlicht: (2026)
Forward Learning with Differential Privacy
von: Feng, Mingqian, et al.
Veröffentlicht: (2025)
von: Feng, Mingqian, et al.
Veröffentlicht: (2025)
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
Can Sound Replace Vision in LLaVA With Token Substitution?
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
A Versatile Multimodal Agent for Multimedia Content Generation
von: Zhang, Daoan, et al.
Veröffentlicht: (2026)
von: Zhang, Daoan, et al.
Veröffentlicht: (2026)
LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
Harnessing the Computation Redundancy in ViTs to Boost Adversarial Transferability
von: Liu, Jiani, et al.
Veröffentlicht: (2025)
von: Liu, Jiani, et al.
Veröffentlicht: (2025)
VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity
von: Bi, Jing, et al.
Veröffentlicht: (2025)
von: Bi, Jing, et al.
Veröffentlicht: (2025)
Why Instruction-Based Unlearning Fails in Diffusion Models?
von: Zhang, Zeliang, et al.
Veröffentlicht: (2026)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025) -
GaussianStyle: Gaussian Head Avatar via StyleGAN
von: Liu, Pinxin, et al.
Veröffentlicht: (2024) -
Forward Learning for Gradient-based Black-box Saliency Map Generation
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024) -
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024) -
Generative AI for Cel-Animation: A Survey
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)