AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Xinlong, Ding, Yue, Lin, Weihong, Hua, Jingyun, Yao, Linli, Shi, Yang, Li, Bozhou, Zhang, Yuanxing, Liu, Qiang, Wan, Pengfei, Wang, Liang, Tan, Tieniu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models
by: Chen, Xinlong, et al.
Published: (2026)
by: Chen, Xinlong, et al.
Published: (2026)
VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
by: Yao, Linli, et al.
Published: (2026)
by: Yao, Linli, et al.
Published: (2026)
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models
by: Ding, Yue, et al.
Published: (2026)
by: Ding, Yue, et al.
Published: (2026)
The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers
by: Li, Bozhou, et al.
Published: (2026)
by: Li, Bozhou, et al.
Published: (2026)
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
by: Wang, Yuchi, et al.
Published: (2025)
by: Wang, Yuchi, et al.
Published: (2025)
Generative Frame Sampler for Long Video Understanding
by: Yao, Linli, et al.
Published: (2025)
by: Yao, Linli, et al.
Published: (2025)
Edit As You Wish: Video Caption Editing with Multi-grained User Control
by: Yao, Linli, et al.
Published: (2023)
by: Yao, Linli, et al.
Published: (2023)
False Starts: The Segregated Lives of PreschoolersBy CaseyStockstill, New York: New York University Press, 2023. 232 pp. USD $28 (paperback). ISBN : 978‐1‐47‐981504‐3
by: Yuanxing Tan
Published: (2025)
by: Yuanxing Tan
Published: (2025)
ViDiC: Video Difference Captioning
by: Wu, Jiangtao, et al.
Published: (2025)
by: Wu, Jiangtao, et al.
Published: (2025)
On higher regulators of Picard modular surfaces
by: Shi, Linli
Published: (2025)
by: Shi, Linli
Published: (2025)
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
by: Tang, Yuqi, et al.
Published: (2026)
by: Tang, Yuqi, et al.
Published: (2026)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
by: Xu, Le, et al.
Published: (2025)
by: Xu, Le, et al.
Published: (2025)
UBiSS: A Unified Framework for Bimodal Semantic Summarization of Videos
by: Mei, Yuting, et al.
Published: (2024)
by: Mei, Yuting, et al.
Published: (2024)
Temporal Reasoning Transfer from Text to Video
by: Li, Lei, et al.
Published: (2024)
by: Li, Lei, et al.
Published: (2024)
MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction
by: Sun, Xiaokun, et al.
Published: (2026)
by: Sun, Xiaokun, et al.
Published: (2026)
LTCA: Long-range Temporal Context Attention for Referring Video Object Segmentation
by: Yan, Cilin, et al.
Published: (2025)
by: Yan, Cilin, et al.
Published: (2025)
IF-VidCap: Can Video Caption Models Follow Instructions?
by: Li, Shihao, et al.
Published: (2025)
by: Li, Shihao, et al.
Published: (2025)
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
by: Liu, Tengfei, et al.
Published: (2026)
by: Liu, Tengfei, et al.
Published: (2026)
Vibe AIGC: A New Paradigm for Content Generation via Agentic Orchestration
by: Liu, Jiaheng, et al.
Published: (2026)
by: Liu, Jiaheng, et al.
Published: (2026)
Audio-Sync Video Generation with Multi-Stream Temporal Control
by: Weng, Shuchen, et al.
Published: (2025)
by: Weng, Shuchen, et al.
Published: (2025)
VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
by: Wang, Qunzhong, et al.
Published: (2025)
by: Wang, Qunzhong, et al.
Published: (2025)
Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding
by: Ding, Ning, et al.
Published: (2025)
by: Ding, Ning, et al.
Published: (2025)
A Comprehensive Survey on Test-Time Adaptation under Distribution Shifts
by: Liang, Jian, et al.
Published: (2023)
by: Liang, Jian, et al.
Published: (2023)
LocCa: Visual Pretraining with Location-aware Captioners
by: Wan, Bo, et al.
Published: (2024)
by: Wan, Bo, et al.
Published: (2024)
Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
by: Wu, Junfei, et al.
Published: (2024)
by: Wu, Junfei, et al.
Published: (2024)
NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative
by: Nadeem, Asmar, et al.
Published: (2024)
by: Nadeem, Asmar, et al.
Published: (2024)
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
by: Xu, Yifan, et al.
Published: (2024)
by: Xu, Yifan, et al.
Published: (2024)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
DeltaEdit: Exploring Text-free Training for Text-Driven Image Manipulation
by: Lyu, Yueming, et al.
Published: (2023)
by: Lyu, Yueming, et al.
Published: (2023)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
by: Ge, Shiping, et al.
Published: (2024)
by: Ge, Shiping, et al.
Published: (2024)
Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks
by: Zeng, Bohan, et al.
Published: (2026)
by: Zeng, Bohan, et al.
Published: (2026)
RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
Similar Items
-
DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models
by: Chen, Xinlong, et al.
Published: (2026) -
VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
by: Chen, Xinlong, et al.
Published: (2025) -
VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
by: Chen, Xinlong, et al.
Published: (2025) -
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
by: Yao, Linli, et al.
Published: (2026) -
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
by: Chen, Xinlong, et al.
Published: (2025)