VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Qiang, Gao, Xinyuan, Dong, SongLin, Han, Jizhou, Li, Jiangyang, He, Yuhang, Gong, Yihong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unleashing the Potential of All Test Samples: Mean-Shift Guided Test-Time Adaptation
by: Han, Jizhou, et al.
Published: (2025)
by: Han, Jizhou, et al.
Published: (2025)
Hallucination Localization in Video Captioning
by: Nakada, Shota, et al.
Published: (2025)
by: Nakada, Shota, et al.
Published: (2025)
Cap2Sum: Learning to Summarize Videos by Generating Captions
by: Zhao, Cairong, et al.
Published: (2024)
by: Zhao, Cairong, et al.
Published: (2024)
Consistent Supervised-Unsupervised Alignment for Generalized Category Discovery
by: Han, Jizhou, et al.
Published: (2025)
by: Han, Jizhou, et al.
Published: (2025)
DualCP: Rehearsal-Free Domain-Incremental Learning via Dual-Level Concept Prototype
by: Wang, Qiang, et al.
Published: (2025)
by: Wang, Qiang, et al.
Published: (2025)
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
by: Sun, Zeyi, et al.
Published: (2025)
by: Sun, Zeyi, et al.
Published: (2025)
GOAL: Geometrically Optimal Alignment for Continual Generalized Category Discovery
by: Han, Jizhou, et al.
Published: (2026)
by: Han, Jizhou, et al.
Published: (2026)
Learning Like Humans: Analogical Concept Learning for Generalized Category Discovery
by: Han, Jizhou, et al.
Published: (2026)
by: Han, Jizhou, et al.
Published: (2026)
Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions
by: Shi, Junchang, et al.
Published: (2025)
by: Shi, Junchang, et al.
Published: (2025)
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
Mining the Social Fabric: Unveiling Communities for Fake News Detection in Short Videos
by: Gong, Haisong, et al.
Published: (2025)
by: Gong, Haisong, et al.
Published: (2025)
Beyond CLIP Generalization: Against Forward&Backward Forgetting Adapter for Continual Learning of Vision-Language Models
by: Dong, Songlin, et al.
Published: (2025)
by: Dong, Songlin, et al.
Published: (2025)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
by: Wang, Xinran, et al.
Published: (2026)
by: Wang, Xinran, et al.
Published: (2026)
Co-Director: Agentic Generative Video Storytelling
by: Song, Yale, et al.
Published: (2026)
by: Song, Yale, et al.
Published: (2026)
PolySmart @ TRECVid 2024 Video Captioning (VTT)
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
Enhancing Neural Adaptive Wireless Video Streaming via Lower-Layer Information Exposure and Online Tuning
by: Zhao, Lingzhi, et al.
Published: (2025)
by: Zhao, Lingzhi, et al.
Published: (2025)
MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation
by: Zhao, Yuan, et al.
Published: (2026)
by: Zhao, Yuan, et al.
Published: (2026)
VARFVV: View-Adaptive Real-Time Interactive Free-View Video Streaming with Edge Computing
by: Hu, Qiang, et al.
Published: (2025)
by: Hu, Qiang, et al.
Published: (2025)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
by: Ge, Shiping, et al.
Published: (2024)
by: Ge, Shiping, et al.
Published: (2024)
TimeLogic Challenge @ CVPR 2026: Strong MLLMs Meet Evidence-Seeking Agents for Temporal-Logic Video Question Answering
by: Xu, Zhaoyang, et al.
Published: (2026)
by: Xu, Zhaoyang, et al.
Published: (2026)
ELLMPEG: An Edge-based Agentic LLM Video Processing Tool
by: Azimi, Zoha, et al.
Published: (2026)
by: Azimi, Zoha, et al.
Published: (2026)
When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video Recommendation
by: Chen, Siran, et al.
Published: (2025)
by: Chen, Siran, et al.
Published: (2025)
Video Summarization: Towards Entity-Aware Captions
by: Ayyubi, Hammad A., et al.
Published: (2023)
by: Ayyubi, Hammad A., et al.
Published: (2023)
Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health Understanding
by: Zhou, Zhiyuan, et al.
Published: (2026)
by: Zhou, Zhiyuan, et al.
Published: (2026)
NewsCaption: Named-Entity aware Captioning for Out-of-Context Media
by: Singh, Anurag, et al.
Published: (2024)
by: Singh, Anurag, et al.
Published: (2024)
Multimodal Interaction Modeling via Self-Supervised Multi-Task Learning for Review Helpfulness Prediction
by: Gong, HongLin, et al.
Published: (2024)
by: Gong, HongLin, et al.
Published: (2024)
Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus Adaptation
by: Shi, Zhaofeng, et al.
Published: (2025)
by: Shi, Zhaofeng, et al.
Published: (2025)
Edit As You Wish: Video Caption Editing with Multi-grained User Control
by: Yao, Linli, et al.
Published: (2023)
by: Yao, Linli, et al.
Published: (2023)
Multimodal Representation Learning and Fusion
by: Jin, Qihang, et al.
Published: (2025)
by: Jin, Qihang, et al.
Published: (2025)
A Subjective Quality Evaluation of 3D Mesh with Dynamic Level of Detail in Virtual Reality
by: Nguyen, Duc, et al.
Published: (2024)
by: Nguyen, Duc, et al.
Published: (2024)
Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation
by: He, Liu, et al.
Published: (2024)
by: He, Liu, et al.
Published: (2024)
Personalized Playback Technology: How Short Video Services Create Excellent User Experience
by: Deng, Weihui, et al.
Published: (2024)
by: Deng, Weihui, et al.
Published: (2024)
Generalizing Video DeepFake Detection by Self-generated Audio-Visual Pseudo-Fakes
by: Wei, Zihe, et al.
Published: (2026)
by: Wei, Zihe, et al.
Published: (2026)
VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
by: Chen, Siran, et al.
Published: (2025)
by: Chen, Siran, et al.
Published: (2025)
Learn by Reasoning: Analogical Weight Generation for Few-Shot Class-Incremental Learning
by: Han, Jizhou, et al.
Published: (2025)
by: Han, Jizhou, et al.
Published: (2025)
Boosting Domain Incremental Learning: Selecting the Optimal Parameters is All You Need
by: Wang, Qiang, et al.
Published: (2025)
by: Wang, Qiang, et al.
Published: (2025)
FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries
by: You, Qijie, et al.
Published: (2026)
by: You, Qijie, et al.
Published: (2026)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
by: Song, Zijie, et al.
Published: (2023)
by: Song, Zijie, et al.
Published: (2023)
Tile-Weighted Rate-Distortion Optimized Packet Scheduling for 360$^\circ$ VR Video Streaming
by: Wang, Haopeng, et al.
Published: (2024)
by: Wang, Haopeng, et al.
Published: (2024)
Adaptive Offloading and Enhancement for Low-Light Video Analytics on Mobile Devices
by: He, Yuanyi, et al.
Published: (2024)
by: He, Yuanyi, et al.
Published: (2024)
Similar Items
-
Unleashing the Potential of All Test Samples: Mean-Shift Guided Test-Time Adaptation
by: Han, Jizhou, et al.
Published: (2025) -
Hallucination Localization in Video Captioning
by: Nakada, Shota, et al.
Published: (2025) -
Cap2Sum: Learning to Summarize Videos by Generating Captions
by: Zhao, Cairong, et al.
Published: (2024) -
Consistent Supervised-Unsupervised Alignment for Generalized Category Discovery
by: Han, Jizhou, et al.
Published: (2025) -
DualCP: Rehearsal-Free Domain-Incremental Learning via Dual-Level Concept Prototype
by: Wang, Qiang, et al.
Published: (2025)