IF-VidCap: Can Video Caption Models Follow Instructions?
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Shihao, Zhang, Yuanxing, Wu, Jiangtao, Lei, Zhide, He, Yiwen, Wen, Runzhe, Liao, Chenxi, Jiang, Chengkang, Ping, An, Gao, Shuo, Wang, Suhan, Bian, Zhaozhou, Zhou, Zijun, Xie, Jingyi, Zhou, Jiayi, Wang, Jing, Yao, Yifan, Xie, Weihao, Tan, Yingshui, Wang, Yanghai, Xie, Qianqian, Zhang, Zhaoxiang, Liu, Jiaheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ViDiC: Video Difference Captioning
by: Wu, Jiangtao, et al.
Published: (2025)
by: Wu, Jiangtao, et al.
Published: (2025)
VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
by: Cao, Zhe, et al.
Published: (2025)
by: Cao, Zhe, et al.
Published: (2025)
MeaCap: Memory-Augmented Zero-shot Image Captioning
by: Zeng, Zequn, et al.
Published: (2024)
by: Zeng, Zequn, et al.
Published: (2024)
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
by: Pan, Yaning, et al.
Published: (2025)
by: Pan, Yaning, et al.
Published: (2025)
Vibe AIGC: A New Paradigm for Content Generation via Agentic Orchestration
by: Liu, Jiaheng, et al.
Published: (2026)
by: Liu, Jiaheng, et al.
Published: (2026)
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
by: Peng, Tianhao, et al.
Published: (2025)
by: Peng, Tianhao, et al.
Published: (2025)
DR$^{3}$-Eval: Towards Realistic and Reproducible Deep Research Evaluation
by: Xie, Qianqian, et al.
Published: (2026)
by: Xie, Qianqian, et al.
Published: (2026)
VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
by: Fan, Tiehan, et al.
Published: (2024)
by: Fan, Tiehan, et al.
Published: (2024)
Vid2Sim: Realistic and Interactive Simulation from Video for Urban Navigation
by: Xie, Ziyang, et al.
Published: (2025)
by: Xie, Ziyang, et al.
Published: (2025)
VidTwin: Video VAE with Decoupled Structure and Dynamics
by: Wang, Yuchi, et al.
Published: (2024)
by: Wang, Yuchi, et al.
Published: (2024)
InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction
by: Wu, Yuhui, et al.
Published: (2025)
by: Wu, Yuhui, et al.
Published: (2025)
FingerCap: Fine-grained Finger-level Hand Motion Captioning
by: Shen, Xin, et al.
Published: (2025)
by: Shen, Xin, et al.
Published: (2025)
Convexity of 2-convex translating and expanding solitons to the mean curvature flow in $\mathbb{R}^{n+1}$
by: Xie, Junming, et al.
Published: (2022)
by: Xie, Junming, et al.
Published: (2022)
CapS-Adapter: Caption-based MultiModal Adapter in Zero-Shot Classification
by: Wang, Qijie, et al.
Published: (2024)
by: Wang, Qijie, et al.
Published: (2024)
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
by: Chai, Wenhao, et al.
Published: (2024)
by: Chai, Wenhao, et al.
Published: (2024)
A Multiplex Approach Against Disturbance Propagation in Nonlinear Networks with Delays
by: Xie, Shihao, et al.
Published: (2022)
by: Xie, Shihao, et al.
Published: (2022)
CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models
by: Zhang, Alexander, et al.
Published: (2025)
by: Zhang, Alexander, et al.
Published: (2025)
SnapCap: Efficient Snapshot Compressive Video Captioning
by: Sun, Jianqiao, et al.
Published: (2024)
by: Sun, Jianqiao, et al.
Published: (2024)
Human or LLM as Standardized Patients? A Comparative Study for Medical Education
by: Zhang, Bingquan, et al.
Published: (2025)
by: Zhang, Bingquan, et al.
Published: (2025)
Designing Unimodular Waveforms for MIMO Radar Based on Manifold Optimization Method
by: Zhao, Xuyang, et al.
Published: (2024)
by: Zhao, Xuyang, et al.
Published: (2024)
Cap2Sum: Learning to Summarize Videos by Generating Captions
by: Zhao, Cairong, et al.
Published: (2024)
by: Zhao, Cairong, et al.
Published: (2024)
Unitary modules over the gap-$p$ Virasoro algebras
by: Xu, Chengkang
Published: (2026)
by: Xu, Chengkang
Published: (2026)
$δ$-biderivations of Virasoro related algebras
by: Xu, Chengkang
Published: (2026)
by: Xu, Chengkang
Published: (2026)
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
by: Liu, Yanqing, et al.
Published: (2024)
by: Liu, Yanqing, et al.
Published: (2024)
Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories
by: Wang, Jiaming, et al.
Published: (2026)
by: Wang, Jiaming, et al.
Published: (2026)
Logic-based switching finite-time stabilization with applications in mechanical systems
by: Zheng, Shiqi, et al.
Published: (2020)
by: Zheng, Shiqi, et al.
Published: (2020)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
by: Jeon, MinJu, et al.
Published: (2025)
by: Jeon, MinJu, et al.
Published: (2025)
Unlocking Hidden Potential in Point Cloud Networks with Attention-Guided Grouping-Feature Coordination
by: Xie, Shangzhuo, et al.
Published: (2025)
by: Xie, Shangzhuo, et al.
Published: (2025)
A comparison of citation-based clustering and topic modeling for science mapping
by: Xie, Qianqian, et al.
Published: (2023)
by: Xie, Qianqian, et al.
Published: (2023)
VoCap: Video Object Captioning and Segmentation from Any Prompt
by: Uijlings, Jasper, et al.
Published: (2025)
by: Uijlings, Jasper, et al.
Published: (2025)
Orion-RAG: Path-Aligned Hybrid Retrieval for Graphless Data
by: Chen, Zhen, et al.
Published: (2026)
by: Chen, Zhen, et al.
Published: (2026)
OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward
by: Zhong, Chunlin, et al.
Published: (2025)
by: Zhong, Chunlin, et al.
Published: (2025)
Boundedness of multilinear Littlewood--Paley operators with convolution type kernels on products of BMO spaces
by: Zhang, Runzhe, et al.
Published: (2025)
by: Zhang, Runzhe, et al.
Published: (2025)
Measuring Investor Learning in Private Markets: A Sequential LLM-Bayesian Analysis of Expert Network Calls
by: Chai, Yidong, et al.
Published: (2025)
by: Chai, Yidong, et al.
Published: (2025)
AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
by: Ren, Yiming, et al.
Published: (2025)
by: Ren, Yiming, et al.
Published: (2025)
Decoding Emotional Trajectories: A Temporal-Semantic Network Approach for Latent Depression Assessment in Social Media
by: Kuang, Junwei, et al.
Published: (2023)
by: Kuang, Junwei, et al.
Published: (2023)
ControlCap: Controllable Region-level Captioning
by: Zhao, Yuzhong, et al.
Published: (2024)
by: Zhao, Yuzhong, et al.
Published: (2024)
Change Captioning in Remote Sensing: Evolution to SAT-Cap -- A Single-Stage Transformer Approach
by: Wang, Yuduo, et al.
Published: (2025)
by: Wang, Yuduo, et al.
Published: (2025)
Similar Items
-
ViDiC: Video Difference Captioning
by: Wu, Jiangtao, et al.
Published: (2025) -
VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
by: Chen, Xinlong, et al.
Published: (2025) -
T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
by: Cao, Zhe, et al.
Published: (2025) -
MeaCap: Memory-Augmented Zero-shot Image Captioning
by: Zeng, Zequn, et al.
Published: (2024) -
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
by: Pan, Yaning, et al.
Published: (2025)