Addressing the ID-Matching Challenge in Long Video Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Zhantao, Wang, Huangji, Feng, Ruili, Zhang, Han, Hu, Yuting, Zhu, Shangwen, Li, Junyan, Liu, Yu, Cheng, Fan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs
by: Yang, Zhantao, et al.
Published: (2024)
by: Yang, Zhantao, et al.
Published: (2024)
MAMBO-G: Magnitude-Aware Mitigation for Boosted Guidance
by: Zhu, Shangwen, et al.
Published: (2025)
by: Zhu, Shangwen, et al.
Published: (2025)
Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
by: Zhu, Shangwen, et al.
Published: (2026)
by: Zhu, Shangwen, et al.
Published: (2026)
Lipschitz Singularities in Diffusion Models
by: Yang, Zhantao, et al.
Published: (2023)
by: Yang, Zhantao, et al.
Published: (2023)
TIE: Time Interval Encoding for Video Generation over Events
by: Shu, Zhilei, et al.
Published: (2026)
by: Shu, Zhilei, et al.
Published: (2026)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
Video ReCap: Recursive Captioning of Hour-Long Videos
by: Islam, Md Mohaiminul, et al.
Published: (2024)
by: Islam, Md Mohaiminul, et al.
Published: (2024)
Fine-Grained Captioning of Long Videos through Scene Graph Consolidation
by: Chu, Sanghyeok, et al.
Published: (2025)
by: Chu, Sanghyeok, et al.
Published: (2025)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
RAIN: Real-time Animation of Infinite Video Stream
by: Shu, Zhilei, et al.
Published: (2024)
by: Shu, Zhilei, et al.
Published: (2024)
Exploiting Auxiliary Caption for Video Grounding
by: Li, Hongxiang, et al.
Published: (2023)
by: Li, Hongxiang, et al.
Published: (2023)
Retrieval-Augmented Egocentric Video Captioning
by: Xu, Jilan, et al.
Published: (2024)
by: Xu, Jilan, et al.
Published: (2024)
MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
by: Ju, Xuan, et al.
Published: (2024)
by: Ju, Xuan, et al.
Published: (2024)
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
by: Zheng, Guangcong, et al.
Published: (2025)
by: Zheng, Guangcong, et al.
Published: (2025)
Progress-Aware Video Frame Captioning
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
Synthetic-To-Real Video Person Re-ID
by: Zhang, Xiangqun, et al.
Published: (2024)
by: Zhang, Xiangqun, et al.
Published: (2024)
GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration
by: Xu, Wan, et al.
Published: (2025)
by: Xu, Wan, et al.
Published: (2025)
Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
by: Tang, Yunlong, et al.
Published: (2025)
by: Tang, Yunlong, et al.
Published: (2025)
ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
by: Yang, Timing, et al.
Published: (2025)
by: Yang, Timing, et al.
Published: (2025)
Live Video Captioning
by: Blanco-Fernández, Eduardo, et al.
Published: (2024)
by: Blanco-Fernández, Eduardo, et al.
Published: (2024)
RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection
by: Chen, Fangyi, et al.
Published: (2024)
by: Chen, Fangyi, et al.
Published: (2024)
Implicit and Explicit Commonsense for Multi-sentence Video Captioning
by: Chou, Shih-Han, et al.
Published: (2023)
by: Chou, Shih-Han, et al.
Published: (2023)
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
by: Xue, Zhucun, et al.
Published: (2025)
by: Xue, Zhucun, et al.
Published: (2025)
Towards Long-Form Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2026)
by: Gu, Xin, et al.
Published: (2026)
VReID-XFD: Video-based Person Re-identification at Extreme Far Distance Challenge Results
by: Hambarde, Kailash A., et al.
Published: (2026)
by: Hambarde, Kailash A., et al.
Published: (2026)
CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning
by: Lin, Zihan, et al.
Published: (2026)
by: Lin, Zihan, et al.
Published: (2026)
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
by: Zhang, Xu, et al.
Published: (2026)
by: Zhang, Xu, et al.
Published: (2026)
Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning
by: Cheng, Qinchuan, et al.
Published: (2026)
by: Cheng, Qinchuan, et al.
Published: (2026)
Technical Report for Soccernet 2023 -- Dense Video Captioning
by: Ruan, Zheng, et al.
Published: (2024)
by: Ruan, Zheng, et al.
Published: (2024)
Segment and Caption Anything
by: Huang, Xiaoke, et al.
Published: (2023)
by: Huang, Xiaoke, et al.
Published: (2023)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
by: Perrett, Toby, et al.
Published: (2024)
by: Perrett, Toby, et al.
Published: (2024)
AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-Identification
by: Nguyen, Huy, et al.
Published: (2025)
by: Nguyen, Huy, et al.
Published: (2025)
FantasyID: Face Knowledge Enhanced ID-Preserving Video Generation
by: Zhang, Yunpeng, et al.
Published: (2025)
by: Zhang, Yunpeng, et al.
Published: (2025)
LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos
by: Lin, Chin-Yang, et al.
Published: (2025)
by: Lin, Chin-Yang, et al.
Published: (2025)
Video Summarization: Towards Entity-Aware Captions
by: Ayyubi, Hammad A., et al.
Published: (2023)
by: Ayyubi, Hammad A., et al.
Published: (2023)
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024)
by: Kazakos, Evangelos, et al.
Published: (2024)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
ID-Eraser: Proactive Defense Against Face Swapping via Identity Perturbation
by: Luo, Junyan, et al.
Published: (2026)
by: Luo, Junyan, et al.
Published: (2026)
ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
by: Chen, Lin, et al.
Published: (2024)
by: Chen, Lin, et al.
Published: (2024)
Similar Items
-
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs
by: Yang, Zhantao, et al.
Published: (2024) -
MAMBO-G: Magnitude-Aware Mitigation for Boosted Guidance
by: Zhu, Shangwen, et al.
Published: (2025) -
Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
by: Zhu, Shangwen, et al.
Published: (2026) -
Lipschitz Singularities in Diffusion Models
by: Yang, Zhantao, et al.
Published: (2023) -
TIE: Time Interval Encoding for Video Generation over Events
by: Shu, Zhilei, et al.
Published: (2026)