SOVC: Subject-Oriented Video Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Teng, Chang, Ma, Yunchuan, Li, Guorong, Qi, Yuankai, Qing, Laiyu, Huang, Qingming |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning
by: Ma, Yunchuan, et al.
Published: (2024)
by: Ma, Yunchuan, et al.
Published: (2024)
Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization
by: Ma, Yunchuan, et al.
Published: (2026)
by: Ma, Yunchuan, et al.
Published: (2026)
Boosting Point-supervised Temporal Action Localization via Text Refinement and Alignment
by: Ma, Yunchuan, et al.
Published: (2026)
by: Ma, Yunchuan, et al.
Published: (2026)
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
by: Tian, Mingkai, et al.
Published: (2025)
by: Tian, Mingkai, et al.
Published: (2025)
SDVPT: Semantic-Driven Visual Prompt Tuning for Open-World Object Counting
by: Zhao, Yiming, et al.
Published: (2025)
by: Zhao, Yiming, et al.
Published: (2025)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
by: Zheng, Zelin, et al.
Published: (2026)
by: Zheng, Zelin, et al.
Published: (2026)
Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization
by: Tao, Zhuo, et al.
Published: (2025)
by: Tao, Zhuo, et al.
Published: (2025)
Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Distractors-Immune Representation Learning with Cross-modal Contrastive Regularization for Change Captioning
by: Tu, Yunbin, et al.
Published: (2024)
by: Tu, Yunbin, et al.
Published: (2024)
Context-aware Difference Distilling for Multi-change Captioning
by: Tu, Yunbin, et al.
Published: (2024)
by: Tu, Yunbin, et al.
Published: (2024)
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning
by: Qiu, Tianheng, et al.
Published: (2025)
by: Qiu, Tianheng, et al.
Published: (2025)
Uncertainty-boosted Robust Video Activity Anticipation
by: Qi, Zhaobo, et al.
Published: (2024)
by: Qi, Zhaobo, et al.
Published: (2024)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
by: Tu, Yunbin, et al.
Published: (2024)
by: Tu, Yunbin, et al.
Published: (2024)
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
by: Xue, Zhucun, et al.
Published: (2025)
by: Xue, Zhucun, et al.
Published: (2025)
PolySmart @ TRECVid 2024 Video Captioning (VTT)
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing
by: Cong, Gaoxiang, et al.
Published: (2026)
by: Cong, Gaoxiang, et al.
Published: (2026)
CANeRV: Content Adaptive Neural Representation for Video Compression
by: Tang, Lv, et al.
Published: (2025)
by: Tang, Lv, et al.
Published: (2025)
A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D Scenes
by: Yu, Ting, et al.
Published: (2024)
by: Yu, Ting, et al.
Published: (2024)
Text-only Synthesis for Image Captioning
by: Zhou, Qing, et al.
Published: (2024)
by: Zhou, Qing, et al.
Published: (2024)
Question-guided Visual Compression with Memory Feedback for Long-Term Video Understanding
by: Yamao, Sosuke, et al.
Published: (2026)
by: Yamao, Sosuke, et al.
Published: (2026)
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
Retrieval-Augmented Egocentric Video Captioning
by: Xu, Jilan, et al.
Published: (2024)
by: Xu, Jilan, et al.
Published: (2024)
Exploiting Auxiliary Caption for Video Grounding
by: Li, Hongxiang, et al.
Published: (2023)
by: Li, Hongxiang, et al.
Published: (2023)
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
by: Zheng, Guangcong, et al.
Published: (2025)
by: Zheng, Guangcong, et al.
Published: (2025)
Live Video Captioning
by: Blanco-Fernández, Eduardo, et al.
Published: (2024)
by: Blanco-Fernández, Eduardo, et al.
Published: (2024)
Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in Video
by: Qi, Zhaobo, et al.
Published: (2024)
by: Qi, Zhaobo, et al.
Published: (2024)
MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer
by: Liu, Penghui, et al.
Published: (2025)
by: Liu, Penghui, et al.
Published: (2025)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
by: Yao, Linli, et al.
Published: (2026)
by: Yao, Linli, et al.
Published: (2026)
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
by: Meng, Desen, et al.
Published: (2025)
by: Meng, Desen, et al.
Published: (2025)
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024)
by: Kazakos, Evangelos, et al.
Published: (2024)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
by: Mkhallati, Hassan, et al.
Published: (2023)
by: Mkhallati, Hassan, et al.
Published: (2023)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
by: Jeon, MinJu, et al.
Published: (2025)
by: Jeon, MinJu, et al.
Published: (2025)
A Flexible Recursive Network for Video Stereo Matching Based on Residual Estimation
by: Zhao, Youchen, et al.
Published: (2024)
by: Zhao, Youchen, et al.
Published: (2024)
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning
by: Tang, Yolo Yunlong, et al.
Published: (2023)
by: Tang, Yolo Yunlong, et al.
Published: (2023)
ViDiC: Video Difference Captioning
by: Wu, Jiangtao, et al.
Published: (2025)
by: Wu, Jiangtao, et al.
Published: (2025)
VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models
by: Huang, Chi-Pin, et al.
Published: (2025)
by: Huang, Chi-Pin, et al.
Published: (2025)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
by: Tang, Yunlong, et al.
Published: (2025)
by: Tang, Yunlong, et al.
Published: (2025)
Similar Items
-
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning
by: Ma, Yunchuan, et al.
Published: (2024) -
Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization
by: Ma, Yunchuan, et al.
Published: (2026) -
Boosting Point-supervised Temporal Action Localization via Text Refinement and Alignment
by: Ma, Yunchuan, et al.
Published: (2026) -
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
by: Tian, Mingkai, et al.
Published: (2025) -
SDVPT: Semantic-Driven Visual Prompt Tuning for Open-World Object Counting
by: Zhao, Yiming, et al.
Published: (2025)