Exploiting Auxiliary Caption for Video Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hongxiang, Cao, Meng, Cheng, Xuxin, Zhu, Zhihong, Li, Yaowei, Zou, Yuexian |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game Theory
by: Li, Hongxiang, et al.
Published: (2023)
by: Li, Hongxiang, et al.
Published: (2023)
DisPose: Disentangling Pose Guidance for Controllable Human Image Animation
by: Li, Hongxiang, et al.
Published: (2024)
by: Li, Hongxiang, et al.
Published: (2024)
Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning
by: Yang, Bang, et al.
Published: (2024)
by: Yang, Bang, et al.
Published: (2024)
VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification
by: Zhuang, Xianwei, et al.
Published: (2025)
by: Zhuang, Xianwei, et al.
Published: (2025)
BlobCtrl: Taming Controllable Blob for Element-level Image Editing
by: Li, Yaowei, et al.
Published: (2025)
by: Li, Yaowei, et al.
Published: (2025)
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024)
by: Kazakos, Evangelos, et al.
Published: (2024)
Uncertainty-aware sign language video retrieval with probability distribution modeling
by: Wu, Xuan, et al.
Published: (2024)
by: Wu, Xuan, et al.
Published: (2024)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
Image Conductor: Precision Control for Interactive Video Synthesis
by: Li, Yaowei, et al.
Published: (2024)
by: Li, Yaowei, et al.
Published: (2024)
Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation
by: Tang, Lexiang, et al.
Published: (2025)
by: Tang, Lexiang, et al.
Published: (2025)
Towards Spoken Language Understanding via Multi-level Multi-grained Contrastive Learning
by: Cheng, Xuxin, et al.
Published: (2024)
by: Cheng, Xuxin, et al.
Published: (2024)
Large-scale Pre-training for Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2025)
by: Kazakos, Evangelos, et al.
Published: (2025)
Addressing the ID-Matching Challenge in Long Video Captioning
by: Yang, Zhantao, et al.
Published: (2025)
by: Yang, Zhantao, et al.
Published: (2025)
Improving Deep Representation Learning via Auxiliary Learnable Target Coding
by: Liu, Kangjun, et al.
Published: (2023)
by: Liu, Kangjun, et al.
Published: (2023)
Textual Inversion and Self-supervised Refinement for Radiology Report Generation
by: Luo, Yuanjiang, et al.
Published: (2024)
by: Luo, Yuanjiang, et al.
Published: (2024)
Exploiting Multiple Sequence Lengths in Fast End to End Training for Image Captioning
by: Hu, Jia Cheng, et al.
Published: (2022)
by: Hu, Jia Cheng, et al.
Published: (2022)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
XMeCap: Meme Caption Generation with Sub-Image Adaptability
by: Chen, Yuyan, et al.
Published: (2024)
by: Chen, Yuyan, et al.
Published: (2024)
Dual-path Collaborative Generation Network for Emotional Video Captioning
by: Ye, Cheng, et al.
Published: (2024)
by: Ye, Cheng, et al.
Published: (2024)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
by: Li, Weiming, et al.
Published: (2025)
by: Li, Weiming, et al.
Published: (2025)
VideoGuard: Protecting Video Content from Unauthorized Editing
by: Cao, Junjie, et al.
Published: (2025)
by: Cao, Junjie, et al.
Published: (2025)
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
by: Meng, Desen, et al.
Published: (2025)
by: Meng, Desen, et al.
Published: (2025)
BrushEdit: All-In-One Image Inpainting and Editing
by: Li, Yaowei, et al.
Published: (2024)
by: Li, Yaowei, et al.
Published: (2024)
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
by: Li, Xiangtai, et al.
Published: (2025)
by: Li, Xiangtai, et al.
Published: (2025)
CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture Generation
by: Fang, Fengyi, et al.
Published: (2025)
by: Fang, Fengyi, et al.
Published: (2025)
ZeroNLG: Aligning and Autoencoding Domains for Zero-Shot Multimodal and Multilingual Natural Language Generation
by: Yang, Bang, et al.
Published: (2023)
by: Yang, Bang, et al.
Published: (2023)
MV-CC: Mask Enhanced Video Model for Remote Sensing Change Caption
by: Liu, Ruixun, et al.
Published: (2024)
by: Liu, Ruixun, et al.
Published: (2024)
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
by: Zhang, Shi-Xue, et al.
Published: (2025)
by: Zhang, Shi-Xue, et al.
Published: (2025)
SOVC: Subject-Oriented Video Captioning
by: Teng, Chang, et al.
Published: (2023)
by: Teng, Chang, et al.
Published: (2023)
CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model
by: Yao, Ziyu, et al.
Published: (2025)
by: Yao, Ziyu, et al.
Published: (2025)
Live Video Captioning
by: Blanco-Fernández, Eduardo, et al.
Published: (2024)
by: Blanco-Fernández, Eduardo, et al.
Published: (2024)
Infusing Environmental Captions for Long-Form Video Language Grounding
by: Lee, Hyogun, et al.
Published: (2024)
by: Lee, Hyogun, et al.
Published: (2024)
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning
by: Lu, Yifan, et al.
Published: (2023)
by: Lu, Yifan, et al.
Published: (2023)
SGCap: Decoding Semantic Group for Zero-shot Video Captioning
by: Pan, Zeyu, et al.
Published: (2025)
by: Pan, Zeyu, et al.
Published: (2025)
RefVSR++: Exploiting Reference Inputs for Reference-based Video Super-resolution
by: Zou, Han, et al.
Published: (2023)
by: Zou, Han, et al.
Published: (2023)
ViDiC: Video Difference Captioning
by: Wu, Jiangtao, et al.
Published: (2025)
by: Wu, Jiangtao, et al.
Published: (2025)
Cross-Modal Conditioned Reconstruction for Language-guided Medical Image Segmentation
by: Huang, Xiaoshuang, et al.
Published: (2024)
by: Huang, Xiaoshuang, et al.
Published: (2024)
Spectral Tail Auxiliary Learning for AI-Generated Image Detection
by: Li, Xingyi, et al.
Published: (2026)
by: Li, Xingyi, et al.
Published: (2026)
An End-to-End Framework for Video Multi-Person Pose Estimation
by: Wei, Zhihong
Published: (2025)
by: Wei, Zhihong
Published: (2025)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
Similar Items
-
G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game Theory
by: Li, Hongxiang, et al.
Published: (2023) -
DisPose: Disentangling Pose Guidance for Controllable Human Image Animation
by: Li, Hongxiang, et al.
Published: (2024) -
Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning
by: Yang, Bang, et al.
Published: (2024) -
VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification
by: Zhuang, Xianwei, et al.
Published: (2025) -
BlobCtrl: Taming Controllable Blob for Element-level Image Editing
by: Li, Yaowei, et al.
Published: (2025)