Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark
Fuente:
arXiv
Guardado en:
| Autores principales: | Xi, Zeyu, Shi, Ge, Li, Xuefen, Yan, Junchi, Li, Zun, Wu, Lifang, Liu, Zilin, Wang, Liang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning
por: Xi, Zeyu, et al.
Publicado: (2025)
por: Xi, Zeyu, et al.
Publicado: (2025)
VicKAM: Visual Conceptual Knowledge Guided Action Map for Weakly Supervised Group Activity Recognition
por: Wang, Zhuming, et al.
Publicado: (2025)
por: Wang, Zhuming, et al.
Publicado: (2025)
Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning
por: You, Xiaoxing, et al.
Publicado: (2025)
por: You, Xiaoxing, et al.
Publicado: (2025)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
por: Wu, Peiran, et al.
Publicado: (2025)
por: Wu, Peiran, et al.
Publicado: (2025)
Context-aware Difference Distilling for Multi-change Captioning
por: Tu, Yunbin, et al.
Publicado: (2024)
por: Tu, Yunbin, et al.
Publicado: (2024)
SGCap: Decoding Semantic Group for Zero-shot Video Captioning
por: Pan, Zeyu, et al.
Publicado: (2025)
por: Pan, Zeyu, et al.
Publicado: (2025)
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning
por: Lu, Yifan, et al.
Publicado: (2023)
por: Lu, Yifan, et al.
Publicado: (2023)
Video Summarization: Towards Entity-Aware Captions
por: Ayyubi, Hammad A., et al.
Publicado: (2023)
por: Ayyubi, Hammad A., et al.
Publicado: (2023)
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
por: Hu, Shiyu, et al.
Publicado: (2024)
por: Hu, Shiyu, et al.
Publicado: (2024)
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
por: Zhang, Shi-Xue, et al.
Publicado: (2025)
por: Zhang, Shi-Xue, et al.
Publicado: (2025)
Benchmarking and Improving Detail Image Caption
por: Dong, Hongyuan, et al.
Publicado: (2024)
por: Dong, Hongyuan, et al.
Publicado: (2024)
Caries DETR: Tooth Structure-aware Prior and Lesion-aware Dynamic Loss Refinement for DETR Based Caries Detection
por: Liu, Xuefen, et al.
Publicado: (2026)
por: Liu, Xuefen, et al.
Publicado: (2026)
Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images
por: Lu, Zimao, et al.
Publicado: (2025)
por: Lu, Zimao, et al.
Publicado: (2025)
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
por: Liu, Yanqing, et al.
Publicado: (2024)
por: Liu, Yanqing, et al.
Publicado: (2024)
Structural Pruning via Spatial-aware Information Redundancy for Semantic Segmentation
por: Wu, Dongyue, et al.
Publicado: (2024)
por: Wu, Dongyue, et al.
Publicado: (2024)
DragEntity: Trajectory Guided Video Generation using Entity and Positional Relationships
por: Wan, Zhang, et al.
Publicado: (2024)
por: Wan, Zhang, et al.
Publicado: (2024)
ViDiC: Video Difference Captioning
por: Wu, Jiangtao, et al.
Publicado: (2025)
por: Wu, Jiangtao, et al.
Publicado: (2025)
Exploiting Auxiliary Caption for Video Grounding
por: Li, Hongxiang, et al.
Publicado: (2023)
por: Li, Hongxiang, et al.
Publicado: (2023)
AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
por: Chen, Xinlong, et al.
Publicado: (2025)
por: Chen, Xinlong, et al.
Publicado: (2025)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
por: Li, Kunchang, et al.
Publicado: (2023)
por: Li, Kunchang, et al.
Publicado: (2023)
Streaming Dense Video Captioning
por: Zhou, Xingyi, et al.
Publicado: (2024)
por: Zhou, Xingyi, et al.
Publicado: (2024)
RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
por: Song, Jiahe, et al.
Publicado: (2025)
por: Song, Jiahe, et al.
Publicado: (2025)
OPCap:Object-aware Prompting Captioning
por: Huang, Feiyang
Publicado: (2024)
por: Huang, Feiyang
Publicado: (2024)
SOVC: Subject-Oriented Video Captioning
por: Teng, Chang, et al.
Publicado: (2023)
por: Teng, Chang, et al.
Publicado: (2023)
PolySmart @ TRECVid 2024 Video Captioning (VTT)
por: Wu, Jiaxin, et al.
Publicado: (2024)
por: Wu, Jiaxin, et al.
Publicado: (2024)
SGA-INTERACT: A 3D Skeleton-based Benchmark for Group Activity Understanding in Modern Basketball Tactic
por: Yang, Yuchen, et al.
Publicado: (2025)
por: Yang, Yuchen, et al.
Publicado: (2025)
Envisioning Class Entity Reasoning by Large Language Models for Few-shot Learning
por: Liu, Mushui, et al.
Publicado: (2024)
por: Liu, Mushui, et al.
Publicado: (2024)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
por: Fan, Tiehan, et al.
Publicado: (2024)
por: Fan, Tiehan, et al.
Publicado: (2024)
Describe Anything: Detailed Localized Image and Video Captioning
por: Lian, Long, et al.
Publicado: (2025)
por: Lian, Long, et al.
Publicado: (2025)
Semantic-CC: Boosting Remote Sensing Image Change Captioning via Foundational Knowledge and Semantic Guidance
por: Zhu, Yongshuo, et al.
Publicado: (2024)
por: Zhu, Yongshuo, et al.
Publicado: (2024)
Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model
por: AlJunaid, Reem, et al.
Publicado: (2025)
por: AlJunaid, Reem, et al.
Publicado: (2025)
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
por: Zhang, Xu, et al.
Publicado: (2026)
por: Zhang, Xu, et al.
Publicado: (2026)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
por: Tang, Yunlong, et al.
Publicado: (2025)
por: Tang, Yunlong, et al.
Publicado: (2025)
Technical Report for Soccernet 2023 -- Dense Video Captioning
por: Ruan, Zheng, et al.
Publicado: (2024)
por: Ruan, Zheng, et al.
Publicado: (2024)
Live Video Captioning
por: Blanco-Fernández, Eduardo, et al.
Publicado: (2024)
por: Blanco-Fernández, Eduardo, et al.
Publicado: (2024)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
por: Wu, Shengqiong, et al.
Publicado: (2025)
por: Wu, Shengqiong, et al.
Publicado: (2025)
Distractors-Immune Representation Learning with Cross-modal Contrastive Regularization for Change Captioning
por: Tu, Yunbin, et al.
Publicado: (2024)
por: Tu, Yunbin, et al.
Publicado: (2024)
Addressing the ID-Matching Challenge in Long Video Captioning
por: Yang, Zhantao, et al.
Publicado: (2025)
por: Yang, Zhantao, et al.
Publicado: (2025)
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
por: Chai, Wenhao, et al.
Publicado: (2024)
por: Chai, Wenhao, et al.
Publicado: (2024)
AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation
por: Qiu, Lu, et al.
Publicado: (2025)
por: Qiu, Lu, et al.
Publicado: (2025)
Ejemplares similares
-
Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning
por: Xi, Zeyu, et al.
Publicado: (2025) -
VicKAM: Visual Conceptual Knowledge Guided Action Map for Weakly Supervised Group Activity Recognition
por: Wang, Zhuming, et al.
Publicado: (2025) -
Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning
por: You, Xiaoxing, et al.
Publicado: (2025) -
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
por: Wu, Peiran, et al.
Publicado: (2025) -
Context-aware Difference Distilling for Multi-change Captioning
por: Tu, Yunbin, et al.
Publicado: (2024)