Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Xi, Zeyu, Sun, Haoying, Wu, Yaofei, Yan, Junchi, Zhang, Haoran, Wu, Lifang, Wang, Liang, Chen, Changwen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark
by: Xi, Zeyu, et al.
Published: (2024)
by: Xi, Zeyu, et al.
Published: (2024)
Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models
by: Jiang, Songtao, et al.
Published: (2024)
by: Jiang, Songtao, et al.
Published: (2024)
Text Data-Centric Image Captioning with Interactive Prompts
by: Wang, Yiyu, et al.
Published: (2024)
by: Wang, Yiyu, et al.
Published: (2024)
Are Made and Missed Different? An analysis of Field Goal Attempts of Professional Basketball Players via Depth Based Testing Procedure
by: Qi, Kai, et al.
Published: (2024)
by: Qi, Kai, et al.
Published: (2024)
VicKAM: Visual Conceptual Knowledge Guided Action Map for Weakly Supervised Group Activity Recognition
by: Wang, Zhuming, et al.
Published: (2025)
by: Wang, Zhuming, et al.
Published: (2025)
CAMD: Coverage-Aware Multimodal Decoding for Efficient Reasoning of Multimodal Large Language Models
by: Guo, Huijie, et al.
Published: (2026)
by: Guo, Huijie, et al.
Published: (2026)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
by: Tang, Yunlong, et al.
Published: (2025)
by: Tang, Yunlong, et al.
Published: (2025)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models
by: Xing, Wenbin, et al.
Published: (2026)
by: Xing, Wenbin, et al.
Published: (2026)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
by: Qiu, Zongyang, et al.
Published: (2025)
by: Qiu, Zongyang, et al.
Published: (2025)
Offensive Lineup Analysis in Basketball with Clustering Players Based on Shooting Style and Offensive Role
by: Yamada, Kazuhiro, et al.
Published: (2024)
by: Yamada, Kazuhiro, et al.
Published: (2024)
Study of Mental Health among Male and Female Basketball Players of Nagpur
by: Chaudhary, Sanjay
Published: (2026)
by: Chaudhary, Sanjay
Published: (2026)
VideoDreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning on Language-Video Foundation Models
by: Chen, Hong, et al.
Published: (2023)
by: Chen, Hong, et al.
Published: (2023)
MoPE: Mixture of Prompt Experts for Parameter-Efficient and Scalable Multimodal Fusion
by: Jiang, Ruixiang, et al.
Published: (2024)
by: Jiang, Ruixiang, et al.
Published: (2024)
National Players vs. Foreign Players: what distinguishes their game performances? A study in the Portuguese Basketball League
by: Eduardo Guimarãe
Published: (2018)
by: Eduardo Guimarãe
Published: (2018)
TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors
by: Cheng, Wei-Yuan, et al.
Published: (2026)
by: Cheng, Wei-Yuan, et al.
Published: (2026)
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search
by: Yu, Linhao, et al.
Published: (2025)
by: Yu, Linhao, et al.
Published: (2025)
DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models
by: Chen, Xinlong, et al.
Published: (2026)
by: Chen, Xinlong, et al.
Published: (2026)
PresentAgent: Multimodal Agent for Presentation Video Generation
by: Shi, Jingwei, et al.
Published: (2025)
by: Shi, Jingwei, et al.
Published: (2025)
PlayBest: Professional Basketball Player Behavior Synthesis via Planning with Diffusion
by: Chen, Xiusi, et al.
Published: (2023)
by: Chen, Xiusi, et al.
Published: (2023)
Patch as Node: Human-Centric Graph Representation Learning for Multimodal Action Recognition
by: Liang, Zeyu, et al.
Published: (2025)
by: Liang, Zeyu, et al.
Published: (2025)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
Are Large Vision Language Models Good Game Players?
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
BayesPrompt: Prompting Large-Scale Pre-Trained Language Models on Few-shot Inference via Debiased Domain Abstraction
by: Li, Jiangmeng, et al.
Published: (2024)
by: Li, Jiangmeng, et al.
Published: (2024)
Effectively Enhancing Vision Language Large Models by Prompt Augmentation and Caption Utilization
by: Zhao, Minyi, et al.
Published: (2024)
by: Zhao, Minyi, et al.
Published: (2024)
Youth Basketball Players' Awareness and Experiences of Sports‐Related Traumatic Dental Injuries and Mouthguards in Turkey: A Cross‐Sectional Study
by: Hamide Comert, et al.
Published: (2025)
by: Hamide Comert, et al.
Published: (2025)
A New Framework to Estimate Return on Investment for Player Salaries in the National Basketball Association
by: Lautier, Jackson P.
Published: (2023)
by: Lautier, Jackson P.
Published: (2023)
A New Framework to Estimate Return on Investment for Player Salaries in the National Basketball Association
by: Jackson P. Lautier
Published: (2025)
by: Jackson P. Lautier
Published: (2025)
RETRACTION: Construction and Simulation of a Multiattribute Training Data Mining Model for Basketball Players Based on Big Data
by: Wireless Communications and Mobile Computing
Published: (2025)
by: Wireless Communications and Mobile Computing
Published: (2025)
ForgerySleuth: Empowering Multimodal Large Language Models for Image Manipulation Detection
by: Sun, Zhihao, et al.
Published: (2024)
by: Sun, Zhihao, et al.
Published: (2024)
Can AI Prompt Humans? Multimodal Agents Prompt Players' Game Actions and Show Consequences to Raise Sustainability Awareness
by: Zhang, Qinshi, et al.
Published: (2024)
by: Zhang, Qinshi, et al.
Published: (2024)
ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning
by: Yashima, Daichi, et al.
Published: (2026)
by: Yashima, Daichi, et al.
Published: (2026)
Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization
by: Wu, Jiulong, et al.
Published: (2025)
by: Wu, Jiulong, et al.
Published: (2025)
GeoMix: Towards Geometry-Aware Data Augmentation
by: Zhao, Wentao, et al.
Published: (2024)
by: Zhao, Wentao, et al.
Published: (2024)
RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
by: Song, Jiahe, et al.
Published: (2025)
by: Song, Jiahe, et al.
Published: (2025)
yley123/MatSynth-Captions-Text-Prompts-for-Material-Videos: MatSynth-Captions-Text-Prompts-for-Material-Videos
by: BOWEN
Published: (2025)
by: BOWEN
Published: (2025)
Bootstrapping Heterogeneous Graph Representation Learning via Large Language Models: A Generalized Approach
by: Gao, Hang, et al.
Published: (2024)
by: Gao, Hang, et al.
Published: (2024)
CompCap: Improving Multimodal Large Language Models with Composite Captions
by: Chen, Xiaohui, et al.
Published: (2024)
by: Chen, Xiaohui, et al.
Published: (2024)
Effect of Contextual Interference and Differential Learning on Motor Skill Development and Motivation in Novice Basketball Players
by: Ghazal Shamshiri, et al.
Published: (2025)
by: Ghazal Shamshiri, et al.
Published: (2025)
Similar Items
-
Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark
by: Xi, Zeyu, et al.
Published: (2024) -
Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models
by: Jiang, Songtao, et al.
Published: (2024) -
Text Data-Centric Image Captioning with Interactive Prompts
by: Wang, Yiyu, et al.
Published: (2024) -
Are Made and Missed Different? An analysis of Field Goal Attempts of Professional Basketball Players via Depth Based Testing Procedure
by: Qi, Kai, et al.
Published: (2024) -
VicKAM: Visual Conceptual Knowledge Guided Action Map for Weakly Supervised Group Activity Recognition
by: Wang, Zhuming, et al.
Published: (2025)