Large VLM-based Stylized Sports Captioning
Fuente:
arXiv
Salvato in:
| Autori principali: | Dhar, Sauptik, Buoncristiani, Nicholas, Anakata, Joe, Zhang, Haoyu, Munson, Michelle |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Breakdance Video classification in the age of Generative AI
di: Dhar, Sauptik, et al.
Pubblicazione: (2025)
di: Dhar, Sauptik, et al.
Pubblicazione: (2025)
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
di: Deria, Ankan, et al.
Pubblicazione: (2025)
di: Deria, Ankan, et al.
Pubblicazione: (2025)
Balanced Image Stylization with Style Matching Score
di: Jiang, Yuxin, et al.
Pubblicazione: (2025)
di: Jiang, Yuxin, et al.
Pubblicazione: (2025)
Stylized Synthetic Augmentation further improves Corruption Robustness
di: Siedel, Georg, et al.
Pubblicazione: (2025)
di: Siedel, Georg, et al.
Pubblicazione: (2025)
CASHG: Context-Aware Stylized Online Handwriting Generation
di: Shin, Jinsu, et al.
Pubblicazione: (2026)
di: Shin, Jinsu, et al.
Pubblicazione: (2026)
TrafficVLM: A Controllable Visual Language Model for Traffic Video Captioning
di: Dinh, Quang Minh, et al.
Pubblicazione: (2024)
di: Dinh, Quang Minh, et al.
Pubblicazione: (2024)
DocVLM: Make Your VLM an Efficient Reader
di: Nacson, Mor Shpigel, et al.
Pubblicazione: (2024)
di: Nacson, Mor Shpigel, et al.
Pubblicazione: (2024)
ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area
di: Li, Junxian, et al.
Pubblicazione: (2024)
di: Li, Junxian, et al.
Pubblicazione: (2024)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
di: Wu, Zhenkai, et al.
Pubblicazione: (2025)
di: Wu, Zhenkai, et al.
Pubblicazione: (2025)
InsTex: Indoor Scenes Stylized Texture Synthesis
di: Zhang, Yunfan, et al.
Pubblicazione: (2025)
di: Zhang, Yunfan, et al.
Pubblicazione: (2025)
CommonForms: A Large, Diverse Dataset for Form Field Detection
di: Barrow, Joe
Pubblicazione: (2025)
di: Barrow, Joe
Pubblicazione: (2025)
StyDeSty: Min-Max Stylization and Destylization for Single Domain Generalization
di: Liu, Songhua, et al.
Pubblicazione: (2024)
di: Liu, Songhua, et al.
Pubblicazione: (2024)
Pretrained Image-Text Models are Secretly Video Captioners
di: Zhang, Chunhui, et al.
Pubblicazione: (2025)
di: Zhang, Chunhui, et al.
Pubblicazione: (2025)
Modeling Image-Caption Rating from Comparative Judgments
di: Minni, Kezia, et al.
Pubblicazione: (2026)
di: Minni, Kezia, et al.
Pubblicazione: (2026)
RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
di: Hu, Suhang, et al.
Pubblicazione: (2025)
di: Hu, Suhang, et al.
Pubblicazione: (2025)
BendVLM: Test-Time Debiasing of Vision-Language Embeddings
di: Gerych, Walter, et al.
Pubblicazione: (2024)
di: Gerych, Walter, et al.
Pubblicazione: (2024)
From Pixels to Prose: A Large Dataset of Dense Image Captions
di: Singla, Vasu, et al.
Pubblicazione: (2024)
di: Singla, Vasu, et al.
Pubblicazione: (2024)
Suppressing VLM Hallucinations with Spectral Representation Filtering
di: Ali, Ameen, et al.
Pubblicazione: (2025)
di: Ali, Ameen, et al.
Pubblicazione: (2025)
VIBE: Can a VLM Read the Room?
di: Chakraborty, Tania, et al.
Pubblicazione: (2025)
di: Chakraborty, Tania, et al.
Pubblicazione: (2025)
TRISHUL: Towards Region Identification and Screen Hierarchy Understanding for Large VLM based GUI Agents
di: Singh, Kunal, et al.
Pubblicazione: (2025)
di: Singh, Kunal, et al.
Pubblicazione: (2025)
Pixels to Prose: Understanding the art of Image Captioning
di: Singh, Hrishikesh, et al.
Pubblicazione: (2024)
di: Singh, Hrishikesh, et al.
Pubblicazione: (2024)
LILAC: Long-sequence Incremental Low-latency Arbitrary Motion Stylization via Streaming VAE-Diffusion with Causal Decoding
di: Ren, Peng, et al.
Pubblicazione: (2025)
di: Ren, Peng, et al.
Pubblicazione: (2025)
CompCap: Improving Multimodal Large Language Models with Composite Captions
di: Chen, Xiaohui, et al.
Pubblicazione: (2024)
di: Chen, Xiaohui, et al.
Pubblicazione: (2024)
Learning to Rank Caption Chains for Video-Text Alignment
di: Blume, Ansel, et al.
Pubblicazione: (2026)
di: Blume, Ansel, et al.
Pubblicazione: (2026)
Differentially Private Representation Learning via Image Captioning
di: Sander, Tom, et al.
Pubblicazione: (2024)
di: Sander, Tom, et al.
Pubblicazione: (2024)
SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
di: Sivakumar, Anushka, et al.
Pubblicazione: (2025)
di: Sivakumar, Anushka, et al.
Pubblicazione: (2025)
BMIP: Bi-directional Modality Interaction Prompt Learning for VLM
di: Lv, Song-Lin, et al.
Pubblicazione: (2025)
di: Lv, Song-Lin, et al.
Pubblicazione: (2025)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
di: Merchant, Nicholas, et al.
Pubblicazione: (2025)
di: Merchant, Nicholas, et al.
Pubblicazione: (2025)
STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning
di: Gong, Yanpei, et al.
Pubblicazione: (2026)
di: Gong, Yanpei, et al.
Pubblicazione: (2026)
Stylized Structural Patterns for Improved Neural Network Pre-training
di: Salehi, Farnood, et al.
Pubblicazione: (2025)
di: Salehi, Farnood, et al.
Pubblicazione: (2025)
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
di: Xu, Yi, et al.
Pubblicazione: (2024)
di: Xu, Yi, et al.
Pubblicazione: (2024)
ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation
di: Yanuka, Moran, et al.
Pubblicazione: (2024)
di: Yanuka, Moran, et al.
Pubblicazione: (2024)
Infusing Environmental Captions for Long-Form Video Language Grounding
di: Lee, Hyogun, et al.
Pubblicazione: (2024)
di: Lee, Hyogun, et al.
Pubblicazione: (2024)
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
di: Vasu, Pavan Kumar Anasosalu, et al.
Pubblicazione: (2024)
di: Vasu, Pavan Kumar Anasosalu, et al.
Pubblicazione: (2024)
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
di: Lai, Zhengfeng, et al.
Pubblicazione: (2024)
di: Lai, Zhengfeng, et al.
Pubblicazione: (2024)
Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
di: Li, Muyang, et al.
Pubblicazione: (2026)
di: Li, Muyang, et al.
Pubblicazione: (2026)
GraphVLM: Benchmarking Vision Language Models for Multimodal Graph Learning
di: Liu, Jiajin, et al.
Pubblicazione: (2026)
di: Liu, Jiajin, et al.
Pubblicazione: (2026)
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
di: Bao, Chen, et al.
Pubblicazione: (2024)
di: Bao, Chen, et al.
Pubblicazione: (2024)
Time-Archival Camera Virtualization for Sports and Visual Performances
di: Zhang, Yunxiao, et al.
Pubblicazione: (2026)
di: Zhang, Yunxiao, et al.
Pubblicazione: (2026)
Progressive Alignment with VLM-LLM Feature to Augment Defect Classification for the ASE Dataset
di: Hsu, Chih-Chung, et al.
Pubblicazione: (2024)
di: Hsu, Chih-Chung, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Breakdance Video classification in the age of Generative AI
di: Dhar, Sauptik, et al.
Pubblicazione: (2025) -
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
di: Deria, Ankan, et al.
Pubblicazione: (2025) -
Balanced Image Stylization with Style Matching Score
di: Jiang, Yuxin, et al.
Pubblicazione: (2025) -
Stylized Synthetic Augmentation further improves Corruption Robustness
di: Siedel, Georg, et al.
Pubblicazione: (2025) -
CASHG: Context-Aware Stylized Online Handwriting Generation
di: Shin, Jinsu, et al.
Pubblicazione: (2026)