Large VLM-based Stylized Sports Captioning
Fuente:
arXiv
Guardado en:
| Autores principales: | Dhar, Sauptik, Buoncristiani, Nicholas, Anakata, Joe, Zhang, Haoyu, Munson, Michelle |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Breakdance Video classification in the age of Generative AI
por: Dhar, Sauptik, et al.
Publicado: (2025)
por: Dhar, Sauptik, et al.
Publicado: (2025)
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
por: Deria, Ankan, et al.
Publicado: (2025)
por: Deria, Ankan, et al.
Publicado: (2025)
Balanced Image Stylization with Style Matching Score
por: Jiang, Yuxin, et al.
Publicado: (2025)
por: Jiang, Yuxin, et al.
Publicado: (2025)
Stylized Synthetic Augmentation further improves Corruption Robustness
por: Siedel, Georg, et al.
Publicado: (2025)
por: Siedel, Georg, et al.
Publicado: (2025)
CASHG: Context-Aware Stylized Online Handwriting Generation
por: Shin, Jinsu, et al.
Publicado: (2026)
por: Shin, Jinsu, et al.
Publicado: (2026)
TrafficVLM: A Controllable Visual Language Model for Traffic Video Captioning
por: Dinh, Quang Minh, et al.
Publicado: (2024)
por: Dinh, Quang Minh, et al.
Publicado: (2024)
DocVLM: Make Your VLM an Efficient Reader
por: Nacson, Mor Shpigel, et al.
Publicado: (2024)
por: Nacson, Mor Shpigel, et al.
Publicado: (2024)
ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area
por: Li, Junxian, et al.
Publicado: (2024)
por: Li, Junxian, et al.
Publicado: (2024)
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
por: Wu, Zhenkai, et al.
Publicado: (2025)
por: Wu, Zhenkai, et al.
Publicado: (2025)
InsTex: Indoor Scenes Stylized Texture Synthesis
por: Zhang, Yunfan, et al.
Publicado: (2025)
por: Zhang, Yunfan, et al.
Publicado: (2025)
CommonForms: A Large, Diverse Dataset for Form Field Detection
por: Barrow, Joe
Publicado: (2025)
por: Barrow, Joe
Publicado: (2025)
StyDeSty: Min-Max Stylization and Destylization for Single Domain Generalization
por: Liu, Songhua, et al.
Publicado: (2024)
por: Liu, Songhua, et al.
Publicado: (2024)
Pretrained Image-Text Models are Secretly Video Captioners
por: Zhang, Chunhui, et al.
Publicado: (2025)
por: Zhang, Chunhui, et al.
Publicado: (2025)
Modeling Image-Caption Rating from Comparative Judgments
por: Minni, Kezia, et al.
Publicado: (2026)
por: Minni, Kezia, et al.
Publicado: (2026)
RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
por: Hu, Suhang, et al.
Publicado: (2025)
por: Hu, Suhang, et al.
Publicado: (2025)
BendVLM: Test-Time Debiasing of Vision-Language Embeddings
por: Gerych, Walter, et al.
Publicado: (2024)
por: Gerych, Walter, et al.
Publicado: (2024)
From Pixels to Prose: A Large Dataset of Dense Image Captions
por: Singla, Vasu, et al.
Publicado: (2024)
por: Singla, Vasu, et al.
Publicado: (2024)
Suppressing VLM Hallucinations with Spectral Representation Filtering
por: Ali, Ameen, et al.
Publicado: (2025)
por: Ali, Ameen, et al.
Publicado: (2025)
VIBE: Can a VLM Read the Room?
por: Chakraborty, Tania, et al.
Publicado: (2025)
por: Chakraborty, Tania, et al.
Publicado: (2025)
TRISHUL: Towards Region Identification and Screen Hierarchy Understanding for Large VLM based GUI Agents
por: Singh, Kunal, et al.
Publicado: (2025)
por: Singh, Kunal, et al.
Publicado: (2025)
Pixels to Prose: Understanding the art of Image Captioning
por: Singh, Hrishikesh, et al.
Publicado: (2024)
por: Singh, Hrishikesh, et al.
Publicado: (2024)
LILAC: Long-sequence Incremental Low-latency Arbitrary Motion Stylization via Streaming VAE-Diffusion with Causal Decoding
por: Ren, Peng, et al.
Publicado: (2025)
por: Ren, Peng, et al.
Publicado: (2025)
CompCap: Improving Multimodal Large Language Models with Composite Captions
por: Chen, Xiaohui, et al.
Publicado: (2024)
por: Chen, Xiaohui, et al.
Publicado: (2024)
Learning to Rank Caption Chains for Video-Text Alignment
por: Blume, Ansel, et al.
Publicado: (2026)
por: Blume, Ansel, et al.
Publicado: (2026)
Differentially Private Representation Learning via Image Captioning
por: Sander, Tom, et al.
Publicado: (2024)
por: Sander, Tom, et al.
Publicado: (2024)
SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
por: Sivakumar, Anushka, et al.
Publicado: (2025)
por: Sivakumar, Anushka, et al.
Publicado: (2025)
BMIP: Bi-directional Modality Interaction Prompt Learning for VLM
por: Lv, Song-Lin, et al.
Publicado: (2025)
por: Lv, Song-Lin, et al.
Publicado: (2025)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
por: Merchant, Nicholas, et al.
Publicado: (2025)
por: Merchant, Nicholas, et al.
Publicado: (2025)
STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning
por: Gong, Yanpei, et al.
Publicado: (2026)
por: Gong, Yanpei, et al.
Publicado: (2026)
Stylized Structural Patterns for Improved Neural Network Pre-training
por: Salehi, Farnood, et al.
Publicado: (2025)
por: Salehi, Farnood, et al.
Publicado: (2025)
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
por: Xu, Yi, et al.
Publicado: (2024)
por: Xu, Yi, et al.
Publicado: (2024)
ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation
por: Yanuka, Moran, et al.
Publicado: (2024)
por: Yanuka, Moran, et al.
Publicado: (2024)
Infusing Environmental Captions for Long-Form Video Language Grounding
por: Lee, Hyogun, et al.
Publicado: (2024)
por: Lee, Hyogun, et al.
Publicado: (2024)
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2024)
por: Vasu, Pavan Kumar Anasosalu, et al.
Publicado: (2024)
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
por: Lai, Zhengfeng, et al.
Publicado: (2024)
por: Lai, Zhengfeng, et al.
Publicado: (2024)
Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
por: Li, Muyang, et al.
Publicado: (2026)
por: Li, Muyang, et al.
Publicado: (2026)
GraphVLM: Benchmarking Vision Language Models for Multimodal Graph Learning
por: Liu, Jiajin, et al.
Publicado: (2026)
por: Liu, Jiajin, et al.
Publicado: (2026)
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
por: Bao, Chen, et al.
Publicado: (2024)
por: Bao, Chen, et al.
Publicado: (2024)
Time-Archival Camera Virtualization for Sports and Visual Performances
por: Zhang, Yunxiao, et al.
Publicado: (2026)
por: Zhang, Yunxiao, et al.
Publicado: (2026)
Progressive Alignment with VLM-LLM Feature to Augment Defect Classification for the ASE Dataset
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
por: Hsu, Chih-Chung, et al.
Publicado: (2024)
Ejemplares similares
-
Breakdance Video classification in the age of Generative AI
por: Dhar, Sauptik, et al.
Publicado: (2025) -
Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
por: Deria, Ankan, et al.
Publicado: (2025) -
Balanced Image Stylization with Style Matching Score
por: Jiang, Yuxin, et al.
Publicado: (2025) -
Stylized Synthetic Augmentation further improves Corruption Robustness
por: Siedel, Georg, et al.
Publicado: (2025) -
CASHG: Context-Aware Stylized Online Handwriting Generation
por: Shin, Jinsu, et al.
Publicado: (2026)