Saved in:
| Main Authors: | Garg, Kapil, Tang, Xinru, Heo, Jimin, Morgan, Dwayne R., Gergle, Darren, Sudderth, Erik B., Piper, Anne Marie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.08917 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VIPaint: Image Inpainting with Pre-Trained Diffusion Models via Variational Inference
by: Agarwal, Sakshi, et al.
Published: (2024)
by: Agarwal, Sakshi, et al.
Published: (2024)
Linguistic Similarity Within Centralized FLOSS Development
by: Gaughan, Matthew, et al.
Published: (2026)
by: Gaughan, Matthew, et al.
Published: (2026)
Reimagining Sign Language Technologies: Analyzing Translation Work of Chinese Deaf Online Content Creators
by: Tang, Xinru, et al.
Published: (2026)
by: Tang, Xinru, et al.
Published: (2026)
Designing for Collective Access: In Search of a Solution to Accessible Communication in a Mixed-Ability Non-Profit
by: Tang, Xinru, et al.
Published: (2026)
by: Tang, Xinru, et al.
Published: (2026)
The Accessibility Paradox: How Blind and Low Vision Employees Experience and Negotiate Accessibility in the Technology Industry
by: Marathe, Aparajita, et al.
Published: (2025)
by: Marathe, Aparajita, et al.
Published: (2025)
Beyond Words: An Experimental Study of Signaling in Crowdfunding
by: Dambanemuya, Henry K., et al.
Published: (2022)
by: Dambanemuya, Henry K., et al.
Published: (2022)
When Testing AI Tests Us: Safeguarding Mental Health on the Digital Frontlines
by: Pendse, Sachin R., et al.
Published: (2025)
by: Pendse, Sachin R., et al.
Published: (2025)
Avoid Wasted Annotation Costs in Open-set Active Learning with Pre-trained Vision-Language Model
by: Heo, Jaehyuk, et al.
Published: (2024)
by: Heo, Jaehyuk, et al.
Published: (2024)
The Agony of Opacity: Foundations for Reflective Interpretability in AI-Mediated Mental Health Support
by: Pendse, Sachin R., et al.
Published: (2025)
by: Pendse, Sachin R., et al.
Published: (2025)
Differentiable and Stable Long-Range Tracking of Multiple Posterior Modes
by: Younis, Ali, et al.
Published: (2024)
by: Younis, Ali, et al.
Published: (2024)
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
CXMArena: Unified Dataset to benchmark performance in realistic CXM Scenarios
by: Garg, Raghav, et al.
Published: (2025)
by: Garg, Raghav, et al.
Published: (2025)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
by: Liu, Zheng, et al.
Published: (2024)
by: Liu, Zheng, et al.
Published: (2024)
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
by: Tang, Changli, et al.
Published: (2025)
by: Tang, Changli, et al.
Published: (2025)
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
by: Zhang, Xinsong, et al.
Published: (2025)
by: Zhang, Xinsong, et al.
Published: (2025)
Understanding How Paper Writers Use AI-Generated Captions in Figure Caption Writing
by: Yin, Ho, et al.
Published: (2025)
by: Yin, Ho, et al.
Published: (2025)
Learning to be Smooth: An End-to-End Differentiable Particle Smoother
by: Younis, Ali, et al.
Published: (2025)
by: Younis, Ali, et al.
Published: (2025)
Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioning
by: Baek, Seung Hyup, et al.
Published: (2026)
by: Baek, Seung Hyup, et al.
Published: (2026)
How Users Experience Closed Captions on Live Television: Quality Metrics Remain a Challenge
by: Chavez, Mariana Arroyo, et al.
Published: (2024)
by: Chavez, Mariana Arroyo, et al.
Published: (2024)
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
The Perils of Chart Deception: How Misleading Visualizations Affect Vision-Language Models
by: Mahbub, Ridwan, et al.
Published: (2025)
by: Mahbub, Ridwan, et al.
Published: (2025)
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization
by: Tang, Changli, et al.
Published: (2024)
by: Tang, Changli, et al.
Published: (2024)
Linear Alignment of Vision-language Models for Image Captioning
by: Paischer, Fabian, et al.
Published: (2023)
by: Paischer, Fabian, et al.
Published: (2023)
How Quality Affects Deep Neural Networks in Fine-Grained Image Classification
by: Smith, Joseph, et al.
Published: (2024)
by: Smith, Joseph, et al.
Published: (2024)
How Does the Spatial Distribution of Pre-training Data Affect Geospatial Foundation Models?
by: Purohit, Mirali, et al.
Published: (2025)
by: Purohit, Mirali, et al.
Published: (2025)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
by: Lavoie, Samuel, et al.
Published: (2024)
by: Lavoie, Samuel, et al.
Published: (2024)
Large-scale Pre-training for Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2025)
by: Kazakos, Evangelos, et al.
Published: (2025)
DreamLIP: Language-Image Pre-training with Long Captions
by: Zheng, Kecheng, et al.
Published: (2024)
by: Zheng, Kecheng, et al.
Published: (2024)
Dynamic Neural Style Transfer for Artistic Image Generation using VGG19
by: Kashyap, Kapil, et al.
Published: (2025)
by: Kashyap, Kapil, et al.
Published: (2025)
Long-term Pre-training for Temporal Action Detection with Transformers
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Imagine How To Change: Explicit Procedure Modeling for Change Captioning
by: Sun, Jiayang, et al.
Published: (2026)
by: Sun, Jiayang, et al.
Published: (2026)
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
by: Byun, Sanghyun, et al.
Published: (2025)
by: Byun, Sanghyun, et al.
Published: (2025)
How people use Copilot for Health
by: Costa-Gomes, Beatriz, et al.
Published: (2026)
by: Costa-Gomes, Beatriz, et al.
Published: (2026)
VIVECaption: A Split Approach to Caption Quality Improvement
by: Ananth, Varun, et al.
Published: (2026)
by: Ananth, Varun, et al.
Published: (2026)
BRACE: A Benchmark for Robust Audio Caption Quality Evaluation
by: Guo, Tianyu, et al.
Published: (2025)
by: Guo, Tianyu, et al.
Published: (2025)
ViTOC: Vision Transformer and Object-aware Captioner
by: Huang, Feiyang
Published: (2024)
by: Huang, Feiyang
Published: (2024)
OmniCaptioner: One Captioner to Rule Them All
by: Lu, Yiting, et al.
Published: (2025)
by: Lu, Yiting, et al.
Published: (2025)
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
by: Qi, Yayun, et al.
Published: (2024)
by: Qi, Yayun, et al.
Published: (2024)
Visually-Aware Context Modeling for News Image Captioning
by: Qu, Tingyu, et al.
Published: (2023)
by: Qu, Tingyu, et al.
Published: (2023)
Carrot and Stick: Inducing Self-Motivation with Positive & Negative Feedback
by: Sohn, Jimin, et al.
Published: (2024)
by: Sohn, Jimin, et al.
Published: (2024)
Similar Items
-
VIPaint: Image Inpainting with Pre-Trained Diffusion Models via Variational Inference
by: Agarwal, Sakshi, et al.
Published: (2024) -
Linguistic Similarity Within Centralized FLOSS Development
by: Gaughan, Matthew, et al.
Published: (2026) -
Reimagining Sign Language Technologies: Analyzing Translation Work of Chinese Deaf Online Content Creators
by: Tang, Xinru, et al.
Published: (2026) -
Designing for Collective Access: In Search of a Solution to Accessible Communication in a Mixed-Ability Non-Profit
by: Tang, Xinru, et al.
Published: (2026) -
The Accessibility Paradox: How Blind and Low Vision Employees Experience and Negotiate Accessibility in the Technology Industry
by: Marathe, Aparajita, et al.
Published: (2025)