Semantic and Expressive Variation in Image Captions Across Languages
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Andre, Santy, Sebastin, Hwang, Jena D., Zhang, Amy X., Krishna, Ranjay |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Incentives Backfire, Data Stops Being Human
by: Santy, Sebastin, et al.
Published: (2025)
by: Santy, Sebastin, et al.
Published: (2025)
BLIP: Facilitating the Exploration of Undesirable Consequences of Digital Technologies
by: Pang, Rock Yuren, et al.
Published: (2024)
by: Pang, Rock Yuren, et al.
Published: (2024)
Language Models as Critical Thinking Tools: A Case Study of Philosophers
by: Ye, Andre, et al.
Published: (2024)
by: Ye, Andre, et al.
Published: (2024)
Investigating Disability Representations in Text-to-Image Models
by: Tian, Yang, et al.
Published: (2026)
by: Tian, Yang, et al.
Published: (2026)
Agonistic Image Generation: Unsettling the Hegemony of Intention
by: Shaw, Andrew, et al.
Published: (2025)
by: Shaw, Andrew, et al.
Published: (2025)
Confidence Contours: Uncertainty-Aware Annotation for Medical Semantic Segmentation
by: Ye, Andre, et al.
Published: (2023)
by: Ye, Andre, et al.
Published: (2023)
Learning Multimodal Cues of Children's Uncertainty
by: Cheng, Qi, et al.
Published: (2024)
by: Cheng, Qi, et al.
Published: (2024)
WAXAL-NET: Finetuned Edge ASR Across 19 African Languages
by: Olufemi, Victor Tolulope, et al.
Published: (2026)
by: Olufemi, Victor Tolulope, et al.
Published: (2026)
Vision-Language Models Suppress Female Representations Under Ambiguous Input
by: Marin-Llobet, Arnau, et al.
Published: (2026)
by: Marin-Llobet, Arnau, et al.
Published: (2026)
Signformer is all you need: Towards Edge AI for Sign Language
by: Yang, Eta
Published: (2024)
by: Yang, Eta
Published: (2024)
A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation
by: Betala, Siddharth, et al.
Published: (2025)
by: Betala, Siddharth, et al.
Published: (2025)
Human-Centred Evaluation of Text-to-Image Generation Models for Self-expression of Mental Distress: A Dataset Based on GPT-4o
by: He, Sui, et al.
Published: (2025)
by: He, Sui, et al.
Published: (2025)
DepMamba: Progressive Fusion Mamba for Multimodal Depression Detection
by: Ye, Jiaxin, et al.
Published: (2024)
by: Ye, Jiaxin, et al.
Published: (2024)
CAF-Mamba: Mamba-Based Cross-Modal Adaptive Attention Fusion for Multimodal Depression Detection
by: Zhou, Bowen, et al.
Published: (2026)
by: Zhou, Bowen, et al.
Published: (2026)
Towards Geographic Inclusion in the Evaluation of Text-to-Image Models
by: Hall, Melissa, et al.
Published: (2024)
by: Hall, Melissa, et al.
Published: (2024)
Intuitions of Machine Learning Researchers about Transfer Learning for Medical Image Classification
by: Lu, Yucheng, et al.
Published: (2025)
by: Lu, Yucheng, et al.
Published: (2025)
Cooperative Speech, Semantic Competence, and AI
by: Almotahari, Mahrad
Published: (2025)
by: Almotahari, Mahrad
Published: (2025)
Morae: Proactively Pausing UI Agents for User Choices
by: Peng, Yi-Hao, et al.
Published: (2025)
by: Peng, Yi-Hao, et al.
Published: (2025)
Real Time Captioning of Sign Language Gestures in Video Meetings
by: Mukherjee, Sharanya, et al.
Published: (2025)
by: Mukherjee, Sharanya, et al.
Published: (2025)
"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
by: Garg, Kapil, et al.
Published: (2025)
by: Garg, Kapil, et al.
Published: (2025)
Navigating the Conceptual Multiverse
by: Ye, Andre, et al.
Published: (2026)
by: Ye, Andre, et al.
Published: (2026)
What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metric
by: Kerkouri, Mohamed Amine, et al.
Published: (2026)
by: Kerkouri, Mohamed Amine, et al.
Published: (2026)
BTPD: A Multilingual Hand-curated Dataset of Bengali Transnational Political Discourse Across Online Communities
by: Das, Dipto, et al.
Published: (2025)
by: Das, Dipto, et al.
Published: (2025)
Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning
by: Huang, Ting-Hao 'Kenneth', et al.
Published: (2025)
by: Huang, Ting-Hao 'Kenneth', et al.
Published: (2025)
How Can Large Language Models Enable Better Socially Assistive Human-Robot Interaction: A Brief Survey
by: Shi, Zhonghao, et al.
Published: (2024)
by: Shi, Zhonghao, et al.
Published: (2024)
When Algorithms Meet Artists: Semantic Compression of Artists' Concerns in the Public AI-Art Debate
by: Mukherjee-Gandhi, Ariya, et al.
Published: (2025)
by: Mukherjee-Gandhi, Ariya, et al.
Published: (2025)
A Review on Large Language Models for Visual Analytics
by: Agarwal, Navya Sonal, et al.
Published: (2025)
by: Agarwal, Navya Sonal, et al.
Published: (2025)
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
by: Subedi, Krishna
Published: (2025)
by: Subedi, Krishna
Published: (2025)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
by: Verma, Arnav, et al.
Published: (2025)
by: Verma, Arnav, et al.
Published: (2025)
SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches
by: Lin, Haichuan, et al.
Published: (2025)
by: Lin, Haichuan, et al.
Published: (2025)
Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty
by: Zhou, Kaitlyn, et al.
Published: (2024)
by: Zhou, Kaitlyn, et al.
Published: (2024)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
Impacts of Anthropomorphizing Large Language Models in Learning Environments
by: Schaaff, Kristina, et al.
Published: (2024)
by: Schaaff, Kristina, et al.
Published: (2024)
A Call to Arms: AI Should be Critical for Social Media Analysis of Conflict Zones
by: Abedin, Afia, et al.
Published: (2023)
by: Abedin, Afia, et al.
Published: (2023)
Improved Digital Therapy for Developmental Pediatrics Using Domain-Specific Artificial Intelligence: Machine Learning Study
by: Washington, Peter, et al.
Published: (2020)
by: Washington, Peter, et al.
Published: (2020)
A Comparison of Human and Machine Learning Errors in Face Recognition
by: Estévez-Almenzar, Marina, et al.
Published: (2025)
by: Estévez-Almenzar, Marina, et al.
Published: (2025)
Classification of the lunar surface pattern by AI architectures: Does AI see a rabbit in the Moon?
by: Shoji, Daigo
Published: (2023)
by: Shoji, Daigo
Published: (2023)
AIDEN: Design and Pilot Study of an AI Assistant for the Visually Impaired
by: Marquez-Carpintero, Luis, et al.
Published: (2025)
by: Marquez-Carpintero, Luis, et al.
Published: (2025)
Beyond Questionnaires: Video Analysis for Social Anxiety Detection
by: Sahu, Nilesh Kumar, et al.
Published: (2024)
by: Sahu, Nilesh Kumar, et al.
Published: (2024)
The Cadaver in the Machine: The Social Practices of Measurement and Validation in Motion Capture Technology
by: Harvey, Emma, et al.
Published: (2024)
by: Harvey, Emma, et al.
Published: (2024)
Similar Items
-
When Incentives Backfire, Data Stops Being Human
by: Santy, Sebastin, et al.
Published: (2025) -
BLIP: Facilitating the Exploration of Undesirable Consequences of Digital Technologies
by: Pang, Rock Yuren, et al.
Published: (2024) -
Language Models as Critical Thinking Tools: A Case Study of Philosophers
by: Ye, Andre, et al.
Published: (2024) -
Investigating Disability Representations in Text-to-Image Models
by: Tian, Yang, et al.
Published: (2026) -
Agonistic Image Generation: Unsettling the Hegemony of Intention
by: Shaw, Andrew, et al.
Published: (2025)