A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Betala, Siddharth, Raj, Kushan, Betala, Vipul, Saswade, Rohan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Brotherhood at WMT 2024: Leveraging LLM-Generated Contextual Conversations for Cross-Lingual Image Captioning
von: Betala, Siddharth, et al.
Veröffentlicht: (2024)
von: Betala, Siddharth, et al.
Veröffentlicht: (2024)
MEGAnno+: A Human-LLM Collaborative Annotation System
von: Kim, Hannah, et al.
Veröffentlicht: (2024)
von: Kim, Hannah, et al.
Veröffentlicht: (2024)
Machine Translation in the Wild: User Reaction to Xiaohongshu's Built-In Translation Feature
von: He, Sui
Veröffentlicht: (2026)
von: He, Sui
Veröffentlicht: (2026)
A Picture Is Worth a Thousand Words: Exploring Diagram and Video-Based OOP Exercises to Counter LLM Over-Reliance
von: Cipriano, Bruno Pereira, et al.
Veröffentlicht: (2024)
von: Cipriano, Bruno Pereira, et al.
Veröffentlicht: (2024)
Context-Aware Monolingual Human Evaluation of Machine Translation
von: Picinini, Silvio, et al.
Veröffentlicht: (2025)
von: Picinini, Silvio, et al.
Veröffentlicht: (2025)
Audio-Based Crowd-Sourced Evaluation of Machine Translation Quality
von: Haq, Sami Ul, et al.
Veröffentlicht: (2025)
von: Haq, Sami Ul, et al.
Veröffentlicht: (2025)
A Picture is Worth a Thousand Prompts? Efficacy of Iterative Human-Driven Prompt Refinement in Image Regeneration Tasks
von: Trinh, Khoi, et al.
Veröffentlicht: (2025)
von: Trinh, Khoi, et al.
Veröffentlicht: (2025)
Enhancing Public Speaking Skills in Engineering Students Through AI
von: Harsh, Amol, et al.
Veröffentlicht: (2025)
von: Harsh, Amol, et al.
Veröffentlicht: (2025)
Introducing Quality Estimation to Machine Translation Post-editing Workflow: An Empirical Study on Its Usefulness
von: Liu, Siqi, et al.
Veröffentlicht: (2025)
von: Liu, Siqi, et al.
Veröffentlicht: (2025)
When Large Language Models are Reliable for Judging Empathic Communication
von: Kumar, Aakriti, et al.
Veröffentlicht: (2025)
von: Kumar, Aakriti, et al.
Veröffentlicht: (2025)
Effects of Collaboration on the Performance of Interactive Theme Discovery Systems
von: Chen, Alvin Po-Chun, et al.
Veröffentlicht: (2024)
von: Chen, Alvin Po-Chun, et al.
Veröffentlicht: (2024)
Translation Analytics for Freelancers II: Benchmarking Local LLMs for Confidential Translation Workflows
von: Balashov, Yuri, et al.
Veröffentlicht: (2026)
von: Balashov, Yuri, et al.
Veröffentlicht: (2026)
Understanding How Paper Writers Use AI-Generated Captions in Figure Caption Writing
von: Yin, Ho, et al.
Veröffentlicht: (2025)
von: Yin, Ho, et al.
Veröffentlicht: (2025)
Captioning Visualizations with Large Language Models (CVLLM): A Tutorial
von: Carenini, Giuseppe, et al.
Veröffentlicht: (2024)
von: Carenini, Giuseppe, et al.
Veröffentlicht: (2024)
Prompting ChatGPT for Translation: A Comparative Analysis of Translation Brief and Persona Prompts
von: He, Sui
Veröffentlicht: (2024)
von: He, Sui
Veröffentlicht: (2024)
Towards Multimodal Social Conversations with Robots: Using Vision-Language Models
von: Janssens, Ruben, et al.
Veröffentlicht: (2025)
von: Janssens, Ruben, et al.
Veröffentlicht: (2025)
AI vs. Human Judgment of Content Moderation: LLM-as-a-Judge and Ethics-Based Response Refusals
von: Pasch, Stefan
Veröffentlicht: (2025)
von: Pasch, Stefan
Veröffentlicht: (2025)
Emojinize: Enriching Any Text with Emoji Translations
von: Klein, Lars Henning, et al.
Veröffentlicht: (2024)
von: Klein, Lars Henning, et al.
Veröffentlicht: (2024)
Pearmut: Human Evaluation of Translation Made Trivial
von: Zouhar, Vilém, et al.
Veröffentlicht: (2026)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2026)
From Scratch to Fine-Tuned: A Comparative Study of Transformer Training Strategies for Legal Machine Translation
von: Barman, Amit, et al.
Veröffentlicht: (2025)
von: Barman, Amit, et al.
Veröffentlicht: (2025)
Media of Langue: The Interface for Exploring Word Translation Network/Space
von: Muramoto, Goki, et al.
Veröffentlicht: (2023)
von: Muramoto, Goki, et al.
Veröffentlicht: (2023)
E-THER: A Multimodal Dataset for Empathic AI -- Towards Emotional Mismatch Awareness
von: Tahir, Sharjeel, et al.
Veröffentlicht: (2025)
von: Tahir, Sharjeel, et al.
Veröffentlicht: (2025)
ReactGenie: A Development Framework for Complex Multimodal Interactions Using Large Language Models
von: Yang, Jackie Junrui, et al.
Veröffentlicht: (2023)
von: Yang, Jackie Junrui, et al.
Veröffentlicht: (2023)
Repairs in a Block World: A New Benchmark for Handling User Corrections with Multi-Modal Language Models
von: Chiyah-Garcia, Javier, et al.
Veröffentlicht: (2024)
von: Chiyah-Garcia, Javier, et al.
Veröffentlicht: (2024)
Lost Before Translation: Social Information Transmission and Survival in AI-AI Communication
von: Ghafouri, Bijean, et al.
Veröffentlicht: (2026)
von: Ghafouri, Bijean, et al.
Veröffentlicht: (2026)
MathBuddy: A Multimodal System for Affective Math Tutoring
von: Kar, Debanjana, et al.
Veröffentlicht: (2025)
von: Kar, Debanjana, et al.
Veröffentlicht: (2025)
Machine Learning for Enhancing Deliberation in Online Political Discussions and Participatory Processes: A Survey
von: Behrendt, Maike, et al.
Veröffentlicht: (2025)
von: Behrendt, Maike, et al.
Veröffentlicht: (2025)
RAGExplorer: A Visual Analytics System for the Comparative Diagnosis of RAG Systems
von: Tian, Haoyu, et al.
Veröffentlicht: (2026)
von: Tian, Haoyu, et al.
Veröffentlicht: (2026)
Combine Virtual Reality and Machine-Learning to Identify the Presence of Dyslexia: A Cross-Linguistic Approach
von: Materazzini, Michele, et al.
Veröffentlicht: (2025)
von: Materazzini, Michele, et al.
Veröffentlicht: (2025)
Guiding Generative Storytelling with Knowledge Graphs
von: Pan, Zhijun, et al.
Veröffentlicht: (2025)
von: Pan, Zhijun, et al.
Veröffentlicht: (2025)
Successfully Guiding Humans with Imperfect Instructions by Highlighting Potential Errors and Suggesting Corrections
von: Zhao, Lingjun, et al.
Veröffentlicht: (2024)
von: Zhao, Lingjun, et al.
Veröffentlicht: (2024)
Questionnaires for Everyone: Streamlining Cross-Cultural Questionnaire Adaptation with GPT-Based Translation Quality Evaluation
von: Haavisto, Otso, et al.
Veröffentlicht: (2024)
von: Haavisto, Otso, et al.
Veröffentlicht: (2024)
Multimodal Dialog Systems with Dual Knowledge-enhanced Generative Pretrained Language Model
von: Chen, Xiaolin, et al.
Veröffentlicht: (2022)
von: Chen, Xiaolin, et al.
Veröffentlicht: (2022)
Exploring Personalized Health Support through Data-Driven, Theory-Guided LLMs: A Case Study in Sleep Health
von: Wang, Xingbo, et al.
Veröffentlicht: (2025)
von: Wang, Xingbo, et al.
Veröffentlicht: (2025)
Unsupervised Word-level Quality Estimation for Machine Translation Through the Lens of Annotators (Dis)agreement
von: Sarti, Gabriele, et al.
Veröffentlicht: (2025)
von: Sarti, Gabriele, et al.
Veröffentlicht: (2025)
FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems
von: Liao, Borui, et al.
Veröffentlicht: (2025)
von: Liao, Borui, et al.
Veröffentlicht: (2025)
AwareLLM: A Proactive Multimodal Ecosystem for Personalized Human-AI Collaboration to Enhance Productivity
von: Rao, Amog, et al.
Veröffentlicht: (2026)
von: Rao, Amog, et al.
Veröffentlicht: (2026)
EmphasisChecker: A Tool for Guiding Chart and Caption Emphasis
von: Kim, Dae Hyun, et al.
Veröffentlicht: (2023)
von: Kim, Dae Hyun, et al.
Veröffentlicht: (2023)
ACE: A LLM-based Negotiation Coaching System
von: Shea, Ryan, et al.
Veröffentlicht: (2024)
von: Shea, Ryan, et al.
Veröffentlicht: (2024)
Facilitating Self-Guided Mental Health Interventions Through Human-Language Model Interaction: A Case Study of Cognitive Restructuring
von: Sharma, Ashish, et al.
Veröffentlicht: (2023)
von: Sharma, Ashish, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Brotherhood at WMT 2024: Leveraging LLM-Generated Contextual Conversations for Cross-Lingual Image Captioning
von: Betala, Siddharth, et al.
Veröffentlicht: (2024) -
MEGAnno+: A Human-LLM Collaborative Annotation System
von: Kim, Hannah, et al.
Veröffentlicht: (2024) -
Machine Translation in the Wild: User Reaction to Xiaohongshu's Built-In Translation Feature
von: He, Sui
Veröffentlicht: (2026) -
A Picture Is Worth a Thousand Words: Exploring Diagram and Video-Based OOP Exercises to Counter LLM Over-Reliance
von: Cipriano, Bruno Pereira, et al.
Veröffentlicht: (2024) -
Context-Aware Monolingual Human Evaluation of Machine Translation
von: Picinini, Silvio, et al.
Veröffentlicht: (2025)