Measuring How (Not Just Whether) VLMs Build Common Ground
Fuente:
arXiv
Saved in:
| Main Authors: | Imai, Saki, İnan, Mert, Sicilia, Anthony, Alikhani, Malihe |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SiLVERScore: Semantically-Aware Embeddings for Sign Language Generation Evaluation
by: Imai, Saki, et al.
Published: (2025)
by: Imai, Saki, et al.
Published: (2025)
Accounting for Sycophancy in Language Model Uncertainty Estimation
by: Sicilia, Anthony, et al.
Published: (2024)
by: Sicilia, Anthony, et al.
Published: (2024)
Generating Signed Language Instructions in Large-Scale Dialogue Systems
by: İnan, Mert, et al.
Published: (2024)
by: İnan, Mert, et al.
Published: (2024)
Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others in Conversation Forecasting
by: Sicilia, Anthony, et al.
Published: (2024)
by: Sicilia, Anthony, et al.
Published: (2024)
Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors
by: Sicilia, Anthony, et al.
Published: (2024)
by: Sicilia, Anthony, et al.
Published: (2024)
MixDPO: Modeling Preference Strength for Pluralistic Alignment
by: Imai, Saki, et al.
Published: (2026)
by: Imai, Saki, et al.
Published: (2026)
Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generation
by: İnan, Mert, et al.
Published: (2025)
by: İnan, Mert, et al.
Published: (2025)
HumBEL: A Human-in-the-Loop Approach for Evaluating Demographic Factors of Language Models in Human-Machine Conversations
by: Sicilia, Anthony, et al.
Published: (2023)
by: Sicilia, Anthony, et al.
Published: (2023)
BASIL: Bayesian Assessment of Sycophancy in LLMs
by: Atwell, Katherine, et al.
Published: (2025)
by: Atwell, Katherine, et al.
Published: (2025)
Modeling Intensification for Sign Language Generation: A Computational Approach
by: İnan, Mert, et al.
Published: (2022)
by: İnan, Mert, et al.
Published: (2022)
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
by: Sicilia, Anthony, et al.
Published: (2024)
by: Sicilia, Anthony, et al.
Published: (2024)
How Pragmatics Shape Articulation: A Computational Case Study in STEM ASL Discourse
by: Imai, Saki, et al.
Published: (2025)
by: Imai, Saki, et al.
Published: (2025)
Contextual ASR Error Handling with LLMs Augmentation for Goal-Oriented Conversational AI
by: Asano, Yuya, et al.
Published: (2025)
by: Asano, Yuya, et al.
Published: (2025)
An Active Learning Framework for Inclusive Generation by Large Language Models
by: Hassan, Sabit, et al.
Published: (2024)
by: Hassan, Sabit, et al.
Published: (2024)
Active Learning for Robust and Representative LLM Generation in Safety-Critical Scenarios
by: Hassan, Sabit, et al.
Published: (2024)
by: Hassan, Sabit, et al.
Published: (2024)
Including Facial Expressions in Contextual Embeddings for Sign Language Generation
by: Viegas, Carla, et al.
Published: (2022)
by: Viegas, Carla, et al.
Published: (2022)
Learning to Generate Context-Sensitive Backchannel Smiles for Embodied AI Agents with Applications in Mental Health Dialogues
by: Bilalpur, Maneesh, et al.
Published: (2024)
by: Bilalpur, Maneesh, et al.
Published: (2024)
Better Slow than Sorry: Introducing Positive Friction for Reliable Dialogue Systems
by: İnan, Mert, et al.
Published: (2025)
by: İnan, Mert, et al.
Published: (2025)
Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences
by: Ashkinaze, Joshua, et al.
Published: (2025)
by: Ashkinaze, Joshua, et al.
Published: (2025)
Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
by: Zhang, Jiayi, et al.
Published: (2025)
by: Zhang, Jiayi, et al.
Published: (2025)
Change My View? The Dynamics of Persuasion and Polarization in Online Discourse
by: Freeborn, David, et al.
Published: (2026)
by: Freeborn, David, et al.
Published: (2026)
"Nothing about us without us": Perspectives of Global Deaf and Hard-of-hearing Community Members on Sign Language Technologies
by: Atwell, Katherine, et al.
Published: (2025)
by: Atwell, Katherine, et al.
Published: (2025)
WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language Models
by: Ning, Kangyun, et al.
Published: (2024)
by: Ning, Kangyun, et al.
Published: (2024)
Frame of Reference: Addressing the Challenges of Common Ground Representation in Situational Dialogs
by: Mohapatra, Biswesh, et al.
Published: (2026)
by: Mohapatra, Biswesh, et al.
Published: (2026)
Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning
by: Xu, Ningning, et al.
Published: (2025)
by: Xu, Ningning, et al.
Published: (2025)
Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation
by: Parikh, Aditya, et al.
Published: (2026)
by: Parikh, Aditya, et al.
Published: (2026)
Distributed Partial Information Puzzles: Examining Common Ground Construction Under Epistemic Asymmetry
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
DAGverse: Building Document-Grounded Semantic DAGs from Scientific Papers
by: Wan, Shu, et al.
Published: (2026)
by: Wan, Shu, et al.
Published: (2026)
How to Understand Named Entities: Using Common Sense for News Captioning
by: Xu, Ning, et al.
Published: (2024)
by: Xu, Ning, et al.
Published: (2024)
How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?
by: Sil, Pritam, et al.
Published: (2026)
by: Sil, Pritam, et al.
Published: (2026)
Large Language Models Implicitly Learn to See and Hear Just By Reading
by: Verma, Prateek, et al.
Published: (2025)
by: Verma, Prateek, et al.
Published: (2025)
Retrieve Only Relevant Tables Whether Few or Many: Adaptive Table Retrieval Method
by: Kim, Taehee, et al.
Published: (2026)
by: Kim, Taehee, et al.
Published: (2026)
Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance
by: Chen, Jingyi, et al.
Published: (2025)
by: Chen, Jingyi, et al.
Published: (2025)
CEGI: Measuring the trade-off between efficiency and carbon emissions for SLMs and VLMs
by: Kumar, Abhas, et al.
Published: (2024)
by: Kumar, Abhas, et al.
Published: (2024)
Using Machine Mental Imagery for Representing Common Ground in Situated Dialogue
by: Mohapatra, Biswesh, et al.
Published: (2026)
by: Mohapatra, Biswesh, et al.
Published: (2026)
Human-centered explanation does not fit all: The interplay of sociotechnical, cognitive, and individual factors in the effect AI explanations in algorithmic decision-making
by: Ahn, Yongsu, et al.
Published: (2025)
by: Ahn, Yongsu, et al.
Published: (2025)
LLMs are Not Just Next Token Predictors
by: Downes, Stephen M., et al.
Published: (2024)
by: Downes, Stephen M., et al.
Published: (2024)
How Well Do Large Language Models Truly Ground?
by: Lee, Hyunji, et al.
Published: (2023)
by: Lee, Hyunji, et al.
Published: (2023)
Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning
by: Sun, Yiqun, et al.
Published: (2025)
by: Sun, Yiqun, et al.
Published: (2025)
Improved GUI Grounding via Iterative Narrowing
by: Nguyen, Anthony
Published: (2024)
by: Nguyen, Anthony
Published: (2024)
Similar Items
-
SiLVERScore: Semantically-Aware Embeddings for Sign Language Generation Evaluation
by: Imai, Saki, et al.
Published: (2025) -
Accounting for Sycophancy in Language Model Uncertainty Estimation
by: Sicilia, Anthony, et al.
Published: (2024) -
Generating Signed Language Instructions in Large-Scale Dialogue Systems
by: İnan, Mert, et al.
Published: (2024) -
Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others in Conversation Forecasting
by: Sicilia, Anthony, et al.
Published: (2024) -
Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors
by: Sicilia, Anthony, et al.
Published: (2024)