PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions
Fuente:
arXiv
Saved in:
| Main Authors: | Ananthram, Amith, Stengel-Eskin, Elias, Bradford, Lorena A., Demarest, Julia, Purvis, Adam, Krut, Keith, Stein, Robert, Pantalony, Rina Elster, Bansal, Mohit, McKeown, Kathleen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
See It from My Perspective: How Language Affects Cultural Bias in Image Understanding
by: Ananthram, Amith, et al.
Published: (2024)
by: Ananthram, Amith, et al.
Published: (2024)
Mining Contextualized Visual Associations from Images for Creativity Understanding
by: Sahu, Ananya, et al.
Published: (2025)
by: Sahu, Ananya, et al.
Published: (2025)
Enhancing Multimodal Affective Analysis with Learned Live Comment Features
by: Deng, Zhaoyuan, et al.
Published: (2024)
by: Deng, Zhaoyuan, et al.
Published: (2024)
Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style
by: Limpijankit, Marvin, et al.
Published: (2026)
by: Limpijankit, Marvin, et al.
Published: (2026)
Social Orientation: A New Feature for Dialogue Analysis
by: Morrill, Todd, et al.
Published: (2024)
by: Morrill, Todd, et al.
Published: (2024)
Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
by: Prasad, Archiki, et al.
Published: (2023)
by: Prasad, Archiki, et al.
Published: (2023)
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
by: Patil, Vaidehi, et al.
Published: (2025)
by: Patil, Vaidehi, et al.
Published: (2025)
ReGAL: Refactoring Programs to Discover Generalizable Abstractions
by: Stengel-Eskin, Elias, et al.
Published: (2024)
by: Stengel-Eskin, Elias, et al.
Published: (2024)
LACIE: Listener-Aware Finetuning for Confidence Calibration in Large Language Models
by: Stengel-Eskin, Elias, et al.
Published: (2024)
by: Stengel-Eskin, Elias, et al.
Published: (2024)
The Sum Leaks More Than Its Parts: Compositional Privacy Risks and Mitigations in Multi-Agent Collaboration
by: Patil, Vaidehi, et al.
Published: (2025)
by: Patil, Vaidehi, et al.
Published: (2025)
Teaching Models to Balance Resisting and Accepting Persuasion
by: Stengel-Eskin, Elias, et al.
Published: (2024)
by: Stengel-Eskin, Elias, et al.
Published: (2024)
Language Models Identify Ambiguities and Exploit Loopholes
by: Choi, Jio, et al.
Published: (2025)
by: Choi, Jio, et al.
Published: (2025)
LASeR: Learning to Adaptively Select Reward Models with Multi-Armed Bandits
by: Nguyen, Duy, et al.
Published: (2024)
by: Nguyen, Duy, et al.
Published: (2024)
Soft Self-Consistency Improves Language Model Agents
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
by: Pothiraj, Atin, et al.
Published: (2025)
by: Pothiraj, Atin, et al.
Published: (2025)
RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
by: Niu, Tianyi, et al.
Published: (2025)
by: Niu, Tianyi, et al.
Published: (2025)
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
by: Nguyen, Duy, et al.
Published: (2025)
by: Nguyen, Duy, et al.
Published: (2025)
Are language models rational? The case of coherence norms and belief revision
by: Hofweber, Thomas, et al.
Published: (2024)
by: Hofweber, Thomas, et al.
Published: (2024)
AdaCAD: Adaptively Decoding to Balance Conflicts between Contextual and Parametric Knowledge
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Multi-Attribute Steering of Language Models via Targeted Intervention
by: Nguyen, Duy, et al.
Published: (2025)
by: Nguyen, Duy, et al.
Published: (2025)
DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
by: Khan, Zaid, et al.
Published: (2024)
by: Khan, Zaid, et al.
Published: (2024)
Retrieval-Augmented Generation with Conflicting Evidence
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
by: Wan, David, et al.
Published: (2024)
by: Wan, David, et al.
Published: (2024)
Data Caricatures: On the Representation of African American Language in Pretraining Corpora
by: Deas, Nicholas, et al.
Published: (2025)
by: Deas, Nicholas, et al.
Published: (2025)
The food of trout in New South Wales, 1933–1934
by: McKeown, Keith C.
Published: (1934)
by: McKeown, Keith C.
Published: (1934)
Catalogue of the Cerambycidae (Coleoptera) of Australia
by: McKeown, Keith C.
Published: (1947)
by: McKeown, Keith C.
Published: (1947)
step2point dataset: Detailed shower simulation for data representation studies
by: Zaborowska, Anna, et al.
Published: (2025)
by: Zaborowska, Anna, et al.
Published: (2025)
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
by: Xiao, Hanqi, et al.
Published: (2025)
by: Xiao, Hanqi, et al.
Published: (2025)
Summarization of Opinionated Political Documents with Varied Perspectives
by: Deas, Nicholas, et al.
Published: (2024)
by: Deas, Nicholas, et al.
Published: (2024)
Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions
by: Deas, Nicholas, et al.
Published: (2025)
by: Deas, Nicholas, et al.
Published: (2025)
MAGDi: Structured Distillation of Multi-Agent Interaction Graphs Improves Reasoning in Smaller Language Models
by: Chen, Justin Chih-Yao, et al.
Published: (2024)
by: Chen, Justin Chih-Yao, et al.
Published: (2024)
GenerationPrograms: Fine-grained Attribution with Executable Programs
by: Wan, David, et al.
Published: (2025)
by: Wan, David, et al.
Published: (2025)
Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns
by: Xiao, Hanqi, et al.
Published: (2025)
by: Xiao, Hanqi, et al.
Published: (2025)
Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems
by: Khan, Zaid, et al.
Published: (2025)
by: Khan, Zaid, et al.
Published: (2025)
One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided Exploration
by: Khan, Zaid, et al.
Published: (2025)
by: Khan, Zaid, et al.
Published: (2025)
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
by: Wan, David, et al.
Published: (2025)
by: Wan, David, et al.
Published: (2025)
MAMM-Refine: A Recipe for Improving Faithfulness in Generation with Multi-Agent Collaboration
by: Wan, David, et al.
Published: (2025)
by: Wan, David, et al.
Published: (2025)
Fundamental Problems With Model Editing: How Should Rational Belief Revision Work in LLMs?
by: Hase, Peter, et al.
Published: (2024)
by: Hase, Peter, et al.
Published: (2024)
God's Babies
by: McKeown, John
Published: (2018)
by: McKeown, John
Published: (2018)
The Busemann Process and Steep Highways in Directed First Passage Percolation
by: McKeown, Sam
Published: (2025)
by: McKeown, Sam
Published: (2025)
Similar Items
-
See It from My Perspective: How Language Affects Cultural Bias in Image Understanding
by: Ananthram, Amith, et al.
Published: (2024) -
Mining Contextualized Visual Associations from Images for Creativity Understanding
by: Sahu, Ananya, et al.
Published: (2025) -
Enhancing Multimodal Affective Analysis with Learned Live Comment Features
by: Deng, Zhaoyuan, et al.
Published: (2024) -
Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style
by: Limpijankit, Marvin, et al.
Published: (2026) -
Social Orientation: A New Feature for Dialogue Analysis
by: Morrill, Todd, et al.
Published: (2024)