WinoViz: Probing Visual Properties of Objects Under Different States
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jin, Woojeong, Srinivasan, Tejas, Thomason, Jesse, Ren, Xiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adjust for Trust: Mitigating Trust-Induced Inappropriate Reliance on AI Assistance
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2025)
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2025)
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
von: He, Keyu, et al.
Veröffentlicht: (2025)
von: He, Keyu, et al.
Veröffentlicht: (2025)
From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered
von: Devic, Siddartha, et al.
Veröffentlicht: (2025)
von: Devic, Siddartha, et al.
Veröffentlicht: (2025)
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2024)
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2024)
WinoWhat: A Parallel Corpus of Paraphrased WinoGrande Sentences with Common Sense Categorization
von: Gevers, Ine, et al.
Veröffentlicht: (2025)
von: Gevers, Ine, et al.
Veröffentlicht: (2025)
Better Slow than Sorry: Introducing Positive Friction for Reliable Dialogue Systems
von: İnan, Mert, et al.
Veröffentlicht: (2025)
von: İnan, Mert, et al.
Veröffentlicht: (2025)
Large Language Models Do Multi-Label Classification Differently
von: Ma, Marcus, et al.
Veröffentlicht: (2025)
von: Ma, Marcus, et al.
Veröffentlicht: (2025)
LegalViz: Legal Text Visualization by Text To Diagram Generation
von: Onami, Eri, et al.
Veröffentlicht: (2025)
von: Onami, Eri, et al.
Veröffentlicht: (2025)
When Parts Are Greater Than Sums: Individual LLM Components Can Outperform Full Models
von: Chang, Ting-Yun, et al.
Veröffentlicht: (2024)
von: Chang, Ting-Yun, et al.
Veröffentlicht: (2024)
Do Localization Methods Actually Localize Memorized Data in LLMs? A Tale of Two Benchmarks
von: Chang, Ting-Yun, et al.
Veröffentlicht: (2023)
von: Chang, Ting-Yun, et al.
Veröffentlicht: (2023)
Phonological Representation Learning for Isolated Signs Improves Out-of-Vocabulary Generalization
von: Kezar, Lee, et al.
Veröffentlicht: (2025)
von: Kezar, Lee, et al.
Veröffentlicht: (2025)
PSALM-V: Automating Symbolic Planning in Interactive Visual Environments with Large Language Models
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2025)
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2025)
Compare without Despair: Reliable Preference Evaluation with Generation Separability
von: Ghosh, Sayan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sayan, et al.
Veröffentlicht: (2024)
PDDL-Mind: Large Language Models are Capable on Belief Reasoning with Reliable State Tracking
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026)
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026)
WinoPron: Revisiting English Winogender Schemas for Consistency, Coverage, and Grammatical Case
von: Gautam, Vagrant, et al.
Veröffentlicht: (2024)
von: Gautam, Vagrant, et al.
Veröffentlicht: (2024)
Efficient End-to-End Visual Document Understanding with Rationale Distillation
von: Zhu, Wang, et al.
Veröffentlicht: (2023)
von: Zhu, Wang, et al.
Veröffentlicht: (2023)
Words that make SENSE: Sensorimotor Norms in Learned Lexical Token Representations
von: Gupta, Abhinav, et al.
Veröffentlicht: (2026)
von: Gupta, Abhinav, et al.
Veröffentlicht: (2026)
Can VLMs Recall Factual Associations From Visual References?
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2025)
von: Ashok, Dhananjay, et al.
Veröffentlicht: (2025)
Estonian WinoGrande Dataset: Comparative Analysis of LLM Performance on Human and Machine Translation
von: Ojastu, Marii, et al.
Veröffentlicht: (2025)
von: Ojastu, Marii, et al.
Veröffentlicht: (2025)
Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models
von: Felkner, Virginia K., et al.
Veröffentlicht: (2023)
von: Felkner, Virginia K., et al.
Veröffentlicht: (2023)
Iterative Formalization and Planning in Partially Observable Environments
von: Gong, Liancheng, et al.
Veröffentlicht: (2025)
von: Gong, Liancheng, et al.
Veröffentlicht: (2025)
Language Models can Infer Action Semantics for Symbolic Planners from Environment Feedback
von: Zhu, Wang, et al.
Veröffentlicht: (2024)
von: Zhu, Wang, et al.
Veröffentlicht: (2024)
ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language Models
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
Generating Contextually-Relevant Navigation Instructions for Blind and Low Vision People
von: Merchant, Zain, et al.
Veröffentlicht: (2024)
von: Merchant, Zain, et al.
Veröffentlicht: (2024)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
von: Wang, Siting, et al.
Veröffentlicht: (2025)
von: Wang, Siting, et al.
Veröffentlicht: (2025)
TwoStep: Multi-agent Task Planning using Classical Planners and Large Language Models
von: Bai, David, et al.
Veröffentlicht: (2024)
von: Bai, David, et al.
Veröffentlicht: (2024)
VizTrust: A Visual Analytics Tool for Capturing User Trust Dynamics in Human-AI Communication
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
From Chat Logs to Collective Insights: Aggregative Question Answering
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?
von: Meshram, Pragati Shuddhodhan, et al.
Veröffentlicht: (2024)
von: Meshram, Pragati Shuddhodhan, et al.
Veröffentlicht: (2024)
TempViz: On the Evaluation of Temporal Knowledge in Text-to-Image Models
von: Holtermann, Carolin, et al.
Veröffentlicht: (2026)
von: Holtermann, Carolin, et al.
Veröffentlicht: (2026)
Instruction-following Evaluation through Verbalizer Manipulation
von: Li, Shiyang, et al.
Veröffentlicht: (2023)
von: Li, Shiyang, et al.
Veröffentlicht: (2023)
ClimateViz: A Benchmark for Statistical Reasoning and Fact Verification on Scientific Charts
von: Su, Ruiran, et al.
Veröffentlicht: (2025)
von: Su, Ruiran, et al.
Veröffentlicht: (2025)
The American Sign Language Knowledge Graph: Infusing ASL Models with Linguistic Knowledge
von: Kezar, Lee, et al.
Veröffentlicht: (2024)
von: Kezar, Lee, et al.
Veröffentlicht: (2024)
Approximating Language Model Training Data from Weights
von: Morris, John X., et al.
Veröffentlicht: (2025)
von: Morris, John X., et al.
Veröffentlicht: (2025)
Demystifying Language Model Forgetting with Low-rank Example Associations
von: Jin, Xisen, et al.
Veröffentlicht: (2024)
von: Jin, Xisen, et al.
Veröffentlicht: (2024)
What Will My Model Forget? Forecasting Forgotten Examples in Language Model Refinement
von: Jin, Xisen, et al.
Veröffentlicht: (2024)
von: Jin, Xisen, et al.
Veröffentlicht: (2024)
ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attribution
von: Alqurnawi, Yahia, et al.
Veröffentlicht: (2026)
von: Alqurnawi, Yahia, et al.
Veröffentlicht: (2026)
ChatGPT for automated grading of short answer questions in mechanical ventilation
von: Jade, Tejas, et al.
Veröffentlicht: (2025)
von: Jade, Tejas, et al.
Veröffentlicht: (2025)
Details Make a Difference: Object State-Sensitive Neurorobotic Task Planning
von: Sun, Xiaowen, et al.
Veröffentlicht: (2024)
von: Sun, Xiaowen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Adjust for Trust: Mitigating Trust-Induced Inappropriate Reliance on AI Assistance
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2025) -
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
von: He, Keyu, et al.
Veröffentlicht: (2025) -
From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered
von: Devic, Siddartha, et al.
Veröffentlicht: (2025) -
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2024) -
WinoWhat: A Parallel Corpus of Paraphrased WinoGrande Sentences with Common Sense Categorization
von: Gevers, Ine, et al.
Veröffentlicht: (2025)