Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Xingjian, Gao, Tianhong, Jin, Suliang, Wang, Tianhao, Ye, Teng, Adar, Eytan, Mei, Qiaozhu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring Bridges Between Algorithmic and AI-generated Art
von: Wu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Wu, Jiaqi, et al.
Veröffentlicht: (2024)
Assessing Affective Objectives for Communicative Visualizations
von: Lee-Robbins, Elsie, et al.
Veröffentlicht: (2026)
von: Lee-Robbins, Elsie, et al.
Veröffentlicht: (2026)
QuizRank: Picking Images by Quizzing VLMs
von: Ji, Tenghao, et al.
Veröffentlicht: (2025)
von: Ji, Tenghao, et al.
Veröffentlicht: (2025)
Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations
von: Liang, Chen, et al.
Veröffentlicht: (2026)
von: Liang, Chen, et al.
Veröffentlicht: (2026)
WigglyEyes: Inferring Eye Movements from Keypress Data
von: Zhu, Yujun, et al.
Veröffentlicht: (2024)
von: Zhu, Yujun, et al.
Veröffentlicht: (2024)
Imagine a dragon made of seaweed: How images enhance learning in Wikipedia
von: Silva, Anita, et al.
Veröffentlicht: (2024)
von: Silva, Anita, et al.
Veröffentlicht: (2024)
EyeSpy: Inferring Eye Gaze via Side-Channel Attacks Against Foveated Rendering
von: Maynard, Paul, et al.
Veröffentlicht: (2026)
von: Maynard, Paul, et al.
Veröffentlicht: (2026)
Feminist Interaction Techniques: Deterring Non-Consensual Screenshots with Interaction Techniques
von: Qiwei, Li, et al.
Veröffentlicht: (2024)
von: Qiwei, Li, et al.
Veröffentlicht: (2024)
Authors' Values and Attitudes Towards AI-bridged Scalable Personalization of Creative Language Arts
von: Kim, Taewook, et al.
Veröffentlicht: (2024)
von: Kim, Taewook, et al.
Veröffentlicht: (2024)
Emoji Promotes Developer Participation and Issue Resolution on GitHub
von: Zhou, Yuhang, et al.
Veröffentlicht: (2023)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2023)
Eyes on the Game: Deciphering Implicit Human Signals to Infer Human Proficiency, Trust, and Intent
von: Hulle, Nikhil, et al.
Veröffentlicht: (2024)
von: Hulle, Nikhil, et al.
Veröffentlicht: (2024)
Seeing Like an AI: How LLMs Apply (and Misapply) Wikipedia Neutrality Norms
von: Ashkinaze, Joshua, et al.
Veröffentlicht: (2024)
von: Ashkinaze, Joshua, et al.
Veröffentlicht: (2024)
When Large Language Models are Reliable for Judging Empathic Communication
von: Kumar, Aakriti, et al.
Veröffentlicht: (2025)
von: Kumar, Aakriti, et al.
Veröffentlicht: (2025)
EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
von: Ashktorab, Zahra, et al.
Veröffentlicht: (2025)
von: Ashktorab, Zahra, et al.
Veröffentlicht: (2025)
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
von: Chiang, Charles, et al.
Veröffentlicht: (2026)
von: Chiang, Charles, et al.
Veröffentlicht: (2026)
Human-Centered Design Recommendations for LLM-as-a-Judge
von: Pan, Qian, et al.
Veröffentlicht: (2024)
von: Pan, Qian, et al.
Veröffentlicht: (2024)
The Laziness of the Crowd: Effort Aversion Among Raters Risks Undermining the Efficacy of X's Community Notes Program
von: Wack, Morgan, et al.
Veröffentlicht: (2026)
von: Wack, Morgan, et al.
Veröffentlicht: (2026)
Attitudes and perceived effectiveness among first-time online instructors during Covid-19
von: Zhang, Owen Xingjian
Veröffentlicht: (2024)
von: Zhang, Owen Xingjian
Veröffentlicht: (2024)
Beyond the Click: A Framework for Inferring Cognitive Traces in Search
von: Zerhoudi, Saber, et al.
Veröffentlicht: (2026)
von: Zerhoudi, Saber, et al.
Veröffentlicht: (2026)
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
von: Szymanski, Annalisa, et al.
Veröffentlicht: (2024)
von: Szymanski, Annalisa, et al.
Veröffentlicht: (2024)
Principled Evaluation with Human Labels: One Rater at a Time and Rater Equivalence
von: Resnick, Paul, et al.
Veröffentlicht: (2021)
von: Resnick, Paul, et al.
Veröffentlicht: (2021)
One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations
von: Lee, Yoonjoo, et al.
Veröffentlicht: (2024)
von: Lee, Yoonjoo, et al.
Veröffentlicht: (2024)
Designing AI Systems that Augment Human Performed vs. Demonstrated Critical Thinking
von: Mei, Katelyn Xiaoying, et al.
Veröffentlicht: (2025)
von: Mei, Katelyn Xiaoying, et al.
Veröffentlicht: (2025)
Decoding Safety Feedback from Diverse Raters: A Data-driven Lens on Responsiveness to Severity
von: Mishra, Pushkar, et al.
Veröffentlicht: (2025)
von: Mishra, Pushkar, et al.
Veröffentlicht: (2025)
Through the Expert's Eyes: Exploring Asynchronous Expert Perspectives and Gaze Visualizations in XR
von: Sayffaerth, Clara, et al.
Veröffentlicht: (2025)
von: Sayffaerth, Clara, et al.
Veröffentlicht: (2025)
Thinking Assistants: LLM-Based Conversational Assistants that Help Users Think By Asking rather than Answering
von: Park, Soya, et al.
Veröffentlicht: (2023)
von: Park, Soya, et al.
Veröffentlicht: (2023)
"A Great Start, But...": Evaluating LLM-Generated Mind Maps for Information Mapping in Video-Based Design
von: He, Tianhao, et al.
Veröffentlicht: (2025)
von: He, Tianhao, et al.
Veröffentlicht: (2025)
MapExplorer: New Content Generation from Low-Dimensional Visualizations
von: Zhang, Xingjian, et al.
Veröffentlicht: (2024)
von: Zhang, Xingjian, et al.
Veröffentlicht: (2024)
Eye2Recall: Exploring the Design of Enhancing Reminiscence Activities via Eye Tracking-Based LLM-Powered Interaction Experience for Older Adults
von: Han, Lei, et al.
Veröffentlicht: (2025)
von: Han, Lei, et al.
Veröffentlicht: (2025)
An LLM-based Simulation Framework for Embodied Conversational Agents in Psychological Counseling
von: Wu, Lixiu, et al.
Veröffentlicht: (2024)
von: Wu, Lixiu, et al.
Veröffentlicht: (2024)
Digital Eyes: Social Implications of XR EyeSight
von: Vergari, Maurizio, et al.
Veröffentlicht: (2024)
von: Vergari, Maurizio, et al.
Veröffentlicht: (2024)
Through Their Eyes: User Perceptions on Sensitive Attribute Inference of Social Media Videos by Visual Language Models
von: Zhang, Shuning, et al.
Veröffentlicht: (2025)
von: Zhang, Shuning, et al.
Veröffentlicht: (2025)
Critical Inker: Scaffolding Critical Thinking in AI-Assisted Writing Through Socratic Questioning
von: Hugenroth, Philipp, et al.
Veröffentlicht: (2026)
von: Hugenroth, Philipp, et al.
Veröffentlicht: (2026)
Talk to Me, Not the Slides: A Real-Time Wearable Assistant for Improving Eye Contact in Presentations
von: Du, Lingyu, et al.
Veröffentlicht: (2026)
von: Du, Lingyu, et al.
Veröffentlicht: (2026)
Trace Mutation in Human-LLM Dialogue: The Transcript as Forensic and Mitigation Surface
von: Bensen, William J.
Veröffentlicht: (2026)
von: Bensen, William J.
Veröffentlicht: (2026)
Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human-AI Collaboration
von: Teng, Zhuyu, et al.
Veröffentlicht: (2026)
von: Teng, Zhuyu, et al.
Veröffentlicht: (2026)
SIAgent: Spatial Interaction Agent via LLM-powered Eye-Hand Motion Intent Understanding in VR
von: Wang, Zhimin, et al.
Veröffentlicht: (2026)
von: Wang, Zhimin, et al.
Veröffentlicht: (2026)
"Can You See Me Think?" Grounding LLM Feedback in Keystrokes and Revision Patterns
von: Zafar, Samra, et al.
Veröffentlicht: (2025)
von: Zafar, Samra, et al.
Veröffentlicht: (2025)
Exploring Communication Dynamics: Eye-tracking Analysis in Pair Programming of Computer Science Education
von: Jang, Wunmin, et al.
Veröffentlicht: (2024)
von: Jang, Wunmin, et al.
Veröffentlicht: (2024)
A-MEM: Agentic Memory for LLM Agents
von: Xu, Wujiang, et al.
Veröffentlicht: (2025)
von: Xu, Wujiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Exploring Bridges Between Algorithmic and AI-generated Art
von: Wu, Jiaqi, et al.
Veröffentlicht: (2024) -
Assessing Affective Objectives for Communicative Visualizations
von: Lee-Robbins, Elsie, et al.
Veröffentlicht: (2026) -
QuizRank: Picking Images by Quizzing VLMs
von: Ji, Tenghao, et al.
Veröffentlicht: (2025) -
Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations
von: Liang, Chen, et al.
Veröffentlicht: (2026) -
WigglyEyes: Inferring Eye Movements from Keypress Data
von: Zhu, Yujun, et al.
Veröffentlicht: (2024)