When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Jane, Shar, Ryan, Pfau, Jacob, Talwalkar, Ameet, He, He, Chen, Valerie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why Do Decision Makers (Not) Use AI? A Cross-Domain Analysis of Factors Impacting AI Adoption
by: Yu, Rebecca, et al.
Published: (2025)
by: Yu, Rebecca, et al.
Published: (2025)
CodingGenie: A Proactive LLM-Powered Programming Assistant
by: Zhao, Sebastian, et al.
Published: (2025)
by: Zhao, Sebastian, et al.
Published: (2025)
Need Help? Designing Proactive AI Assistants for Programming
by: Chen, Valerie, et al.
Published: (2024)
by: Chen, Valerie, et al.
Published: (2024)
Where Does My Model Underperform? A Human Evaluation of Slice Discovery Algorithms
by: Johnson, Nari, et al.
Published: (2023)
by: Johnson, Nari, et al.
Published: (2023)
Dynamic Personalization Through Continuous Feedback Loops in Interactive AI Systems
by: He, Liu
Published: (2026)
by: He, Liu
Published: (2026)
The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers
by: Mozannar, Hussein, et al.
Published: (2024)
by: Mozannar, Hussein, et al.
Published: (2024)
CodeAlignBench: Assessing Code Generation Models on Developer-Preferred Code Adjustments
by: Mehralian, Forough, et al.
Published: (2025)
by: Mehralian, Forough, et al.
Published: (2025)
Learning Personalized Decision Support Policies
by: Bhatt, Umang, et al.
Published: (2023)
by: Bhatt, Umang, et al.
Published: (2023)
What We Talk About When We Talk About Frameworks in HCI
by: Fang, Shitao, et al.
Published: (2026)
by: Fang, Shitao, et al.
Published: (2026)
ReTrace: Interactive Visualizations for Reasoning Traces of Large Reasoning Models
by: Felder, Ludwig, et al.
Published: (2025)
by: Felder, Ludwig, et al.
Published: (2025)
Talking Spell: A Wearable System Enabling Real-Time Anthropomorphic Voice Interaction with Everyday Objects
by: Wang, Xuetong, et al.
Published: (2025)
by: Wang, Xuetong, et al.
Published: (2025)
Modulating Language Model Experiences through Frictions
by: Collins, Katherine M., et al.
Published: (2024)
by: Collins, Katherine M., et al.
Published: (2024)
When LLMs fall short in Deductive Coding: Model Comparison and Human AI Collaboration Workflow Design
by: Li, Zijian, et al.
Published: (2025)
by: Li, Zijian, et al.
Published: (2025)
What Do We Mean When We Talk About Data Storytelling?
by: Yang, Leni, et al.
Published: (2025)
by: Yang, Leni, et al.
Published: (2025)
Seeing to Think? How Source Transparency Design Shapes Interactive Information Seeking and Evaluation in Conversational AI
by: He, Jiangen, et al.
Published: (2026)
by: He, Jiangen, et al.
Published: (2026)
When LLMs Enter Everyday Feminism on Chinese Social Media: Opportunities and Risks for Women's Empowerment
by: Zhang, Runhua, et al.
Published: (2026)
by: Zhang, Runhua, et al.
Published: (2026)
How Beginning Programmers and Code LLMs (Mis)read Each Other
by: Nguyen, Sydney, et al.
Published: (2024)
by: Nguyen, Sydney, et al.
Published: (2024)
Interaction Techniques for Exploratory Data Visualization on Mobile Devices
by: Snyder, Luke S., et al.
Published: (2024)
by: Snyder, Luke S., et al.
Published: (2024)
Look and Talk: Seamless AI Assistant Interaction with Gaze-Triggered Activation
by: Qing, Zhang, et al.
Published: (2025)
by: Qing, Zhang, et al.
Published: (2025)
A Comprehensive Survey of Electrical Stimulation Haptic Feedback in Human-Computer Interaction
by: Yang, Simin, et al.
Published: (2025)
by: Yang, Simin, et al.
Published: (2025)
Facilitating Human Feedback for GenAI Prompt Optimization
by: Sherson, Jacob, et al.
Published: (2024)
by: Sherson, Jacob, et al.
Published: (2024)
From LLM-Driven Trading Card Generation to Procedural Relatedness: A Pokémon Case Study
by: Pfau, Johannes, et al.
Published: (2026)
by: Pfau, Johannes, et al.
Published: (2026)
Does the TalkMoves Codebook Generalize to One-on-One Tutoring and Multimodal Interaction?
by: Focsan, Corina Luca, et al.
Published: (2026)
by: Focsan, Corina Luca, et al.
Published: (2026)
Reject or Not?: A Benchmark for Voice Assistant Query Rejection in Smart Home Scenario and an Improved Method Based on LLMs
by: Men, Huichao, et al.
Published: (2025)
by: Men, Huichao, et al.
Published: (2025)
From Fads to Classics -- Analyzing Video Game Trend Evolutions through Steam Tags
by: Grelier, Nicolas, et al.
Published: (2025)
by: Grelier, Nicolas, et al.
Published: (2025)
When Qualitative Research Meets Large Language Model: Exploring the Potential of QualiGPT as a Tool for Qualitative Coding
by: Zhang, He, et al.
Published: (2024)
by: Zhang, He, et al.
Published: (2024)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
by: Badawi, Abeer, et al.
Published: (2025)
by: Badawi, Abeer, et al.
Published: (2025)
Socially Fluent, Socially Awkward: Artificial Intelligence Relational Talk Backfires in Commercial Interactions
by: Dharmaputri, Stephanie Kwari, et al.
Published: (2026)
by: Dharmaputri, Stephanie Kwari, et al.
Published: (2026)
SPHERE: Scaling Personalized Feedback in Programming Classrooms with Structured Review of LLM Outputs
by: Tang, Xiaohang, et al.
Published: (2024)
by: Tang, Xiaohang, et al.
Published: (2024)
REVA: Supporting LLM-Generated Programming Feedback Validation at Scale Through User Attention-based Adaptation
by: Tang, Xiaohang, et al.
Published: (2025)
by: Tang, Xiaohang, et al.
Published: (2025)
When Peers Outperform AI (and When They Don't): Interaction Quality Over Modality
by: Morris, Caitlin, et al.
Published: (2026)
by: Morris, Caitlin, et al.
Published: (2026)
Exploring Re-inforcement Learning via Human Feedback under User Heterogeneity
by: Shashidhar, Sarvesh, et al.
Published: (2026)
by: Shashidhar, Sarvesh, et al.
Published: (2026)
When and How to Integrate Multimodal Large Language Models in College Psychotherapy: Perspectives from Multi-stakeholders
by: Wang, Jiyao, et al.
Published: (2025)
by: Wang, Jiyao, et al.
Published: (2025)
RECAP: An End-to-End Platform for Capturing, Replaying, and Analyzing AI-Assisted Programming Interactions
by: He, Keyu, et al.
Published: (2026)
by: He, Keyu, et al.
Published: (2026)
Visualizationary: Automating Design Feedback for Visualization Designers using LLMs
by: Shin, Sungbok, et al.
Published: (2024)
by: Shin, Sungbok, et al.
Published: (2024)
Enhancing Interaction with Augmented Reality through Mid-Air Haptic Feedback: Architecture Design and User Feedback
by: Vaquero-Melchor, Diego, et al.
Published: (2025)
by: Vaquero-Melchor, Diego, et al.
Published: (2025)
Audience in the Loop: Viewer Feedback-Driven Content Creation in Micro-drama Production on Social Media
by: Cao, Gengchen, et al.
Published: (2026)
by: Cao, Gengchen, et al.
Published: (2026)
When the Chain Breaks: Interactive Diagnosis of LLM Chain-of-Thought Reasoning Errors
by: Chen, Shiwei, et al.
Published: (2026)
by: Chen, Shiwei, et al.
Published: (2026)
Code Shaping: Iterative Code Editing with Free-form AI-Interpreted Sketching
by: Yen, Ryan, et al.
Published: (2025)
by: Yen, Ryan, et al.
Published: (2025)
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
by: Li, Karen Jia-Hui, et al.
Published: (2025)
by: Li, Karen Jia-Hui, et al.
Published: (2025)
Similar Items
-
Why Do Decision Makers (Not) Use AI? A Cross-Domain Analysis of Factors Impacting AI Adoption
by: Yu, Rebecca, et al.
Published: (2025) -
CodingGenie: A Proactive LLM-Powered Programming Assistant
by: Zhao, Sebastian, et al.
Published: (2025) -
Need Help? Designing Proactive AI Assistants for Programming
by: Chen, Valerie, et al.
Published: (2024) -
Where Does My Model Underperform? A Human Evaluation of Slice Discovery Algorithms
by: Johnson, Nari, et al.
Published: (2023) -
Dynamic Personalization Through Continuous Feedback Loops in Interactive AI Systems
by: He, Liu
Published: (2026)