Teaching People LLM's Errors and Getting it Right
Fuente:
arXiv
Saved in:
| Main Authors: | Stringham, Nathan, Chaleshtori, Fateme Hashemi, Yan, Xinyuan, Xu, Zhichao, Wang, Bei, Marasović, Ana |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chain-of-Thought Unfaithfulness as Disguised Accuracy
by: Bentham, Oliver, et al.
Published: (2024)
by: Bentham, Oliver, et al.
Published: (2024)
On Evaluating Explanation Utility for Human-AI Decision Making in NLP
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024)
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024)
Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
by: Tutek, Martin, et al.
Published: (2025)
by: Tutek, Martin, et al.
Published: (2025)
BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs
by: Woo, Jesse, et al.
Published: (2025)
by: Woo, Jesse, et al.
Published: (2025)
Whispers of Doubt Amidst Echoes of Triumph in NLP Robustness
by: Gupta, Ashim, et al.
Published: (2023)
by: Gupta, Ashim, et al.
Published: (2023)
What Has Been Lost with Synthetic Evaluation?
by: Gill, Alexander, et al.
Published: (2025)
by: Gill, Alexander, et al.
Published: (2025)
Are We on the Right Way to Assessing LLM-as-a-Judge?
by: Feng, Yuanning, et al.
Published: (2025)
by: Feng, Yuanning, et al.
Published: (2025)
Sustainability via LLM Right-sizing
by: Haase, Jennifer, et al.
Published: (2025)
by: Haase, Jennifer, et al.
Published: (2025)
Enhancing LLM-Based Data Annotation with Error Decomposition
by: Xu, Zhen, et al.
Published: (2026)
by: Xu, Zhen, et al.
Published: (2026)
How Can I Get It Right? Using GPT to Rephrase Incorrect Trainee Responses
by: Lin, Jionghao, et al.
Published: (2024)
by: Lin, Jionghao, et al.
Published: (2024)
Apollo: A Lightweight Multilingual Medical LLM towards Democratizing Medical AI to 6B People
by: Wang, Xidong, et al.
Published: (2024)
by: Wang, Xidong, et al.
Published: (2024)
CHILL at SemEval-2025 Task 2: You Can't Just Throw Entities and Hope -- Make Your LLM to Get Them Right
by: Lee, Jaebok, et al.
Published: (2025)
by: Lee, Jaebok, et al.
Published: (2025)
Who Gets the Mic? Investigating Gender Bias in the Speaker Assignment of a Speech-LLM
by: Puhach, Dariia, et al.
Published: (2025)
by: Puhach, Dariia, et al.
Published: (2025)
This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
by: Yun, Hye Sun, et al.
Published: (2026)
by: Yun, Hye Sun, et al.
Published: (2026)
Teaching LLMs According to Their Aptitude: Adaptive Reasoning for Mathematical Problem Solving
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Societal Alignment Frameworks Can Improve LLM Alignment
by: Stańczak, Karolina, et al.
Published: (2025)
by: Stańczak, Karolina, et al.
Published: (2025)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
by: Shi, Zhichao, et al.
Published: (2025)
by: Shi, Zhichao, et al.
Published: (2025)
Detecting Errors through Ensembling Prompts (DEEP): An End-to-End LLM Framework for Detecting Factual Errors
by: Chandler, Alex, et al.
Published: (2024)
by: Chandler, Alex, et al.
Published: (2024)
Krutrim LLM: Multilingual Foundational Model for over a Billion People
by: Kallappa, Aditya, et al.
Published: (2025)
by: Kallappa, Aditya, et al.
Published: (2025)
CEC-Zero: Chinese Error Correction Solution Based on LLM
by: Zhang, Sophie, et al.
Published: (2025)
by: Zhang, Sophie, et al.
Published: (2025)
FocusLLM: Precise Understanding of Long Context by Dynamic Condensing
by: Li, Zhenyu, et al.
Published: (2024)
by: Li, Zhenyu, et al.
Published: (2024)
Asking the Right Questions: Benchmarking Large Language Models in the Development of Clinical Consultation Templates
by: McCoy, Liam G., et al.
Published: (2025)
by: McCoy, Liam G., et al.
Published: (2025)
A Survey on LLM-as-a-Judge
by: Gu, Jiawei, et al.
Published: (2024)
by: Gu, Jiawei, et al.
Published: (2024)
TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement
by: Feng, Zhaopeng, et al.
Published: (2024)
by: Feng, Zhaopeng, et al.
Published: (2024)
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
by: Pan, Wenbo, et al.
Published: (2025)
by: Pan, Wenbo, et al.
Published: (2025)
LLM-First Search: Self-Guided Exploration of the Solution Space
by: Herr, Nathan, et al.
Published: (2025)
by: Herr, Nathan, et al.
Published: (2025)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
Division-of-Thoughts: Harnessing Hybrid Language Model Synergy for Efficient On-Device Agents
by: Shao, Chenyang, et al.
Published: (2025)
by: Shao, Chenyang, et al.
Published: (2025)
LLMs Can Get "Brain Rot": A Pilot Study on Twitter/X
by: Xing, Shuo, et al.
Published: (2025)
by: Xing, Shuo, et al.
Published: (2025)
Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
by: Liu, Geng, et al.
Published: (2026)
by: Liu, Geng, et al.
Published: (2026)
Leveraging What's Overfixed: Post-Correction via LLM Grammatical Error Overcorrection
by: Park, Taehee, et al.
Published: (2025)
by: Park, Taehee, et al.
Published: (2025)
Re-Ex: Revising after Explanation Reduces the Factual Errors in LLM Responses
by: Kim, Juyeon, et al.
Published: (2024)
by: Kim, Juyeon, et al.
Published: (2024)
LLMCL-GEC: Advancing Grammatical Error Correction with LLM-Driven Curriculum Learning
by: Fang, Tao, et al.
Published: (2024)
by: Fang, Tao, et al.
Published: (2024)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
by: Jin, Jiho, et al.
Published: (2026)
by: Jin, Jiho, et al.
Published: (2026)
UltraLogic: Enhancing LLM Reasoning through Large-Scale Data Synthesis and Bipolar Float Reward
by: Liu, Yile, et al.
Published: (2026)
by: Liu, Yile, et al.
Published: (2026)
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
by: Yang, Xintong, et al.
Published: (2026)
by: Yang, Xintong, et al.
Published: (2026)
FlexKBQA: A Flexible LLM-Powered Framework for Few-Shot Knowledge Base Question Answering
by: Li, Zhenyu, et al.
Published: (2023)
by: Li, Zhenyu, et al.
Published: (2023)
Don't Just Demo, Teach Me the Principles: A Principle-Based Multi-Agent Prompting Strategy for Text Classification
by: Wei, Peipei, et al.
Published: (2025)
by: Wei, Peipei, et al.
Published: (2025)
Retrieval-Augmented Guardrails for AI-Drafted Patient-Portal Messages: Error Taxonomy Construction and Large-Scale Evaluation
by: Chen, Wenyuan, et al.
Published: (2025)
by: Chen, Wenyuan, et al.
Published: (2025)
Temporal Consistency for LLM Reasoning Process Error Identification
by: Guo, Jiacheng, et al.
Published: (2025)
by: Guo, Jiacheng, et al.
Published: (2025)
Similar Items
-
Chain-of-Thought Unfaithfulness as Disguised Accuracy
by: Bentham, Oliver, et al.
Published: (2024) -
On Evaluating Explanation Utility for Human-AI Decision Making in NLP
by: Chaleshtori, Fateme Hashemi, et al.
Published: (2024) -
Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
by: Tutek, Martin, et al.
Published: (2025) -
BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs
by: Woo, Jesse, et al.
Published: (2025) -
Whispers of Doubt Amidst Echoes of Triumph in NLP Robustness
by: Gupta, Ashim, et al.
Published: (2023)