No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
Fuente:
arXiv
Saved in:
| Main Authors: | Cencerrado, Iván Vicente Moreno, Masdemont, Arnau Padrés, Hawthorne, Anton Gonzalvez, Africa, David Demitri, Pacchiardi, Lorenzo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Identifying a Circuit for Verb Conjugation in GPT-2
by: Africa, David Demitri
Published: (2025)
by: Africa, David Demitri
Published: (2025)
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
by: Ivanov, Igor, et al.
Published: (2026)
by: Ivanov, Igor, et al.
Published: (2026)
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models
by: Vendrell, Victor Conchello, et al.
Published: (2026)
by: Vendrell, Victor Conchello, et al.
Published: (2026)
M3Kang: Evaluating Multilingual Multimodal Mathematical Reasoning in Vision-Language Models
by: Torres-Camps, Aleix, et al.
Published: (2026)
by: Torres-Camps, Aleix, et al.
Published: (2026)
Steering Awareness: Detecting Activation Steering from Within
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers
by: Wang, Yuhan, et al.
Published: (2026)
by: Wang, Yuhan, et al.
Published: (2026)
Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs
by: Furuhashi, Momoka, et al.
Published: (2025)
by: Furuhashi, Momoka, et al.
Published: (2025)
Estimating the Usefulness of Clarifying Questions and Answers for Conversational Search
by: Sekulić, Ivan, et al.
Published: (2024)
by: Sekulić, Ivan, et al.
Published: (2024)
Changing Answer Order Can Decrease MMLU Accuracy
by: Gupta, Vipul, et al.
Published: (2024)
by: Gupta, Vipul, et al.
Published: (2024)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
by: Wiegreffe, Sarah, et al.
Published: (2024)
by: Wiegreffe, Sarah, et al.
Published: (2024)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
by: Yao, Siyang, et al.
Published: (2026)
by: Yao, Siyang, et al.
Published: (2026)
Answer is All You Need: Instruction-following Text Embedding via Answering the Question
by: Peng, Letian, et al.
Published: (2024)
by: Peng, Letian, et al.
Published: (2024)
Rehearsing Answers to Probable Questions with Perspective-Taking
by: Shih, Yung-Yu, et al.
Published: (2024)
by: Shih, Yung-Yu, et al.
Published: (2024)
DEEPAMBIGQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
by: Ji, Jiabao, et al.
Published: (2025)
by: Ji, Jiabao, et al.
Published: (2025)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
Polyglots or Multitudes? Multilingual LLM Answers to Value-laden Multiple-Choice Questions
by: Labat, Léo, et al.
Published: (2026)
by: Labat, Léo, et al.
Published: (2026)
Reliable Answers for Recurring Questions: Boosting Text-to-SQL Accuracy with Template Constrained Decoding
by: Jivani, Smit, et al.
Published: (2026)
by: Jivani, Smit, et al.
Published: (2026)
Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only
by: Yao, Jihan, et al.
Published: (2024)
by: Yao, Jihan, et al.
Published: (2024)
Towards Self-Contained Answers: Entity-Based Answer Rewriting in Conversational Search
by: Sekulić, Ivan, et al.
Published: (2024)
by: Sekulić, Ivan, et al.
Published: (2024)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
LLMs Provide Unstable Answers to Legal Questions
by: Blair-Stanek, Andrew, et al.
Published: (2025)
by: Blair-Stanek, Andrew, et al.
Published: (2025)
Automatic Question & Answer Generation Using Generative Large Language Model (LLM)
by: Ehsan, Md. Alvee, et al.
Published: (2025)
by: Ehsan, Md. Alvee, et al.
Published: (2025)
Controllable Decontextualization of Yes/No Question and Answers into Factual Statements
by: Mo, Lingbo, et al.
Published: (2024)
by: Mo, Lingbo, et al.
Published: (2024)
Automatic Question-Answer Generation for Long-Tail Knowledge
by: Kumar, Rohan, et al.
Published: (2024)
by: Kumar, Rohan, et al.
Published: (2024)
Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy
by: Shi, Xiaofeng, et al.
Published: (2025)
by: Shi, Xiaofeng, et al.
Published: (2025)
Knowledge-Augmented Question Error Correction for Chinese Question Answer System with QuestionRAG
by: Qiu, Longpeng, et al.
Published: (2025)
by: Qiu, Longpeng, et al.
Published: (2025)
Consistency Training while Mitigating Obfuscation via Rate Matching
by: Imran, Sohaib, et al.
Published: (2026)
by: Imran, Sohaib, et al.
Published: (2026)
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
by: Zhou, Chengliang, et al.
Published: (2025)
by: Zhou, Chengliang, et al.
Published: (2025)
ExpertQA: Expert-Curated Questions and Attributed Answers
by: Malaviya, Chaitanya, et al.
Published: (2023)
by: Malaviya, Chaitanya, et al.
Published: (2023)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
by: Molfese, Francesco Maria, et al.
Published: (2025)
by: Molfese, Francesco Maria, et al.
Published: (2025)
A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities
by: Imamura, Kenji, et al.
Published: (2026)
by: Imamura, Kenji, et al.
Published: (2026)
Graph Guided Question Answer Generation for Procedural Question-Answering
by: Pham, Hai X., et al.
Published: (2024)
by: Pham, Hai X., et al.
Published: (2024)
Evaluating Quality of Answers for Retrieval-Augmented Generation: A Strong LLM Is All You Need
by: Wang, Yang, et al.
Published: (2024)
by: Wang, Yang, et al.
Published: (2024)
EEE-QA: Exploring Effective and Efficient Question-Answer Representations
by: Hu, Zhanghao, et al.
Published: (2024)
by: Hu, Zhanghao, et al.
Published: (2024)
Interpreting Answers to Yes-No Questions in Dialogues from Multiple Domains
by: Wang, Zijie, et al.
Published: (2024)
by: Wang, Zijie, et al.
Published: (2024)
Student Answer Forecasting: Transformer-Driven Answer Choice Prediction for Language Learning
by: Gado, Elena Grazia, et al.
Published: (2024)
by: Gado, Elena Grazia, et al.
Published: (2024)
Scientific QA System with Verifiable Answers
by: Ljajić, Adela, et al.
Published: (2024)
by: Ljajić, Adela, et al.
Published: (2024)
Answer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic Systems
by: Huang, Tianyi, et al.
Published: (2026)
by: Huang, Tianyi, et al.
Published: (2026)
BPQA Dataset: Evaluating How Well Language Models Leverage Blood Pressures to Answer Biomedical Questions
by: Hang, Chi, et al.
Published: (2025)
by: Hang, Chi, et al.
Published: (2025)
Question: How do Large Language Models perform on the Question Answering tasks? Answer:
by: Fischer, Kevin, et al.
Published: (2024)
by: Fischer, Kevin, et al.
Published: (2024)
Similar Items
-
Identifying a Circuit for Verb Conjugation in GPT-2
by: Africa, David Demitri
Published: (2025) -
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
by: Ivanov, Igor, et al.
Published: (2026) -
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models
by: Vendrell, Victor Conchello, et al.
Published: (2026) -
M3Kang: Evaluating Multilingual Multimodal Mathematical Reasoning in Vision-Language Models
by: Torres-Camps, Aleix, et al.
Published: (2026) -
Steering Awareness: Detecting Activation Steering from Within
by: Rivera, Joshua Fonseca, et al.
Published: (2025)