Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Balepur, Nishant, Gu, Feng, Ravichander, Abhilasha, Feng, Shi, Boyd-Graber, Jordan, Rudinger, Rachel |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
par: Balepur, Nishant, et autres
Publié: (2024)
par: Balepur, Nishant, et autres
Publié: (2024)
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
par: Balepur, Nishant, et autres
Publié: (2025)
par: Balepur, Nishant, et autres
Publié: (2025)
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
par: Srikanth, Neha, et autres
Publié: (2026)
par: Srikanth, Neha, et autres
Publié: (2026)
Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
par: Balepur, Nishant, et autres
Publié: (2025)
par: Balepur, Nishant, et autres
Publié: (2025)
KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
par: Shu, Matthew, et autres
Publié: (2024)
par: Shu, Matthew, et autres
Publié: (2024)
Is Your Large Language Model Knowledgeable or a Choices-Only Cheater?
par: Balepur, Nishant, et autres
Publié: (2024)
par: Balepur, Nishant, et autres
Publié: (2024)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
par: Balepur, Nishant, et autres
Publié: (2025)
par: Balepur, Nishant, et autres
Publié: (2025)
Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering
par: Srikanth, Neha, et autres
Publié: (2023)
par: Srikanth, Neha, et autres
Publié: (2023)
It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning
par: Balepur, Nishant, et autres
Publié: (2023)
par: Balepur, Nishant, et autres
Publié: (2023)
Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
par: Balepur, Nishant, et autres
Publié: (2025)
par: Balepur, Nishant, et autres
Publié: (2025)
CFMatch: Aligning Automated Answer Equivalence Evaluation with Expert Judgments For Open-Domain Question Answering
par: Li, Zongxia, et autres
Publié: (2024)
par: Li, Zongxia, et autres
Publié: (2024)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
par: Palta, Shramay, et autres
Publié: (2024)
par: Palta, Shramay, et autres
Publié: (2024)
A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick
par: Balepur, Nishant, et autres
Publié: (2024)
par: Balepur, Nishant, et autres
Publié: (2024)
Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities
par: Balepur, Nishant, et autres
Publié: (2025)
par: Balepur, Nishant, et autres
Publié: (2025)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
par: Balepur, Nishant, et autres
Publié: (2026)
par: Balepur, Nishant, et autres
Publié: (2026)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
par: Gor, Maharshi, et autres
Publié: (2024)
par: Gor, Maharshi, et autres
Publié: (2024)
On the Mutual Influence of Gender and Occupation in LLM Representations
par: An, Haozhe, et autres
Publié: (2025)
par: An, Haozhe, et autres
Publié: (2025)
On the Influence of Gender and Race in Romantic Relationship Prediction from Large Language Models
par: Sancheti, Abhilasha, et autres
Publié: (2024)
par: Sancheti, Abhilasha, et autres
Publié: (2024)
SWE-QA: Can Language Models Answer Repository-level Code Questions?
par: Peng, Weihan, et autres
Publié: (2025)
par: Peng, Weihan, et autres
Publié: (2025)
Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering
par: Calvo-Bartolomé, Lorena, et autres
Publié: (2025)
par: Calvo-Bartolomé, Lorena, et autres
Publié: (2025)
How the Advent of Ubiquitous Large Language Models both Stymie and Turbocharge Dynamic Adversarial Question Generation
par: Sung, Yoo Yeon, et autres
Publié: (2024)
par: Sung, Yoo Yeon, et autres
Publié: (2024)
MODS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document Collections
par: Balepur, Nishant, et autres
Publié: (2025)
par: Balepur, Nishant, et autres
Publié: (2025)
DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute
par: Balepur, Nishant, et autres
Publié: (2026)
par: Balepur, Nishant, et autres
Publié: (2026)
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
par: Gor, Maharshi, et autres
Publié: (2026)
par: Gor, Maharshi, et autres
Publié: (2026)
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
par: Acquaye, Christabel, et autres
Publié: (2026)
par: Acquaye, Christabel, et autres
Publié: (2026)
GenSco: Can Question Decomposition based Passage Alignment improve Question Answering?
par: Fazili, Barah, et autres
Publié: (2024)
par: Fazili, Barah, et autres
Publié: (2024)
Graph Guided Question Answer Generation for Procedural Question-Answering
par: Pham, Hai X., et autres
Publié: (2024)
par: Pham, Hai X., et autres
Publié: (2024)
Question: How do Large Language Models perform on the Question Answering tasks? Answer:
par: Fischer, Kevin, et autres
Publié: (2024)
par: Fischer, Kevin, et autres
Publié: (2024)
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
par: Ravichander, Abhilasha, et autres
Publié: (2025)
par: Ravichander, Abhilasha, et autres
Publié: (2025)
Judge Before Answer: Can MLLM Discern the False Premise in Question?
par: Li, Jidong, et autres
Publié: (2025)
par: Li, Jidong, et autres
Publié: (2025)
SearchRAG: Can Search Engines Be Helpful for LLM-based Medical Question Answering?
par: Shi, Yucheng, et autres
Publié: (2025)
par: Shi, Yucheng, et autres
Publié: (2025)
Evaluate What You Can't Evaluate: Unassessable Quality for Generated Response
par: Liu, Yongkang, et autres
Publié: (2023)
par: Liu, Yongkang, et autres
Publié: (2023)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
par: Yao, Siyang, et autres
Publié: (2026)
par: Yao, Siyang, et autres
Publié: (2026)
What Has Been Lost with Synthetic Evaluation?
par: Gill, Alexander, et autres
Publié: (2025)
par: Gill, Alexander, et autres
Publié: (2025)
Question Answering with LLMs and Learning from Answer Sets
par: Borroto, Manuel, et autres
Publié: (2025)
par: Borroto, Manuel, et autres
Publié: (2025)
PEDANTS: Cheap but Effective and Interpretable Answer Equivalence
par: Li, Zongxia, et autres
Publié: (2024)
par: Li, Zongxia, et autres
Publié: (2024)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
par: Molfese, Francesco Maria, et autres
Publié: (2025)
par: Molfese, Francesco Maria, et autres
Publié: (2025)
The Curious Case of Factuality Finetuning: Models' Internal Beliefs Can Improve Factuality
par: Newman, Benjamin, et autres
Publié: (2025)
par: Newman, Benjamin, et autres
Publié: (2025)
Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering
par: Nachshoni, Eviatar, et autres
Publié: (2025)
par: Nachshoni, Eviatar, et autres
Publié: (2025)
Evaluating Answer Reranking Strategies in Time-sensitive Question Answering
par: Kardan, Mehmet, et autres
Publié: (2025)
par: Kardan, Mehmet, et autres
Publié: (2025)
Documents similaires
-
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
par: Balepur, Nishant, et autres
Publié: (2024) -
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above
par: Balepur, Nishant, et autres
Publié: (2025) -
DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering
par: Srikanth, Neha, et autres
Publié: (2026) -
Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas
par: Balepur, Nishant, et autres
Publié: (2025) -
KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students
par: Shu, Matthew, et autres
Publié: (2024)