QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Belinda Z., Kim, Been, Wang, Zi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sudoku-Bench: Evaluating creative reasoning with Sudoku variants
von: Seely, Jeffrey, et al.
Veröffentlicht: (2025)
von: Seely, Jeffrey, et al.
Veröffentlicht: (2025)
Language models show human-like content effects on reasoning tasks
von: Dasgupta, Ishita, et al.
Veröffentlicht: (2022)
von: Dasgupta, Ishita, et al.
Veröffentlicht: (2022)
Multi-step retrieval and reasoning improves radiology question answering with large language models
von: Wind, Sebastian, et al.
Veröffentlicht: (2025)
von: Wind, Sebastian, et al.
Veröffentlicht: (2025)
Are complicated loss functions necessary for teaching LLMs to reason?
von: Carrino, Gabriele, et al.
Veröffentlicht: (2026)
von: Carrino, Gabriele, et al.
Veröffentlicht: (2026)
VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
von: Jain, Dhruv, et al.
Veröffentlicht: (2025)
von: Jain, Dhruv, et al.
Veröffentlicht: (2025)
XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks
von: Jain, Purvam, et al.
Veröffentlicht: (2026)
von: Jain, Purvam, et al.
Veröffentlicht: (2026)
AgentBench: Evaluating LLMs as Agents
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
LLMs cannot find reasoning errors, but can correct them given the error location
von: Tyen, Gladys, et al.
Veröffentlicht: (2023)
von: Tyen, Gladys, et al.
Veröffentlicht: (2023)
(How) Do Language Models Track State?
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
Can LLMs Follow Simple Rules?
von: Mu, Norman, et al.
Veröffentlicht: (2023)
von: Mu, Norman, et al.
Veröffentlicht: (2023)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
von: Maar, Jim, et al.
Veröffentlicht: (2026)
von: Maar, Jim, et al.
Veröffentlicht: (2026)
Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits
von: Kalibhat, Neha, et al.
Veröffentlicht: (2026)
von: Kalibhat, Neha, et al.
Veröffentlicht: (2026)
Can GRPO Help LLMs Transcend Their Pretraining Origin?
von: Ni, Kangqi, et al.
Veröffentlicht: (2025)
von: Ni, Kangqi, et al.
Veröffentlicht: (2025)
Can LLMs Convert Graphs to Text-Attributed Graphs?
von: Wang, Zehong, et al.
Veröffentlicht: (2024)
von: Wang, Zehong, et al.
Veröffentlicht: (2024)
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
von: Wang, Shouren, et al.
Veröffentlicht: (2025)
von: Wang, Shouren, et al.
Veröffentlicht: (2025)
Enough Coin Flips Can Make LLMs Act Bayesian
von: Gupta, Ritwik, et al.
Veröffentlicht: (2025)
von: Gupta, Ritwik, et al.
Veröffentlicht: (2025)
A Framework to Implement 1+N Multi-task Fine-tuning Pattern in LLMs Using the CGC-LORA Algorithm
von: Song, Chao, et al.
Veröffentlicht: (2024)
von: Song, Chao, et al.
Veröffentlicht: (2024)
seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
von: Ramezanali, Mohammad, et al.
Veröffentlicht: (2025)
von: Ramezanali, Mohammad, et al.
Veröffentlicht: (2025)
Can Post-Training Transform LLMs into Causal Reasoners?
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
When can transformers reason with abstract symbols?
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2023)
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2023)
Artificial Expert Intelligence through PAC-reasoning
von: Shalev-Shwartz, Shai, et al.
Veröffentlicht: (2024)
von: Shalev-Shwartz, Shai, et al.
Veröffentlicht: (2024)
Training Language Models to Explain Their Own Computations
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
von: Li, Belinda Z., et al.
Veröffentlicht: (2025)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
von: Park, Jungsoo, et al.
Veröffentlicht: (2025)
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
von: Yan, Kai, et al.
Veröffentlicht: (2025)
von: Yan, Kai, et al.
Veröffentlicht: (2025)
FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?
von: Ravi, Nikil, et al.
Veröffentlicht: (2026)
von: Ravi, Nikil, et al.
Veröffentlicht: (2026)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
von: Manvi, Rohin, et al.
Veröffentlicht: (2024)
von: Manvi, Rohin, et al.
Veröffentlicht: (2024)
InductionBench: LLMs Fail in the Simplest Complexity Class
von: Hua, Wenyue, et al.
Veröffentlicht: (2025)
von: Hua, Wenyue, et al.
Veröffentlicht: (2025)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
von: Kirchhof, Michael, et al.
Veröffentlicht: (2025)
von: Kirchhof, Michael, et al.
Veröffentlicht: (2025)
Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs
von: Zhou, Wei, et al.
Veröffentlicht: (2026)
von: Zhou, Wei, et al.
Veröffentlicht: (2026)
Is continuous CoT better suited for multi-lingual reasoning?
von: Bashir, Ali Hamza, et al.
Veröffentlicht: (2026)
von: Bashir, Ali Hamza, et al.
Veröffentlicht: (2026)
MetaTool: Facilitating Large Language Models to Master Tools with Meta-task Augmentation
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
FCoReBench: Can Large Language Models Solve Challenging First-Order Combinatorial Reasoning Problems?
von: Mittal, Chinmay, et al.
Veröffentlicht: (2024)
von: Mittal, Chinmay, et al.
Veröffentlicht: (2024)
Neural networks for abstraction and reasoning: Towards broad generalization in machines
von: Bober-Irizar, Mikel, et al.
Veröffentlicht: (2024)
von: Bober-Irizar, Mikel, et al.
Veröffentlicht: (2024)
Agribot: agriculture-specific question answer system
von: Jain, Naman, et al.
Veröffentlicht: (2025)
von: Jain, Naman, et al.
Veröffentlicht: (2025)
Training-free LLM Merging for Multi-task Learning
von: Fu, Zichuan, et al.
Veröffentlicht: (2025)
von: Fu, Zichuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Sudoku-Bench: Evaluating creative reasoning with Sudoku variants
von: Seely, Jeffrey, et al.
Veröffentlicht: (2025) -
Language models show human-like content effects on reasoning tasks
von: Dasgupta, Ishita, et al.
Veröffentlicht: (2022) -
Multi-step retrieval and reasoning improves radiology question answering with large language models
von: Wind, Sebastian, et al.
Veröffentlicht: (2025) -
Are complicated loss functions necessary for teaching LLMs to reason?
von: Carrino, Gabriele, et al.
Veröffentlicht: (2026) -
VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
von: Jain, Dhruv, et al.
Veröffentlicht: (2025)