Saved in:
| Main Authors: | Zhou, Mengtao, Wu, Sifan, Zhang, Huan, Sima, Qi, Liu, Bang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2508.10358 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can LLMs Ask Good Questions?
by: Zhang, Yueheng, et al.
Published: (2025)
by: Zhang, Yueheng, et al.
Published: (2025)
Weak-eval-Strong: Evaluating and Eliciting Lateral Thinking of LLMs with Situation Puzzles
by: Chen, Qi, et al.
Published: (2024)
by: Chen, Qi, et al.
Published: (2024)
What Is Next for LLMs? Next-Generation AI Computing Hardware Using Photonic Chips
by: Li, Renjie, et al.
Published: (2025)
by: Li, Renjie, et al.
Published: (2025)
Improving Clinical Note Generation from Complex Doctor-Patient Conversation
by: Li, Yizhan, et al.
Published: (2024)
by: Li, Yizhan, et al.
Published: (2024)
What Would Happen Next? Predicting Consequences from An Event Causality Graph
by: Zhan, Chuanhong, et al.
Published: (2024)
by: Zhan, Chuanhong, et al.
Published: (2024)
Ask Good Questions for Large Language Models
by: Wu, Qi, et al.
Published: (2025)
by: Wu, Qi, et al.
Published: (2025)
Teaching LLMs to Ask: Self-Querying Category-Theoretic Planning for Under-Specified Reasoning
by: Qu, Shuhui
Published: (2026)
by: Qu, Shuhui
Published: (2026)
PuzzlePlex: Benchmarking Foundation Models on Reasoning and Planning with Puzzles
by: Long, Yitao, et al.
Published: (2025)
by: Long, Yitao, et al.
Published: (2025)
AgentAsk: Multi-Agent Systems Need to Ask
by: Lin, Bohan, et al.
Published: (2025)
by: Lin, Bohan, et al.
Published: (2025)
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
by: Tyagi, Nemika, et al.
Published: (2024)
by: Tyagi, Nemika, et al.
Published: (2024)
Bone Soups: A Seek-and-Soup Model Merging Approach for Controllable Multi-Objective Generation
by: Xie, Guofu, et al.
Published: (2025)
by: Xie, Guofu, et al.
Published: (2025)
AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification
by: Zhang, Xuan, et al.
Published: (2025)
by: Zhang, Xuan, et al.
Published: (2025)
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs
by: Zheng, Xiang, et al.
Published: (2026)
by: Zheng, Xiang, et al.
Published: (2026)
SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
by: Wu, Yutao, et al.
Published: (2025)
by: Wu, Yutao, et al.
Published: (2025)
Fail Fast, or Ask: Mitigating the Deficiencies of Reasoning LLMs with Human-in-the-Loop Systems Engineering
by: Zellinger, Michael J., et al.
Published: (2025)
by: Zellinger, Michael J., et al.
Published: (2025)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
by: Shin, Jisu, et al.
Published: (2024)
by: Shin, Jisu, et al.
Published: (2024)
M^4olGen: Multi-Agent, Multi-Stage Molecular Generation under Precise Multi-Property Constraints
by: Li, Yizhan, et al.
Published: (2026)
by: Li, Yizhan, et al.
Published: (2026)
CircuitSeer: Mining High-Quality Data by Probing Mathematical Reasoning Circuits in LLMs
by: Wang, Shaobo, et al.
Published: (2025)
by: Wang, Shaobo, et al.
Published: (2025)
Grounding Multimodal LLMs to Embodied Agents that Ask for Help with Reinforcement Learning
by: Ramrakhya, Ram, et al.
Published: (2025)
by: Ramrakhya, Ram, et al.
Published: (2025)
SoupLM: Model Integration in Large Language and Multi-Modal Models
by: Bai, Yue, et al.
Published: (2024)
by: Bai, Yue, et al.
Published: (2024)
Students' Feedback Requests and Interactions with the SCRIPT Chatbot: Do They Get What They Ask For?
by: Scholl, Andreas, et al.
Published: (2025)
by: Scholl, Andreas, et al.
Published: (2025)
VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models
by: Xu, Hefei, et al.
Published: (2026)
by: Xu, Hefei, et al.
Published: (2026)
DeepImagine: Learning Biomedical Reasoning via Successive Counterfactual Imagining
by: Zheng, Youze, et al.
Published: (2026)
by: Zheng, Youze, et al.
Published: (2026)
RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs
by: Jin, Hangzhan, et al.
Published: (2025)
by: Jin, Hangzhan, et al.
Published: (2025)
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models
by: Lyu, Zesen, et al.
Published: (2025)
by: Lyu, Zesen, et al.
Published: (2025)
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework
by: Şenol, Ali, et al.
Published: (2026)
by: Şenol, Ali, et al.
Published: (2026)
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
by: Tang, Kexian, et al.
Published: (2025)
by: Tang, Kexian, et al.
Published: (2025)
PuzzleJAX: A Benchmark for Reasoning and Learning
by: Earle, Sam, et al.
Published: (2025)
by: Earle, Sam, et al.
Published: (2025)
Measuring Iterative Temporal Reasoning with Time Puzzles
by: Wang, Zhengxiang, et al.
Published: (2026)
by: Wang, Zhengxiang, et al.
Published: (2026)
What You See is What You Ask: Evaluating Audio Descriptions
by: Kala, Divy, et al.
Published: (2025)
by: Kala, Divy, et al.
Published: (2025)
The Token Games: Evaluating Language Model Reasoning with Puzzle Duels
by: Henniger, Simon, et al.
Published: (2026)
by: Henniger, Simon, et al.
Published: (2026)
HardcoreLogic: Challenging Large Reasoning Models with Long-tail Logic Puzzle Games
by: Liang, Jingcong, et al.
Published: (2025)
by: Liang, Jingcong, et al.
Published: (2025)
PUZZLED: Jailbreaking LLMs through Word-Based Puzzles
by: Ahn, Yelim, et al.
Published: (2025)
by: Ahn, Yelim, et al.
Published: (2025)
Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning
by: Ghosal, Deepanway, et al.
Published: (2024)
by: Ghosal, Deepanway, et al.
Published: (2024)
What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code
by: Zhao, Yuze, et al.
Published: (2026)
by: Zhao, Yuze, et al.
Published: (2026)
Imagination-Limited Q-Learning for Offline Reinforcement Learning
by: Liu, Wenhui, et al.
Published: (2025)
by: Liu, Wenhui, et al.
Published: (2025)
Stepwise Reasoning Error Disruption Attack of LLMs
by: Peng, Jingyu, et al.
Published: (2024)
by: Peng, Jingyu, et al.
Published: (2024)
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
by: Leng, Jixuan, et al.
Published: (2025)
by: Leng, Jixuan, et al.
Published: (2025)
Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action
by: Zhang, Yi, et al.
Published: (2026)
by: Zhang, Yi, et al.
Published: (2026)
Similar Items
-
Can LLMs Ask Good Questions?
by: Zhang, Yueheng, et al.
Published: (2025) -
Weak-eval-Strong: Evaluating and Eliciting Lateral Thinking of LLMs with Situation Puzzles
by: Chen, Qi, et al.
Published: (2024) -
What Is Next for LLMs? Next-Generation AI Computing Hardware Using Photonic Chips
by: Li, Renjie, et al.
Published: (2025) -
Improving Clinical Note Generation from Complex Doctor-Patient Conversation
by: Li, Yizhan, et al.
Published: (2024) -
What Would Happen Next? Predicting Consequences from An Event Causality Graph
by: Zhan, Chuanhong, et al.
Published: (2024)