FINEREASON: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Guizhen, Xu, Weiwen, Zhang, Hao, Chan, Hou Pong, Liu, Chaoqun, Bing, Lidong, Zhao, Deli, Luu, Anh Tuan, Rong, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning
von: Chen, Guizhen, et al.
Veröffentlicht: (2025)
von: Chen, Guizhen, et al.
Veröffentlicht: (2025)
Reasoning Paths Optimization: Learning to Reason and Explore From Diverse Paths
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models
von: Liu, Chaoqun, et al.
Veröffentlicht: (2024)
von: Liu, Chaoqun, et al.
Veröffentlicht: (2024)
Exploring the Potential of Large Language Models in Computational Argumentation
von: Chen, Guizhen, et al.
Veröffentlicht: (2023)
von: Chen, Guizhen, et al.
Veröffentlicht: (2023)
SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia
von: Liu, Chaoqun, et al.
Veröffentlicht: (2025)
von: Liu, Chaoqun, et al.
Veröffentlicht: (2025)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
von: Liu, Chaoqun, et al.
Veröffentlicht: (2025)
von: Liu, Chaoqun, et al.
Veröffentlicht: (2025)
Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks
von: Li, Xingxuan, et al.
Veröffentlicht: (2024)
von: Li, Xingxuan, et al.
Veröffentlicht: (2024)
Babel: Open Multilingual Large Language Models Serving Over 90% of Global Speakers
von: Zhao, Yiran, et al.
Veröffentlicht: (2025)
von: Zhao, Yiran, et al.
Veröffentlicht: (2025)
Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations
von: Xiao, Chenghao, et al.
Veröffentlicht: (2025)
von: Xiao, Chenghao, et al.
Veröffentlicht: (2025)
Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
von: Zhao, Ruochen, et al.
Veröffentlicht: (2024)
von: Zhao, Ruochen, et al.
Veröffentlicht: (2024)
Zero-to-Strong Generalization: Eliciting Strong Capabilities of Large Language Models Iteratively without Gold Labels
von: Liu, Chaoqun, et al.
Veröffentlicht: (2024)
von: Liu, Chaoqun, et al.
Veröffentlicht: (2024)
SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2024)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
von: LASA Team, et al.
Veröffentlicht: (2025)
von: LASA Team, et al.
Veröffentlicht: (2025)
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
von: Yuan, Ruifeng, et al.
Veröffentlicht: (2025)
von: Yuan, Ruifeng, et al.
Veröffentlicht: (2025)
Scaling Language-Centric Omnimodal Representation Learning
von: Xiao, Chenghao, et al.
Veröffentlicht: (2025)
von: Xiao, Chenghao, et al.
Veröffentlicht: (2025)
From Stimuli to Minds: Enhancing Psychological Reasoning in LLMs via Bilateral Reinforcement Learning
von: Feng, Yichao, et al.
Veröffentlicht: (2025)
von: Feng, Yichao, et al.
Veröffentlicht: (2025)
M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
Unsupervised Hallucination Detection by Inspecting Reasoning Processes
von: Srey, Ponhvoan, et al.
Veröffentlicht: (2025)
von: Srey, Ponhvoan, et al.
Veröffentlicht: (2025)
Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness Predictions
von: Nguyen, Thong, et al.
Veröffentlicht: (2022)
von: Nguyen, Thong, et al.
Veröffentlicht: (2022)
How do Large Language Models Handle Multilingualism?
von: Zhao, Yiran, et al.
Veröffentlicht: (2024)
von: Zhao, Yiran, et al.
Veröffentlicht: (2024)
Pruning General Large Language Models into Customized Expert Models
von: Zhao, Yirao, et al.
Veröffentlicht: (2025)
von: Zhao, Yirao, et al.
Veröffentlicht: (2025)
Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs
von: Hu, Man, et al.
Veröffentlicht: (2025)
von: Hu, Man, et al.
Veröffentlicht: (2025)
Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness
von: Cheang, Chi Seng, et al.
Veröffentlicht: (2025)
von: Cheang, Chi Seng, et al.
Veröffentlicht: (2025)
Are LLMs Good Zero-Shot Fallacy Classifiers?
von: Pan, Fengjun, et al.
Veröffentlicht: (2024)
von: Pan, Fengjun, et al.
Veröffentlicht: (2024)
Gradient-Boosted Decision Tree for Listwise Context Model in Multimodal Review Helpfulness Prediction
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
SynTQA: Synergistic Table-based Question Answering via Mixture of Text-to-SQL and E2E TQA
von: Zhang, Siyue, et al.
Veröffentlicht: (2024)
von: Zhang, Siyue, et al.
Veröffentlicht: (2024)
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
von: Tyagi, Nemika, et al.
Veröffentlicht: (2024)
von: Tyagi, Nemika, et al.
Veröffentlicht: (2024)
Read as You See: Guiding Unimodal LLMs for Low-Resource Explainable Harmful Meme Detection
von: Pan, Fengjun, et al.
Veröffentlicht: (2025)
von: Pan, Fengjun, et al.
Veröffentlicht: (2025)
ClozeMath: Improving Mathematical Reasoning in Language Models by Learning to Fill Equations
von: Pham, Quang Hieu, et al.
Veröffentlicht: (2025)
von: Pham, Quang Hieu, et al.
Veröffentlicht: (2025)
Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers
von: Nguyen, Viet-Anh, et al.
Veröffentlicht: (2025)
von: Nguyen, Viet-Anh, et al.
Veröffentlicht: (2025)
InfoCTM: A Mutual Information Maximization Perspective of Cross-Lingual Topic Modeling
von: Wu, Xiaobao, et al.
Veröffentlicht: (2023)
von: Wu, Xiaobao, et al.
Veröffentlicht: (2023)
Can LLMs Generate and Solve Linguistic Olympiad Puzzles?
von: Majmudar, Neh, et al.
Veröffentlicht: (2025)
von: Majmudar, Neh, et al.
Veröffentlicht: (2025)
Cell-o1: Training LLMs to Solve Single-Cell Reasoning Puzzles with Reinforcement Learning
von: Fang, Yin, et al.
Veröffentlicht: (2025)
von: Fang, Yin, et al.
Veröffentlicht: (2025)
ToTRL: Unlock LLM Tree-of-Thoughts Reasoning Potential through Puzzles Solving
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
von: Sun, Yu, et al.
Veröffentlicht: (2025)
von: Sun, Yu, et al.
Veröffentlicht: (2025)
A Solver-in-the-Loop Framework for Improving LLMs on Answer Set Programming for Logic Puzzle Solving
von: Schrader, Timo Pierre, et al.
Veröffentlicht: (2025)
von: Schrader, Timo Pierre, et al.
Veröffentlicht: (2025)
A Survey on Neural Topic Models: Methods, Applications, and Challenges
von: Wu, Xiaobao, et al.
Veröffentlicht: (2024)
von: Wu, Xiaobao, et al.
Veröffentlicht: (2024)
Towards the TopMost: A Topic Modeling System Toolkit
von: Wu, Xiaobao, et al.
Veröffentlicht: (2023)
von: Wu, Xiaobao, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning
von: Chen, Guizhen, et al.
Veröffentlicht: (2025) -
Reasoning Paths Optimization: Learning to Reason and Explore From Diverse Paths
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024) -
Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models
von: Liu, Chaoqun, et al.
Veröffentlicht: (2024) -
Exploring the Potential of Large Language Models in Computational Argumentation
von: Chen, Guizhen, et al.
Veröffentlicht: (2023) -
SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia
von: Liu, Chaoqun, et al.
Veröffentlicht: (2025)