A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
Fuente:
arXiv
Guardado en:
| Autores principales: | Ke, Zixuan, Jiao, Fangkai, Ming, Yifei, Nguyen, Xuan-Phi, Xu, Austin, Long, Do Xuan, Li, Minzhi, Qin, Chengwei, Wang, Peifeng, Savarese, Silvio, Xiong, Caiming, Joty, Shafiq |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Demystifying Domain-adaptive Post-training for Financial LLMs
por: Ke, Zixuan, et al.
Publicado: (2025)
por: Ke, Zixuan, et al.
Publicado: (2025)
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
por: Nguyen, Xuan-Phi, et al.
Publicado: (2025)
por: Nguyen, Xuan-Phi, et al.
Publicado: (2025)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
por: Pandit, Shrey, et al.
Publicado: (2025)
por: Pandit, Shrey, et al.
Publicado: (2025)
Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows
por: Ming, Yifei, et al.
Publicado: (2025)
por: Ming, Yifei, et al.
Publicado: (2025)
MAS-ZERO: Designing Multi-Agent Systems with Zero Supervision
por: Ke, Zixuan, et al.
Publicado: (2025)
por: Ke, Zixuan, et al.
Publicado: (2025)
Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms
por: Pandit, Shrey, et al.
Publicado: (2025)
por: Pandit, Shrey, et al.
Publicado: (2025)
StructTest: Benchmarking LLMs' Reasoning through Compositional Structured Outputs
por: Chen, Hailin, et al.
Publicado: (2024)
por: Chen, Hailin, et al.
Publicado: (2024)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
por: Xu, Austin, et al.
Publicado: (2025)
por: Xu, Austin, et al.
Publicado: (2025)
SFR-RAG: Towards Contextually Faithful LLMs
por: Nguyen, Xuan-Phi, et al.
Publicado: (2024)
por: Nguyen, Xuan-Phi, et al.
Publicado: (2024)
MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
por: Ke, Zixuan, et al.
Publicado: (2026)
por: Ke, Zixuan, et al.
Publicado: (2026)
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
por: Xu, Austin, et al.
Publicado: (2025)
por: Xu, Austin, et al.
Publicado: (2025)
Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts
por: Nguyen, Xuan-Phi, et al.
Publicado: (2026)
por: Nguyen, Xuan-Phi, et al.
Publicado: (2026)
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
por: Ming, Yifei, et al.
Publicado: (2024)
por: Ming, Yifei, et al.
Publicado: (2024)
Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing
por: Jiao, Fangkai, et al.
Publicado: (2024)
por: Jiao, Fangkai, et al.
Publicado: (2024)
Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning
por: Wang, Jiayu, et al.
Publicado: (2025)
por: Wang, Jiayu, et al.
Publicado: (2025)
Direct Judgement Preference Optimization
por: Wang, Peifeng, et al.
Publicado: (2024)
por: Wang, Peifeng, et al.
Publicado: (2024)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
por: Zhou, Yilun, et al.
Publicado: (2025)
por: Zhou, Yilun, et al.
Publicado: (2025)
A Comprehensive Survey of Contamination Detection Methods in Large Language Models
por: Ravaut, Mathieu, et al.
Publicado: (2024)
por: Ravaut, Mathieu, et al.
Publicado: (2024)
ChatGPT's One-year Anniversary: Are Open-Source Large Language Models Catching up?
por: Chen, Hailin, et al.
Publicado: (2023)
por: Chen, Hailin, et al.
Publicado: (2023)
Relevant or Random: Can LLMs Truly Perform Analogical Reasoning?
por: Qin, Chengwei, et al.
Publicado: (2024)
por: Qin, Chengwei, et al.
Publicado: (2024)
Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction
por: Shi, Zhenmei, et al.
Publicado: (2024)
por: Shi, Zhenmei, et al.
Publicado: (2024)
NAACL2025 Tutorial: Adaptation of Large Language Models
por: Ke, Zixuan, et al.
Publicado: (2025)
por: Ke, Zixuan, et al.
Publicado: (2025)
ParaICL: Towards Parallel In-Context Learning
por: Li, Xingxuan, et al.
Publicado: (2024)
por: Li, Xingxuan, et al.
Publicado: (2024)
Preference Optimization for Reasoning with Pseudo Feedback
por: Jiao, Fangkai, et al.
Publicado: (2024)
por: Jiao, Fangkai, et al.
Publicado: (2024)
Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse Prompts
por: Nguyen, Xuan-Phi, et al.
Publicado: (2023)
por: Nguyen, Xuan-Phi, et al.
Publicado: (2023)
Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
por: Pang, Bo, et al.
Publicado: (2025)
por: Pang, Bo, et al.
Publicado: (2025)
CodeXEmbed: A Generalist Embedding Model Family for Multiligual and Multi-task Code Retrieval
por: Liu, Ye, et al.
Publicado: (2024)
por: Liu, Ye, et al.
Publicado: (2024)
Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks
por: Li, Xingxuan, et al.
Publicado: (2024)
por: Li, Xingxuan, et al.
Publicado: (2024)
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
por: Panagopoulou, Artemis, et al.
Publicado: (2023)
por: Panagopoulou, Artemis, et al.
Publicado: (2023)
JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking
por: Niu, Tong, et al.
Publicado: (2024)
por: Niu, Tong, et al.
Publicado: (2024)
VIBEPASS: Can Vibe Coders Really Pass the Vibe Check?
por: Bansal, Srijan, et al.
Publicado: (2026)
por: Bansal, Srijan, et al.
Publicado: (2026)
Beyond Output Matching: Bidirectional Alignment for Enhanced In-Context Learning
por: Qin, Chengwei, et al.
Publicado: (2023)
por: Qin, Chengwei, et al.
Publicado: (2023)
LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
por: Wang, Jiayu, et al.
Publicado: (2025)
por: Wang, Jiayu, et al.
Publicado: (2025)
SweRank: Software Issue Localization with Code Ranking
por: Reddy, Revanth Gangi, et al.
Publicado: (2025)
por: Reddy, Revanth Gangi, et al.
Publicado: (2025)
SkillOrchestra: Learning to Route Agents via Skill Transfer
por: Wang, Jiayu, et al.
Publicado: (2026)
por: Wang, Jiayu, et al.
Publicado: (2026)
Enabling High Data Throughput Reinforcement Learning on GPUs: A Domain Agnostic Framework for Data-Driven Scientific Research
por: Lan, Tian, et al.
Publicado: (2024)
por: Lan, Tian, et al.
Publicado: (2024)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
por: Liao, Baohao, et al.
Publicado: (2025)
por: Liao, Baohao, et al.
Publicado: (2025)
Exploring Self-supervised Logic-enhanced Training for Large Language Models
por: Jiao, Fangkai, et al.
Publicado: (2023)
por: Jiao, Fangkai, et al.
Publicado: (2023)
Lifelong Event Detection with Embedding Space Separation and Compaction
por: Qin, Chengwei, et al.
Publicado: (2024)
por: Qin, Chengwei, et al.
Publicado: (2024)
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
por: Zhou, Honglu, et al.
Publicado: (2025)
por: Zhou, Honglu, et al.
Publicado: (2025)
Ejemplares similares
-
Demystifying Domain-adaptive Post-training for Financial LLMs
por: Ke, Zixuan, et al.
Publicado: (2025) -
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
por: Nguyen, Xuan-Phi, et al.
Publicado: (2025) -
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
por: Pandit, Shrey, et al.
Publicado: (2025) -
Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows
por: Ming, Yifei, et al.
Publicado: (2025) -
MAS-ZERO: Designing Multi-Agent Systems with Zero Supervision
por: Ke, Zixuan, et al.
Publicado: (2025)