PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving
Fuente:
arXiv
Salvato in:
| Autori principali: | Parmar, Mihir, Goyal, Palash, Liu, Xin, Song, Yiwen, Ling, Mingyang, Baral, Chitta, Palangi, Hamid, Pfister, Tomas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review
di: Goyal, Palash, et al.
Pubblicazione: (2026)
di: Goyal, Palash, et al.
Pubblicazione: (2026)
PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem Solving
di: Parmar, Mihir, et al.
Pubblicazione: (2025)
di: Parmar, Mihir, et al.
Pubblicazione: (2025)
LLM-Based Multi-Agent Blackboard System for Information Discovery in Data Science
di: Salemi, Alireza, et al.
Pubblicazione: (2025)
di: Salemi, Alireza, et al.
Pubblicazione: (2025)
HEART: Emotionally-Driven Test-Time Scaling of Language Models
di: Pinto, Gabriela, et al.
Pubblicazione: (2025)
di: Pinto, Gabriela, et al.
Pubblicazione: (2025)
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
di: Tyagi, Nemika, et al.
Pubblicazione: (2024)
di: Tyagi, Nemika, et al.
Pubblicazione: (2024)
GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
di: Handa, Divij, et al.
Pubblicazione: (2025)
di: Handa, Divij, et al.
Pubblicazione: (2025)
Watch and Learn: Learning to Use Computers from Online Videos
di: Song, Chan Hee, et al.
Pubblicazione: (2025)
di: Song, Chan Hee, et al.
Pubblicazione: (2025)
TFRBench: A Reasoning Benchmark for Evaluating Forecasting Systems
di: Ahamed, Md Atik, et al.
Pubblicazione: (2026)
di: Ahamed, Md Atik, et al.
Pubblicazione: (2026)
Synapse: Adaptive Arbitration of Complementary Expertise in Time Series Foundational Models
di: Das, Sarkar Snigdha Sarathi, et al.
Pubblicazione: (2025)
di: Das, Sarkar Snigdha Sarathi, et al.
Pubblicazione: (2025)
Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning
di: Mishra, Venkatesh, et al.
Pubblicazione: (2025)
di: Mishra, Venkatesh, et al.
Pubblicazione: (2025)
CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding
di: Mondal, Ishani, et al.
Pubblicazione: (2026)
di: Mondal, Ishani, et al.
Pubblicazione: (2026)
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
di: Patel, Nisarg, et al.
Pubblicazione: (2024)
di: Patel, Nisarg, et al.
Pubblicazione: (2024)
Reasoning-Aware Training for Time Series Forecasting
di: Ahamed, Md Atik, et al.
Pubblicazione: (2026)
di: Ahamed, Md Atik, et al.
Pubblicazione: (2026)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
di: Parmar, Mihir, et al.
Pubblicazione: (2022)
di: Parmar, Mihir, et al.
Pubblicazione: (2022)
VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
di: Miculicich, Lesly, et al.
Pubblicazione: (2025)
di: Miculicich, Lesly, et al.
Pubblicazione: (2025)
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models
di: RRV, Aswin, et al.
Pubblicazione: (2026)
di: RRV, Aswin, et al.
Pubblicazione: (2026)
LEAF: A Living Benchmark for Event-Augmented Forecasting
di: Tan, Mingtian, et al.
Pubblicazione: (2026)
di: Tan, Mingtian, et al.
Pubblicazione: (2026)
PHANTOM RECALL: When Familiar Puzzles Fool Smart Models
di: Mukhopadhyay, Souradeep, et al.
Pubblicazione: (2025)
di: Mukhopadhyay, Souradeep, et al.
Pubblicazione: (2025)
Triple Preference Optimization: Achieving Better Alignment using a Single Step Optimization
di: Saeidi, Amir, et al.
Pubblicazione: (2024)
di: Saeidi, Amir, et al.
Pubblicazione: (2024)
ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
di: Meng, Rui, et al.
Pubblicazione: (2026)
di: Meng, Rui, et al.
Pubblicazione: (2026)
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
di: Parmar, Mihir, et al.
Pubblicazione: (2024)
di: Parmar, Mihir, et al.
Pubblicazione: (2024)
Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
di: Luo, Man, et al.
Pubblicazione: (2023)
di: Luo, Man, et al.
Pubblicazione: (2023)
Nexus : An Agentic Framework for Time Series Forecasting
di: Das, Sarkar Snigdha Sarathi, et al.
Pubblicazione: (2026)
di: Das, Sarkar Snigdha Sarathi, et al.
Pubblicazione: (2026)
Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark
di: Gupta, Himanshu, et al.
Pubblicazione: (2024)
di: Gupta, Himanshu, et al.
Pubblicazione: (2024)
TarGEN: Targeted Data Generation with Large Language Models
di: Gupta, Himanshu, et al.
Pubblicazione: (2023)
di: Gupta, Himanshu, et al.
Pubblicazione: (2023)
Heterogeneous Swarms: Jointly Optimizing Model Roles and Weights for Multi-LLM Systems
di: Feng, Shangbin, et al.
Pubblicazione: (2025)
di: Feng, Shangbin, et al.
Pubblicazione: (2025)
Exploring Group and Symmetry Principles in Large Language Models
di: Imani, Shima, et al.
Pubblicazione: (2024)
di: Imani, Shima, et al.
Pubblicazione: (2024)
ThinkTuning: Instilling Cognitive Reflections without Distillation
di: RRV, Aswin, et al.
Pubblicazione: (2025)
di: RRV, Aswin, et al.
Pubblicazione: (2025)
VQQA: An Agentic Approach for Video Evaluation and Quality Improvement
di: Song, Yiwen, et al.
Pubblicazione: (2026)
di: Song, Yiwen, et al.
Pubblicazione: (2026)
Investigating VLM Hallucination from a Cognitive Psychology Perspective: A First Step Toward Interpretation with Intriguing Observations
di: Liu, Xiangrui, et al.
Pubblicazione: (2025)
di: Liu, Xiangrui, et al.
Pubblicazione: (2025)
Diffusion Adversarial Post-Training for One-Step Video Generation
di: Lin, Shanchuan, et al.
Pubblicazione: (2025)
di: Lin, Shanchuan, et al.
Pubblicazione: (2025)
PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing
di: Song, Yiwen, et al.
Pubblicazione: (2026)
di: Song, Yiwen, et al.
Pubblicazione: (2026)
SOD: Step-wise On-policy Distillation for Small Language Model Agents
di: Zhong, Qiyong, et al.
Pubblicazione: (2026)
di: Zhong, Qiyong, et al.
Pubblicazione: (2026)
On the Role of Feedback in Test-Time Scaling of Agentic AI Workflows
di: Chakraborty, Souradip, et al.
Pubblicazione: (2025)
di: Chakraborty, Souradip, et al.
Pubblicazione: (2025)
CoSIGN: Few-Step Guidance of ConSIstency Model to Solve General INverse Problems
di: Zhao, Jiankun, et al.
Pubblicazione: (2024)
di: Zhao, Jiankun, et al.
Pubblicazione: (2024)
Step-by-Step Training for Media Technicians
di: Kranch, Doug
Pubblicazione: (1976)
di: Kranch, Doug
Pubblicazione: (1976)
Two-Step Quantum Search Algorithm for Solving Traveling Salesman Problems
di: Sato, Rei, et al.
Pubblicazione: (2024)
di: Sato, Rei, et al.
Pubblicazione: (2024)
Step-Wise Formal Verification for LLM-Based Mathematical Problem Solving
di: Zhou, Kuo, et al.
Pubblicazione: (2025)
di: Zhou, Kuo, et al.
Pubblicazione: (2025)
On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024)
di: Chatterjee, Agneet, et al.
Pubblicazione: (2024)
CRISP: Complex Reasoning with Interpretable Step-based Plans
di: Vetzler, Matan, et al.
Pubblicazione: (2025)
di: Vetzler, Matan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review
di: Goyal, Palash, et al.
Pubblicazione: (2026) -
PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem Solving
di: Parmar, Mihir, et al.
Pubblicazione: (2025) -
LLM-Based Multi-Agent Blackboard System for Information Discovery in Data Science
di: Salemi, Alireza, et al.
Pubblicazione: (2025) -
HEART: Emotionally-Driven Test-Time Scaling of Language Models
di: Pinto, Gabriela, et al.
Pubblicazione: (2025) -
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
di: Tyagi, Nemika, et al.
Pubblicazione: (2024)