Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Tzannetos, Georgios, Kamalaruban, Parameswaran, Singla, Adish |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Proximal Curriculum with Task Correlations for Deep Reinforcement Learning
by: Tzannetos, Georgios, et al.
Published: (2024)
by: Tzannetos, Georgios, et al.
Published: (2024)
Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences
by: Nika, Andi, et al.
Published: (2024)
by: Nika, Andi, et al.
Published: (2024)
Informativeness of Reward Functions in Reinforcement Learning
by: Devidze, Rati, et al.
Published: (2024)
by: Devidze, Rati, et al.
Published: (2024)
Learning Embeddings for Sequential Tasks Using Population of Agents
by: Mahajan, Mridul, et al.
Published: (2023)
by: Mahajan, Mridul, et al.
Published: (2023)
Neural Task Synthesis for Visual Programming
by: Pădurean, Victor-Alexandru, et al.
Published: (2023)
by: Pădurean, Victor-Alexandru, et al.
Published: (2023)
Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback
by: Nika, Andi, et al.
Published: (2026)
by: Nika, Andi, et al.
Published: (2026)
Corruption Robust Offline Reinforcement Learning with Human Feedback
by: Mandal, Debmalya, et al.
Published: (2024)
by: Mandal, Debmalya, et al.
Published: (2024)
Inference-Time Personalized Alignment with a Few User Preference Queries
by: Pădurean, Victor-Alexandru, et al.
Published: (2025)
by: Pădurean, Victor-Alexandru, et al.
Published: (2025)
Policy Teaching via Data Poisoning in Learning from Human Preferences
by: Nika, Andi, et al.
Published: (2025)
by: Nika, Andi, et al.
Published: (2025)
Towards Generalizable Agents in Text-Based Educational Environments: A Study of Integrating RL with LLMs
by: Radmehr, Bahar, et al.
Published: (2024)
by: Radmehr, Bahar, et al.
Published: (2024)
MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems
by: Nöther, Jonathan, et al.
Published: (2026)
by: Nöther, Jonathan, et al.
Published: (2026)
Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms
by: Nöther, Jonathan, et al.
Published: (2025)
by: Nöther, Jonathan, et al.
Published: (2025)
Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints
by: Nöther, Jonathan, et al.
Published: (2025)
by: Nöther, Jonathan, et al.
Published: (2025)
Adversarially Robust Decision Transformer
by: Tang, Xiaohang, et al.
Published: (2024)
by: Tang, Xiaohang, et al.
Published: (2024)
Emergent Bias and Fairness in Multi-Agent Decision Systems
by: Madigan, Maeve, et al.
Published: (2025)
by: Madigan, Maeve, et al.
Published: (2025)
Hints-In-Browser: Benchmarking Language Models for Programming Feedback Generation
by: Kotalwar, Nachiket, et al.
Published: (2024)
by: Kotalwar, Nachiket, et al.
Published: (2024)
Activation Steering for Chain-of-Thought Compression
by: Azizi, Seyedarmin, et al.
Published: (2025)
by: Azizi, Seyedarmin, et al.
Published: (2025)
Corruption-Robust Offline Two-Player Zero-Sum Markov Games
by: Nika, Andi, et al.
Published: (2024)
by: Nika, Andi, et al.
Published: (2024)
Evaluating Fairness in Transaction Fraud Models: Fairness Metrics, Bias Audits, and Challenges
by: Kamalaruban, Parameswaran, et al.
Published: (2024)
by: Kamalaruban, Parameswaran, et al.
Published: (2024)
Formal Models of Active Learning from Contrastive Examples
by: Mansouri, Farnam, et al.
Published: (2025)
by: Mansouri, Farnam, et al.
Published: (2025)
Fairness-Aware Low-Rank Adaptation Under Demographic Privacy Constraints
by: Kamalaruban, Parameswaran, et al.
Published: (2025)
by: Kamalaruban, Parameswaran, et al.
Published: (2025)
Reasoning about Reasoning: BAPO Bounds on Chain-of-Thought Token Complexity in LLMs
by: Tomlinson, Kiran, et al.
Published: (2026)
by: Tomlinson, Kiran, et al.
Published: (2026)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
by: Zhang, Xuan, et al.
Published: (2024)
by: Zhang, Xuan, et al.
Published: (2024)
Chain-of-Thought and Compressed Looped Transformers: A Memory-Budget Separation
by: Zhang, Haozhou
Published: (2026)
by: Zhang, Haozhou
Published: (2026)
Demystifying Long Chain-of-Thought Reasoning in LLMs
by: Yeo, Edward, et al.
Published: (2025)
by: Yeo, Edward, et al.
Published: (2025)
Reasoning Steps as Curriculum: Using Depth of Thought as a Difficulty Signal for Tuning LLMs
by: Jung, Jeesu, et al.
Published: (2025)
by: Jung, Jeesu, et al.
Published: (2025)
Learning Half-Spaces from Perturbed Contrastive Examples
by: Ravari, Aryan Alavi Razavi, et al.
Published: (2026)
by: Ravari, Aryan Alavi Razavi, et al.
Published: (2026)
HMPO: Hybrid Median-length Policy Optimization for Chain-of-Thought Compression
by: Zheng, Minghui, et al.
Published: (2026)
by: Zheng, Minghui, et al.
Published: (2026)
Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO
by: Yu, Bowen, et al.
Published: (2026)
by: Yu, Bowen, et al.
Published: (2026)
Optimal Decision Making Under Strategic Behavior
by: Tsirtsis, Stratis, et al.
Published: (2019)
by: Tsirtsis, Stratis, et al.
Published: (2019)
DarkMind: Latent Chain-of-Thought Backdoor in Customized LLMs
by: Guo, Zhen, et al.
Published: (2025)
by: Guo, Zhen, et al.
Published: (2025)
DMAP: A Distribution Map for Text
by: Kempton, Tom, et al.
Published: (2026)
by: Kempton, Tom, et al.
Published: (2026)
TokenSqueeze: Performance-Preserving Compression for Reasoning LLMs
by: Zhang, Yuxiang, et al.
Published: (2025)
by: Zhang, Yuxiang, et al.
Published: (2025)
Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression
by: Tang, Yuntian, et al.
Published: (2026)
by: Tang, Yuntian, et al.
Published: (2026)
Long Chain-of-Thought Compression via Fine-Grained Group Policy Optimization
by: Han, Xinchen, et al.
Published: (2026)
by: Han, Xinchen, et al.
Published: (2026)
ExpThink: Experience-Guided Reinforcement Learning for Adaptive Chain-of-Thought Compression
by: Bian, Tingcheng, et al.
Published: (2026)
by: Bian, Tingcheng, et al.
Published: (2026)
Understanding Chain-of-Thought in LLMs through Information Theory
by: Ton, Jean-Francois, et al.
Published: (2024)
by: Ton, Jean-Francois, et al.
Published: (2024)
Dialogue Ontology Relation Extraction via Constrained Chain-of-Thought Decoding
by: Vukovic, Renato, et al.
Published: (2024)
by: Vukovic, Renato, et al.
Published: (2024)
e-COP : Episodic Constrained Optimization of Policies
by: Agnihotri, Akhil, et al.
Published: (2024)
by: Agnihotri, Akhil, et al.
Published: (2024)
Reinforcement Learning for Chain of Thought Compression with One-Domain-to-All Generalization
by: Li, Hanyu, et al.
Published: (2025)
by: Li, Hanyu, et al.
Published: (2025)
Similar Items
-
Proximal Curriculum with Task Correlations for Deep Reinforcement Learning
by: Tzannetos, Georgios, et al.
Published: (2024) -
Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences
by: Nika, Andi, et al.
Published: (2024) -
Informativeness of Reward Functions in Reinforcement Learning
by: Devidze, Rati, et al.
Published: (2024) -
Learning Embeddings for Sequential Tasks Using Population of Agents
by: Mahajan, Mridul, et al.
Published: (2023) -
Neural Task Synthesis for Visual Programming
by: Pădurean, Victor-Alexandru, et al.
Published: (2023)