Rethinking Thinking Tokens: LLMs as Improvement Operators
Fuente:
arXiv
Guardado en:
| Autores principales: | Madaan, Lovish, Didolkar, Aniket, Gururangan, Suchin, Quan, John, Silva, Ruan, Salakhutdinov, Ruslan, Zaheer, Manzil, Arora, Sanjeev, Goyal, Anirudh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
por: Didolkar, Aniket, et al.
Publicado: (2025)
por: Didolkar, Aniket, et al.
Publicado: (2025)
LESS: Selecting Influential Data for Targeted Instruction Tuning
por: Xia, Mengzhou, et al.
Publicado: (2024)
por: Xia, Mengzhou, et al.
Publicado: (2024)
Metacognitive Capabilities of LLMs: An Exploration in Mathematical Problem Solving
por: Didolkar, Aniket, et al.
Publicado: (2024)
por: Didolkar, Aniket, et al.
Publicado: (2024)
The Art of Scaling Reinforcement Learning Compute for LLMs
por: Khatri, Devvrit, et al.
Publicado: (2025)
por: Khatri, Devvrit, et al.
Publicado: (2025)
Learning Beyond Pattern Matching? Assaying Mathematical Understanding in LLMs
por: Guo, Siyuan, et al.
Publicado: (2024)
por: Guo, Siyuan, et al.
Publicado: (2024)
Scaling Test-Time Compute for Agentic Coding
por: Kim, Joongwon, et al.
Publicado: (2026)
por: Kim, Joongwon, et al.
Publicado: (2026)
Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
por: Lyu, Kaifeng, et al.
Publicado: (2024)
por: Lyu, Kaifeng, et al.
Publicado: (2024)
Can Models Learn Skill Composition from Examples?
por: Zhao, Haoyu, et al.
Publicado: (2024)
por: Zhao, Haoyu, et al.
Publicado: (2024)
Differentially Private Model Merging
por: Yin, Qichuan, et al.
Publicado: (2026)
por: Yin, Qichuan, et al.
Publicado: (2026)
Masked Generative Priors Improve World Models Sequence Modelling Capabilities
por: Meo, Cristian, et al.
Publicado: (2024)
por: Meo, Cristian, et al.
Publicado: (2024)
A Statistical Framework for Data-dependent Retrieval-Augmented Models
por: Basu, Soumya, et al.
Publicado: (2024)
por: Basu, Soumya, et al.
Publicado: (2024)
Federation over Text: Insight Sharing for Multi-Agent Reasoning
por: Yao, Dixi, et al.
Publicado: (2026)
por: Yao, Dixi, et al.
Publicado: (2026)
Time is Encoded in the Weights of Finetuned Language Models
por: Nylund, Kai, et al.
Publicado: (2023)
por: Nylund, Kai, et al.
Publicado: (2023)
Contrastive Difference Predictive Coding
por: Zheng, Chongyi, et al.
Publicado: (2023)
por: Zheng, Chongyi, et al.
Publicado: (2023)
Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
por: Jayalath, Dulhan, et al.
Publicado: (2025)
por: Jayalath, Dulhan, et al.
Publicado: (2025)
Zero-Shot Object-Centric Representation Learning
por: Didolkar, Aniket, et al.
Publicado: (2024)
por: Didolkar, Aniket, et al.
Publicado: (2024)
On the Impossibility of Retrain Equivalence in Machine Unlearning
por: Yu, Jiatong, et al.
Publicado: (2025)
por: Yu, Jiatong, et al.
Publicado: (2025)
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
por: Kaur, Simran, et al.
Publicado: (2024)
por: Kaur, Simran, et al.
Publicado: (2024)
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore
por: Min, Sewon, et al.
Publicado: (2023)
por: Min, Sewon, et al.
Publicado: (2023)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
por: Huang, Yukun, et al.
Publicado: (2025)
por: Huang, Yukun, et al.
Publicado: (2025)
Unlearning via Sparse Representations
por: Shah, Vedant, et al.
Publicado: (2023)
por: Shah, Vedant, et al.
Publicado: (2023)
InSTA: Towards Internet-Scale Training For Agents
por: Trabucco, Brandon, et al.
Publicado: (2025)
por: Trabucco, Brandon, et al.
Publicado: (2025)
Diversity-driven Data Selection for Language Model Tuning through Sparse Autoencoder
por: Yang, Xianjun, et al.
Publicado: (2025)
por: Yang, Xianjun, et al.
Publicado: (2025)
Effective Data Augmentation With Diffusion Models
por: Trabucco, Brandon, et al.
Publicado: (2023)
por: Trabucco, Brandon, et al.
Publicado: (2023)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
por: Qu, Yuxiao, et al.
Publicado: (2025)
por: Qu, Yuxiao, et al.
Publicado: (2025)
Understanding Visual Concepts Across Models
por: Trabucco, Brandon, et al.
Publicado: (2024)
por: Trabucco, Brandon, et al.
Publicado: (2024)
CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
por: Didolkar, Aniket, et al.
Publicado: (2025)
por: Didolkar, Aniket, et al.
Publicado: (2025)
Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks
por: Dalal, Murtaza, et al.
Publicado: (2024)
por: Dalal, Murtaza, et al.
Publicado: (2024)
POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration
por: Qu, Yuxiao, et al.
Publicado: (2026)
por: Qu, Yuxiao, et al.
Publicado: (2026)
Tree Search for Language Model Agents
por: Koh, Jing Yu, et al.
Publicado: (2024)
por: Koh, Jing Yu, et al.
Publicado: (2024)
AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents
por: Zharmagambetov, Arman, et al.
Publicado: (2025)
por: Zharmagambetov, Arman, et al.
Publicado: (2025)
Quantifying Variance in Evaluation Benchmarks
por: Madaan, Lovish, et al.
Publicado: (2024)
por: Madaan, Lovish, et al.
Publicado: (2024)
Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
por: Tang, Yunhao, et al.
Publicado: (2025)
por: Tang, Yunhao, et al.
Publicado: (2025)
Rethinking the Illusion of Thinking
por: Varela, Iñaki Dellibarda, et al.
Publicado: (2025)
por: Varela, Iñaki Dellibarda, et al.
Publicado: (2025)
Think before you speak: Training Language Models With Pause Tokens
por: Goyal, Sachin, et al.
Publicado: (2023)
por: Goyal, Sachin, et al.
Publicado: (2023)
Deep Reinforcement Learning for Sequential Combinatorial Auctions
por: Ravindranath, Sai Srivatsa, et al.
Publicado: (2024)
por: Ravindranath, Sai Srivatsa, et al.
Publicado: (2024)
Efficient Distributed Optimization under Heavy-Tailed Noise
por: Lee, Su Hyeong, et al.
Publicado: (2025)
por: Lee, Su Hyeong, et al.
Publicado: (2025)
STRIVE: A Think & Improve Approach with Iterative Refinement for Enhancing Question Quality Estimation
por: Deroy, Aniket, et al.
Publicado: (2025)
por: Deroy, Aniket, et al.
Publicado: (2025)
AI-Assisted Generation of Difficult Math Questions
por: Shah, Vedant, et al.
Publicado: (2024)
por: Shah, Vedant, et al.
Publicado: (2024)
Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
por: Duan, Xintong, et al.
Publicado: (2025)
por: Duan, Xintong, et al.
Publicado: (2025)
Ejemplares similares
-
Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
por: Didolkar, Aniket, et al.
Publicado: (2025) -
LESS: Selecting Influential Data for Targeted Instruction Tuning
por: Xia, Mengzhou, et al.
Publicado: (2024) -
Metacognitive Capabilities of LLMs: An Exploration in Mathematical Problem Solving
por: Didolkar, Aniket, et al.
Publicado: (2024) -
The Art of Scaling Reinforcement Learning Compute for LLMs
por: Khatri, Devvrit, et al.
Publicado: (2025) -
Learning Beyond Pattern Matching? Assaying Mathematical Understanding in LLMs
por: Guo, Siyuan, et al.
Publicado: (2024)