The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Aghajohari, Milad, Chitsaz, Kamran, Kazemnejad, Amirhossein, Chandar, Sarath, Sordoni, Alessandro, Courville, Aaron, Reddy, Siva |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VinePPO: Refining Credit Assignment in RL Training of LLMs
von: Kazemnejad, Amirhossein, et al.
Veröffentlicht: (2024)
von: Kazemnejad, Amirhossein, et al.
Veröffentlicht: (2024)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
Are self-explanations from Large Language Models faithful?
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
V-STaR: Training Verifiers for Self-Taught Reasoners
von: Hosseini, Arian, et al.
Veröffentlicht: (2024)
von: Hosseini, Arian, et al.
Veröffentlicht: (2024)
LOQA: Learning with Opponent Q-Learning Awareness
von: Aghajohari, Milad, et al.
Veröffentlicht: (2024)
von: Aghajohari, Milad, et al.
Veröffentlicht: (2024)
Faithfulness Measurable Masked Language Models
von: Madsen, Andreas, et al.
Veröffentlicht: (2023)
von: Madsen, Andreas, et al.
Veröffentlicht: (2023)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
Best Response Shaping
von: Aghajohari, Milad, et al.
Veröffentlicht: (2024)
von: Aghajohari, Milad, et al.
Veröffentlicht: (2024)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
Towards Practical Tool Usage for Continually Learning LLMs
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
Exploring Quantization for Efficient Pre-Training of Transformer Language Models
von: Chitsaz, Kamran, et al.
Veröffentlicht: (2024)
von: Chitsaz, Kamran, et al.
Veröffentlicht: (2024)
Guiding Language Model Reasoning with Planning Tokens
von: Wang, Xinyi, et al.
Veröffentlicht: (2023)
von: Wang, Xinyi, et al.
Veröffentlicht: (2023)
Do Large Language Models Know How Much They Know?
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
Why Don't Prompt-Based Fairness Metrics Correlate?
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
von: Zayed, Abdelrahman, et al.
Veröffentlicht: (2023)
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
von: Tan, Shawn, et al.
Veröffentlicht: (2024)
Interpretability Needs a New Paradigm
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
uTeBC-NLP at SemEval-2024 Task 9: Can LLMs be Lateral Thinkers?
von: Sadeghi, Pouya, et al.
Veröffentlicht: (2024)
von: Sadeghi, Pouya, et al.
Veröffentlicht: (2024)
Improving Context-Aware Preference Modeling for Language Models
von: Pitis, Silviu, et al.
Veröffentlicht: (2024)
von: Pitis, Silviu, et al.
Veröffentlicht: (2024)
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
von: Wang, Zengzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zengzhi, et al.
Veröffentlicht: (2025)
Too Big to Fool: Resisting Deception in Language Models
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
von: Samsami, Mohammad Reza, et al.
Veröffentlicht: (2024)
Forgetting Transformer: Softmax Attention with a Forget Gate
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025)
von: Parthasarathi, Prasanna, et al.
Veröffentlicht: (2025)
LightThinker++: From Reasoning Compression to Memory Management
von: Zhu, Yuqi, et al.
Veröffentlicht: (2026)
von: Zhu, Yuqi, et al.
Veröffentlicht: (2026)
Adaptive Computation Pruning for the Forgetting Transformer
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers
von: Sareen, Kusha, et al.
Veröffentlicht: (2025)
von: Sareen, Kusha, et al.
Veröffentlicht: (2025)
Small Encoders Can Rival Large Decoders in Detecting Groundedness
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
EpiK-Eval: Evaluation for Language Models as Epistemic Models
von: Prato, Gabriele, et al.
Veröffentlicht: (2023)
von: Prato, Gabriele, et al.
Veröffentlicht: (2023)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
NovoMolGen: Rethinking Molecular Language Model Pretraining
von: Chitsaz, Kamran, et al.
Veröffentlicht: (2025)
von: Chitsaz, Kamran, et al.
Veröffentlicht: (2025)
Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
von: Guiroy, Simon, et al.
Veröffentlicht: (2025)
von: Guiroy, Simon, et al.
Veröffentlicht: (2025)
Intelligent Switching for Reset-Free RL
von: Patil, Darshan, et al.
Veröffentlicht: (2024)
von: Patil, Darshan, et al.
Veröffentlicht: (2024)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
von: Huang, Jerry, et al.
Veröffentlicht: (2024)
Lookbehind-SAM: k steps back, 1 step forward
von: Mordido, Gonçalo, et al.
Veröffentlicht: (2023)
von: Mordido, Gonçalo, et al.
Veröffentlicht: (2023)
ROSA: Random Subspace Adaptation for Efficient Fine-Tuning
von: Hameed, Marawan Gamal Abdel, et al.
Veröffentlicht: (2024)
von: Hameed, Marawan Gamal Abdel, et al.
Veröffentlicht: (2024)
BRIDGE: Predicting Human Task Completion Time From Model Performance
von: Liu, Fengyuan, et al.
Veröffentlicht: (2026)
von: Liu, Fengyuan, et al.
Veröffentlicht: (2026)
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
von: Anugraha, David, et al.
Veröffentlicht: (2025)
von: Anugraha, David, et al.
Veröffentlicht: (2025)
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
von: Yadav, Prateek, et al.
Veröffentlicht: (2024)
von: Yadav, Prateek, et al.
Veröffentlicht: (2024)
Squeezing More from the Stream : Learning Representation Online for Streaming Reinforcement Learning
von: Nilaksh, et al.
Veröffentlicht: (2026)
von: Nilaksh, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
VinePPO: Refining Credit Assignment in RL Training of LLMs
von: Kazemnejad, Amirhossein, et al.
Veröffentlicht: (2024) -
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
von: Prato, Gabriele, et al.
Veröffentlicht: (2025) -
Are self-explanations from Large Language Models faithful?
von: Madsen, Andreas, et al.
Veröffentlicht: (2024) -
V-STaR: Training Verifiers for Self-Taught Reasoners
von: Hosseini, Arian, et al.
Veröffentlicht: (2024) -
LOQA: Learning with Opponent Q-Learning Awareness
von: Aghajohari, Milad, et al.
Veröffentlicht: (2024)