Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | McLeish, Sean, Li, Ang, Kirchenbauer, John, Kalra, Dayal Singh, Bartoldson, Brian R., Kailkhura, Bhavya, Schwarzschild, Avi, Geiping, Jonas, Goldstein, Tom, Goldblum, Micah |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
Transformers Can Do Arithmetic with the Right Embeddings
von: McLeish, Sean, et al.
Veröffentlicht: (2024)
von: McLeish, Sean, et al.
Veröffentlicht: (2024)
Benchmarking ChatGPT on Algorithmic Reasoning
von: McLeish, Sean, et al.
Veröffentlicht: (2024)
von: McLeish, Sean, et al.
Veröffentlicht: (2024)
Gemstones: A Model Suite for Multi-Faceted Scaling Laws
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
Multi-Token Prediction via Self-Distillation
von: Kirchenbauer, John, et al.
Veröffentlicht: (2026)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2026)
Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference
von: Lee, Sangyun, et al.
Veröffentlicht: (2026)
von: Lee, Sangyun, et al.
Veröffentlicht: (2026)
Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
When Can You Get Away with Low Memory Adam?
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2025)
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2025)
The CLRS-Text Algorithmic Reasoning Language Benchmark
von: Markeeva, Larisa, et al.
Veröffentlicht: (2024)
von: Markeeva, Larisa, et al.
Veröffentlicht: (2024)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
On the Reliability of Watermarks for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies
von: Bartoldson, Brian R., et al.
Veröffentlicht: (2024)
von: Bartoldson, Brian R., et al.
Veröffentlicht: (2024)
Theoretical rheo-physics of silk: Intermolecular associations reduce the critical specific work for flow-induced crystallisation
von: Schaefer, Charley, et al.
Veröffentlicht: (2021)
von: Schaefer, Charley, et al.
Veröffentlicht: (2021)
AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security
von: Cai, Zikui, et al.
Veröffentlicht: (2025)
von: Cai, Zikui, et al.
Veröffentlicht: (2025)
A Watermark for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
Measuring Style Similarity in Diffusion Models
von: Somepalli, Gowthami, et al.
Veröffentlicht: (2024)
von: Somepalli, Gowthami, et al.
Veröffentlicht: (2024)
Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
von: McDonald, Tavish, et al.
Veröffentlicht: (2025)
von: McDonald, Tavish, et al.
Veröffentlicht: (2025)
Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
Identifying and Evaluating Inactive Heads in Pretrained LLMs
von: Sandoval-Segura, Pedro, et al.
Veröffentlicht: (2025)
von: Sandoval-Segura, Pedro, et al.
Veröffentlicht: (2025)
Has My System Prompt Been Used? Large Language Model Prompt Membership Inference
von: Levin, Roman, et al.
Veröffentlicht: (2025)
von: Levin, Roman, et al.
Veröffentlicht: (2025)
LMD3: Language Model Data Density Dependence
von: Kirchenbauer, John, et al.
Veröffentlicht: (2024)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2024)
Generating Potent Poisons and Backdoors from Scratch with Guided Diffusion
von: Souri, Hossein, et al.
Veröffentlicht: (2024)
von: Souri, Hossein, et al.
Veröffentlicht: (2024)
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
von: Zheng, Haizhong, et al.
Veröffentlicht: (2025)
von: Zheng, Haizhong, et al.
Veröffentlicht: (2025)
Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion
von: Christopher, Jacob K, et al.
Veröffentlicht: (2024)
von: Christopher, Jacob K, et al.
Veröffentlicht: (2024)
ELFS: Label-Free Coreset Selection with Proxy Training Dynamics
von: Zheng, Haizhong, et al.
Veröffentlicht: (2024)
von: Zheng, Haizhong, et al.
Veröffentlicht: (2024)
Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
von: Venkatraman, Siddarth, et al.
Veröffentlicht: (2025)
von: Venkatraman, Siddarth, et al.
Veröffentlicht: (2025)
What do we learn from inverting CLIP models?
von: Kazemi, Hamid, et al.
Veröffentlicht: (2024)
von: Kazemi, Hamid, et al.
Veröffentlicht: (2024)
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
von: Hayes, Kevin David, et al.
Veröffentlicht: (2025)
von: Hayes, Kevin David, et al.
Veröffentlicht: (2025)
Efficient Parallel Samplers for Recurrent-Depth Models and Their Connection to Diffusion Language Models
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
Thinking Deeper, Not Longer: Depth-Recurrent Transformers for Compositional Generalization
von: Chen, Hung-Hsuan
Veröffentlicht: (2026)
von: Chen, Hung-Hsuan
Veröffentlicht: (2026)
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
von: Wang, Zijun, et al.
Veröffentlicht: (2025)
von: Wang, Zijun, et al.
Veröffentlicht: (2025)
Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
Internalizing symptoms and affective vulnerability among heterosexual and sexual minority young adults
von: Alison C. McLeish, et al.
Veröffentlicht: (2024)
von: Alison C. McLeish, et al.
Veröffentlicht: (2024)
Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained Models
von: Wen, Yuxin, et al.
Veröffentlicht: (2024)
von: Wen, Yuxin, et al.
Veröffentlicht: (2024)
Coercing LLMs to do and reveal (almost) anything
von: Geiping, Jonas, et al.
Veröffentlicht: (2024)
von: Geiping, Jonas, et al.
Veröffentlicht: (2024)
FictionalQA: A Dataset for Studying Memorization and Knowledge Acquisition
von: Kirchenbauer, John, et al.
Veröffentlicht: (2025)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2025)
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training
von: Bartoldson, Brian, et al.
Veröffentlicht: (2025)
von: Bartoldson, Brian, et al.
Veröffentlicht: (2025)
Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2026)
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2026)
Why Warmup the Learning Rate? Underlying Mechanisms and Improvements
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2024)
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
von: Geiping, Jonas, et al.
Veröffentlicht: (2025) -
Transformers Can Do Arithmetic with the Right Embeddings
von: McLeish, Sean, et al.
Veröffentlicht: (2024) -
Benchmarking ChatGPT on Algorithmic Reasoning
von: McLeish, Sean, et al.
Veröffentlicht: (2024) -
Gemstones: A Model Suite for Multi-Faceted Scaling Laws
von: McLeish, Sean, et al.
Veröffentlicht: (2025) -
Multi-Token Prediction via Self-Distillation
von: Kirchenbauer, John, et al.
Veröffentlicht: (2026)