Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Geiping, Jonas, McLeish, Sean, Jain, Neel, Kirchenbauer, John, Singh, Siddharth, Bartoldson, Brian R., Kailkhura, Bhavya, Bhatele, Abhinav, Goldstein, Tom |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transformers Can Do Arithmetic with the Right Embeddings
von: McLeish, Sean, et al.
Veröffentlicht: (2024)
von: McLeish, Sean, et al.
Veröffentlicht: (2024)
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
Gemstones: A Model Suite for Multi-Faceted Scaling Laws
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
Benchmarking ChatGPT on Algorithmic Reasoning
von: McLeish, Sean, et al.
Veröffentlicht: (2024)
von: McLeish, Sean, et al.
Veröffentlicht: (2024)
Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference
von: Lee, Sangyun, et al.
Veröffentlicht: (2026)
von: Lee, Sangyun, et al.
Veröffentlicht: (2026)
Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
von: Hans, Abhimanyu, et al.
Veröffentlicht: (2024)
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies
von: Bartoldson, Brian R., et al.
Veröffentlicht: (2024)
von: Bartoldson, Brian R., et al.
Veröffentlicht: (2024)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security
von: Cai, Zikui, et al.
Veröffentlicht: (2025)
von: Cai, Zikui, et al.
Veröffentlicht: (2025)
Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
von: McDonald, Tavish, et al.
Veröffentlicht: (2025)
von: McDonald, Tavish, et al.
Veröffentlicht: (2025)
Theoretical rheo-physics of silk: Intermolecular associations reduce the critical specific work for flow-induced crystallisation
von: Schaefer, Charley, et al.
Veröffentlicht: (2021)
von: Schaefer, Charley, et al.
Veröffentlicht: (2021)
Multi-Token Prediction via Self-Distillation
von: Kirchenbauer, John, et al.
Veröffentlicht: (2026)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2026)
A Watermark for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
von: Zheng, Haizhong, et al.
Veröffentlicht: (2025)
von: Zheng, Haizhong, et al.
Veröffentlicht: (2025)
The CLRS-Text Algorithmic Reasoning Language Benchmark
von: Markeeva, Larisa, et al.
Veröffentlicht: (2024)
von: Markeeva, Larisa, et al.
Veröffentlicht: (2024)
Efficient Parallel Samplers for Recurrent-Depth Models and Their Connection to Diffusion Language Models
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
von: Geiping, Jonas, et al.
Veröffentlicht: (2025)
LMD3: Language Model Data Density Dependence
von: Kirchenbauer, John, et al.
Veröffentlicht: (2024)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2024)
HPC-Coder-V2: Studying Code LLMs Across Low-Resource Parallel Languages
von: Chaturvedi, Aman, et al.
Veröffentlicht: (2024)
von: Chaturvedi, Aman, et al.
Veröffentlicht: (2024)
GenQA: Generating Millions of Instructions from a Handful of Prompts
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
Zero-Shot Vision Encoder Grafting via LLM Surrogates
von: Yue, Kaiyu, et al.
Veröffentlicht: (2025)
von: Yue, Kaiyu, et al.
Veröffentlicht: (2025)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
von: Singh, Siddharth, et al.
Veröffentlicht: (2023)
von: Singh, Siddharth, et al.
Veröffentlicht: (2023)
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
von: Wang, Zijun, et al.
Veröffentlicht: (2025)
von: Wang, Zijun, et al.
Veröffentlicht: (2025)
Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion
von: Christopher, Jacob K, et al.
Veröffentlicht: (2024)
von: Christopher, Jacob K, et al.
Veröffentlicht: (2024)
ELFS: Label-Free Coreset Selection with Proxy Training Dynamics
von: Zheng, Haizhong, et al.
Veröffentlicht: (2024)
von: Zheng, Haizhong, et al.
Veröffentlicht: (2024)
Speculating Experts Accelerates Inference for Mixture-of-Experts
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
von: Ranjan, Aditya K., et al.
Veröffentlicht: (2025)
von: Ranjan, Aditya K., et al.
Veröffentlicht: (2025)
The Big Send-off: Scalable and Performant Collectives for Deep Learning
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
When Can You Get Away with Low Memory Adam?
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2025)
von: Kalra, Dayal Singh, et al.
Veröffentlicht: (2025)
Loki: Low-rank Keys for Efficient Sparse Attention
von: Singhania, Prajwal, et al.
Veröffentlicht: (2024)
von: Singhania, Prajwal, et al.
Veröffentlicht: (2024)
Modeling Code: Is Text All You Need?
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
von: Nichols, Daniel, et al.
Veröffentlicht: (2025)
On the Reliability of Watermarks for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
What do we learn from inverting CLIP models?
von: Kazemi, Hamid, et al.
Veröffentlicht: (2024)
von: Kazemi, Hamid, et al.
Veröffentlicht: (2024)
Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training
von: Bartoldson, Brian, et al.
Veröffentlicht: (2025)
von: Bartoldson, Brian, et al.
Veröffentlicht: (2025)
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
von: Knupp, Jonas, et al.
Veröffentlicht: (2026)
von: Knupp, Jonas, et al.
Veröffentlicht: (2026)
Measuring Style Similarity in Diffusion Models
von: Somepalli, Gowthami, et al.
Veröffentlicht: (2024)
von: Somepalli, Gowthami, et al.
Veröffentlicht: (2024)
Exploiting Sparsity for Long Context Inference: Million Token Contexts on Commodity GPUs
von: Synk, Ryan, et al.
Veröffentlicht: (2025)
von: Synk, Ryan, et al.
Veröffentlicht: (2025)
Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
von: Venkatraman, Siddarth, et al.
Veröffentlicht: (2025)
von: Venkatraman, Siddarth, et al.
Veröffentlicht: (2025)
Internalizing symptoms and affective vulnerability among heterosexual and sexual minority young adults
von: Alison C. McLeish, et al.
Veröffentlicht: (2024)
von: Alison C. McLeish, et al.
Veröffentlicht: (2024)
Scaling Open-Ended Reasoning to Predict the Future
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Transformers Can Do Arithmetic with the Right Embeddings
von: McLeish, Sean, et al.
Veröffentlicht: (2024) -
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
von: McLeish, Sean, et al.
Veröffentlicht: (2025) -
Gemstones: A Model Suite for Multi-Faceted Scaling Laws
von: McLeish, Sean, et al.
Veröffentlicht: (2025) -
Benchmarking ChatGPT on Algorithmic Reasoning
von: McLeish, Sean, et al.
Veröffentlicht: (2024) -
Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference
von: Lee, Sangyun, et al.
Veröffentlicht: (2026)