Thinking into the Future: Latent Lookahead Training for Transformers
Fuente:
arXiv
Guardado en:
| Autores principales: | Noci, Lorenzo, Bachmann, Gregor, Moosavi-Dezfooli, Seyed-Mohsen, Nabi, Moin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Potential of CoT for Reasoning: A Closer Look at Trace Dynamics
por: Bachmann, Gregor, et al.
Publicado: (2026)
por: Bachmann, Gregor, et al.
Publicado: (2026)
Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models
por: Klein, Tassilo, et al.
Publicado: (2024)
por: Klein, Tassilo, et al.
Publicado: (2024)
Rewriting the Budget: A General Framework for Black-Box Attacks Under Cost Asymmetry
por: Salmani, Mahdi, et al.
Publicado: (2025)
por: Salmani, Mahdi, et al.
Publicado: (2025)
Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
por: Movahedi, Sajad, et al.
Publicado: (2024)
por: Movahedi, Sajad, et al.
Publicado: (2024)
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
por: Dong, Yihe, et al.
Publicado: (2025)
por: Dong, Yihe, et al.
Publicado: (2025)
On the Anisotropy of Score-Based Generative Models
por: Floros, Andreas, et al.
Publicado: (2025)
por: Floros, Andreas, et al.
Publicado: (2025)
Revisiting DeepFool: generalization and improvement
por: Abdollahpoorrostam, Alireza, et al.
Publicado: (2023)
por: Abdollahpoorrostam, Alireza, et al.
Publicado: (2023)
Trustworthy Image Super-Resolution via Generative Pseudoinverse
por: Floros, Andreas, et al.
Publicado: (2025)
por: Floros, Andreas, et al.
Publicado: (2025)
Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers
por: Anagnostidis, Sotiris, et al.
Publicado: (2023)
por: Anagnostidis, Sotiris, et al.
Publicado: (2023)
How Good is a Single Basin?
por: Lion, Kai, et al.
Publicado: (2024)
por: Lion, Kai, et al.
Publicado: (2024)
Causal Attention with Lookahead Keys
por: Song, Zhuoqing, et al.
Publicado: (2025)
por: Song, Zhuoqing, et al.
Publicado: (2025)
Learning Private Representations through Entropy-based Adversarial Training
por: Klein, Tassilo, et al.
Publicado: (2025)
por: Klein, Tassilo, et al.
Publicado: (2025)
Scaling Speculative Decoding with Lookahead Reasoning
por: Fu, Yichao, et al.
Publicado: (2025)
por: Fu, Yichao, et al.
Publicado: (2025)
The pitfalls of next-token prediction
por: Bachmann, Gregor, et al.
Publicado: (2024)
por: Bachmann, Gregor, et al.
Publicado: (2024)
Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
por: Fu, Yichao, et al.
Publicado: (2024)
por: Fu, Yichao, et al.
Publicado: (2024)
LORE: Lagrangian-Optimized Robust Embeddings for Visual Encoders
por: Khodabandeh, Borna, et al.
Publicado: (2025)
por: Khodabandeh, Borna, et al.
Publicado: (2025)
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
por: Huang, Zeyi, et al.
Publicado: (2026)
por: Huang, Zeyi, et al.
Publicado: (2026)
Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models
por: Alizadeh, Keivan, et al.
Publicado: (2024)
por: Alizadeh, Keivan, et al.
Publicado: (2024)
Enhancing Latent Computation in Transformers with Latent Tokens
por: Sun, Yuchang, et al.
Publicado: (2025)
por: Sun, Yuchang, et al.
Publicado: (2025)
Too Big to Think: Capacity, Memorization, and Generalization in Pre-Trained Transformers
por: Barron, Joshua, et al.
Publicado: (2025)
por: Barron, Joshua, et al.
Publicado: (2025)
The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models
por: Rizvi-Martel, Michael, et al.
Publicado: (2026)
por: Rizvi-Martel, Michael, et al.
Publicado: (2026)
Stop-Think-AutoRegress: Language Modeling with Latent Diffusion Planning
por: Lovelace, Justin, et al.
Publicado: (2026)
por: Lovelace, Justin, et al.
Publicado: (2026)
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
por: Kennedy, Ian W., et al.
Publicado: (2026)
por: Kennedy, Ian W., et al.
Publicado: (2026)
The Impact of Quantization on the Robustness of Transformer-based Text Classifiers
por: Neshaei, Seyed Parsa, et al.
Publicado: (2024)
por: Neshaei, Seyed Parsa, et al.
Publicado: (2024)
State Stream Transformer (SST) V2: Parallel Training of Nonlinear Recurrence for Latent Space Reasoning
por: Aviss, Thea
Publicado: (2026)
por: Aviss, Thea
Publicado: (2026)
Freely Long-Thinking Transformer (FraiLT)
por: Tabak, Akbay
Publicado: (2024)
por: Tabak, Akbay
Publicado: (2024)
A Dynamic Self-Evolving Extraction System
por: Amin-Naseri, Moin, et al.
Publicado: (2026)
por: Amin-Naseri, Moin, et al.
Publicado: (2026)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
por: Samragh, Mohammad, et al.
Publicado: (2024)
por: Samragh, Mohammad, et al.
Publicado: (2024)
Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
por: Yakushev, George, et al.
Publicado: (2025)
por: Yakushev, George, et al.
Publicado: (2025)
Imagine-then-Plan: Agent Learning from Adaptive Lookahead with World Models
por: Liu, Youwei, et al.
Publicado: (2026)
por: Liu, Youwei, et al.
Publicado: (2026)
Deconstructing Attention: Investigating Design Principles for Effective Language Modeling
por: Xue, Huiyin, et al.
Publicado: (2025)
por: Xue, Huiyin, et al.
Publicado: (2025)
ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces
por: Xu, Xin, et al.
Publicado: (2026)
por: Xu, Xin, et al.
Publicado: (2026)
Latent Adversarial Training Improves the Representation of Refusal
por: Abbas, Alexandra, et al.
Publicado: (2025)
por: Abbas, Alexandra, et al.
Publicado: (2025)
Reinforcement Learning for Latent-Space Thinking in LLMs
por: Özeren, Enes, et al.
Publicado: (2025)
por: Özeren, Enes, et al.
Publicado: (2025)
To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
Thinking in Latents: Adaptive Anchor Refinement for Implicit Reasoning in LLMs
por: Sheshanarayana, Disha, et al.
Publicado: (2026)
por: Sheshanarayana, Disha, et al.
Publicado: (2026)
Tracing the Roots: Leveraging Temporal Dynamics in Diffusion Trajectories for Origin Attribution
por: Floros, Andreas, et al.
Publicado: (2024)
por: Floros, Andreas, et al.
Publicado: (2024)
Position as Probability: Self-Supervised Transformers that Think Past Their Training for Length Extrapolation
por: Lee, Philip Heejun
Publicado: (2025)
por: Lee, Philip Heejun
Publicado: (2025)
LatentChem: From Textual CoT to Latent Thinking in Chemical Reasoning
por: Ye, Xinwu, et al.
Publicado: (2026)
por: Ye, Xinwu, et al.
Publicado: (2026)
Fast Byte Latent Transformer
por: Kallini, Julie, et al.
Publicado: (2026)
por: Kallini, Julie, et al.
Publicado: (2026)
Ejemplares similares
-
The Potential of CoT for Reasoning: A Closer Look at Trace Dynamics
por: Bachmann, Gregor, et al.
Publicado: (2026) -
Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models
por: Klein, Tassilo, et al.
Publicado: (2024) -
Rewriting the Budget: A General Framework for Black-Box Attacks Under Cost Asymmetry
por: Salmani, Mahdi, et al.
Publicado: (2025) -
Geometric Inductive Biases of Deep Networks: The Role of Data and Architecture
por: Movahedi, Sajad, et al.
Publicado: (2024) -
Is Random Attention Sufficient for Sequence Modeling? Disentangling Trainable Components in the Transformer
por: Dong, Yihe, et al.
Publicado: (2025)