N-Gram Induction Heads for In-Context RL: Improving Stability and Reducing Data Needs
Fuente:
arXiv
Guardado en:
| Autores principales: | Zisman, Ilya, Nikulin, Alexander, Sinii, Viacheslav, Tarasov, Denis, Lyubaykin, Nikita, Polubarov, Andrei, Kiselev, Igor, Kurenkov, Vladislav |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Latent Action Learning Requires Supervision in the Presence of Distractors
por: Nikulin, Alexander, et al.
Publicado: (2025)
por: Nikulin, Alexander, et al.
Publicado: (2025)
Yes, Q-learning Helps Offline In-Context RL
por: Tarasov, Denis, et al.
Publicado: (2025)
por: Tarasov, Denis, et al.
Publicado: (2025)
Vintix: Action Model via In-Context Reinforcement Learning
por: Polubarov, Andrey, et al.
Publicado: (2025)
por: Polubarov, Andrey, et al.
Publicado: (2025)
Object-Centric Latent Action Learning
por: Klepach, Albina, et al.
Publicado: (2025)
por: Klepach, Albina, et al.
Publicado: (2025)
Vision-Language Models Unlock Task-Centric Latent Actions
por: Nikulin, Alexander, et al.
Publicado: (2026)
por: Nikulin, Alexander, et al.
Publicado: (2026)
NinA: Normalizing Flows in Action. Training VLA Models with Normalizing Flows
por: Tarasov, Denis, et al.
Publicado: (2025)
por: Tarasov, Denis, et al.
Publicado: (2025)
In-Context Reinforcement Learning for Variable Action Spaces
por: Sinii, Viacheslav, et al.
Publicado: (2023)
por: Sinii, Viacheslav, et al.
Publicado: (2023)
Emergence of In-Context Reinforcement Learning from Noise Distillation
por: Zisman, Ilya, et al.
Publicado: (2023)
por: Zisman, Ilya, et al.
Publicado: (2023)
XLand-MiniGrid: Scalable Meta-Reinforcement Learning Environments in JAX
por: Nikulin, Alexander, et al.
Publicado: (2023)
por: Nikulin, Alexander, et al.
Publicado: (2023)
XLand-100B: A Large-Scale Multi-Task Dataset for In-Context Reinforcement Learning
por: Nikulin, Alexander, et al.
Publicado: (2024)
por: Nikulin, Alexander, et al.
Publicado: (2024)
Vintix II: Decision Pre-Trained Transformer is a Scalable In-Context Reinforcement Learner
por: Polubarov, Andrei, et al.
Publicado: (2026)
por: Polubarov, Andrei, et al.
Publicado: (2026)
Zero-Shot Adaptation of Behavioral Foundation Models to Unseen Dynamics
por: Bobrin, Maksim, et al.
Publicado: (2025)
por: Bobrin, Maksim, et al.
Publicado: (2025)
cadrille: Multi-modal CAD Reconstruction with Reinforcement Learning
por: Kolodiazhnyi, Maksim, et al.
Publicado: (2025)
por: Kolodiazhnyi, Maksim, et al.
Publicado: (2025)
You Do Not Fully Utilize Transformer's Representation Capacity
por: Gerasimov, Gleb, et al.
Publicado: (2025)
por: Gerasimov, Gleb, et al.
Publicado: (2025)
Electrostatics from Laplacian Eigenbasis for Neural Network Interatomic Potentials
por: Zhdanov, Maksim, et al.
Publicado: (2025)
por: Zhdanov, Maksim, et al.
Publicado: (2025)
Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
por: Sinii, Viacheslav, et al.
Publicado: (2025)
por: Sinii, Viacheslav, et al.
Publicado: (2025)
Steering LLM Reasoning Through Bias-Only Adaptation
por: Sinii, Viacheslav, et al.
Publicado: (2025)
por: Sinii, Viacheslav, et al.
Publicado: (2025)
Curvature-Aligned Probing for Local Loss-Landscape Stabilization
por: Kiselev, Nikita, et al.
Publicado: (2026)
por: Kiselev, Nikita, et al.
Publicado: (2026)
Closing the Curvature Gap: Full Transformer Hessians and Their Implications for Scaling Laws
por: Petrov, Egor, et al.
Publicado: (2025)
por: Petrov, Egor, et al.
Publicado: (2025)
On the Emergence of Induction Heads for In-Context Learning
por: Musat, Tiberiu, et al.
Publicado: (2025)
por: Musat, Tiberiu, et al.
Publicado: (2025)
The Role of Deep Learning Regularizations on Actors in Offline RL
por: Tarasov, Denis, et al.
Publicado: (2024)
por: Tarasov, Denis, et al.
Publicado: (2024)
The Differences Between Direct Alignment Algorithms are a Blur
por: Gorbatovski, Alexey, et al.
Publicado: (2025)
por: Gorbatovski, Alexey, et al.
Publicado: (2025)
Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success
por: Bredis, George, et al.
Publicado: (2025)
por: Bredis, George, et al.
Publicado: (2025)
Identifying Semantic Induction Heads to Understand In-Context Learning
por: Ren, Jie, et al.
Publicado: (2024)
por: Ren, Jie, et al.
Publicado: (2024)
Temporal Dependencies in In-Context Learning: The Role of Induction Heads
por: Bajaj, Anooshka, et al.
Publicado: (2026)
por: Bajaj, Anooshka, et al.
Publicado: (2026)
Unraveling the Hessian: A Key to Smooth Convergence in Loss Function Landscapes
por: Kiselev, Nikita, et al.
Publicado: (2024)
por: Kiselev, Nikita, et al.
Publicado: (2024)
ABRA: Agent Benchmark for Radiology Applications
por: Maksudov, Bulat, et al.
Publicado: (2026)
por: Maksudov, Bulat, et al.
Publicado: (2026)
NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
por: Bataev, Vladimir, et al.
Publicado: (2025)
por: Bataev, Vladimir, et al.
Publicado: (2025)
The Evolution of Statistical Induction Heads: In-Context Learning Markov Chains
por: Edelman, Benjamin L., et al.
Publicado: (2024)
por: Edelman, Benjamin L., et al.
Publicado: (2024)
Understanding and Controlling Repetition Neurons and Induction Heads in In-Context Learning
por: Doan, Nhi Hoai, et al.
Publicado: (2025)
por: Doan, Nhi Hoai, et al.
Publicado: (2025)
Selective Induction Heads: How Transformers Select Causal Structures In Context
por: D'Angelo, Francesco, et al.
Publicado: (2025)
por: D'Angelo, Francesco, et al.
Publicado: (2025)
ESSA: Evolutionary Strategies for Scalable Alignment
por: Korotyshova, Daria, et al.
Publicado: (2025)
por: Korotyshova, Daria, et al.
Publicado: (2025)
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
por: Plyusov, Daniil, et al.
Publicado: (2026)
por: Plyusov, Daniil, et al.
Publicado: (2026)
Improved Probabilistic Lower Bounds for Separable Matrices
por: Goshkoder, Daniil, et al.
Publicado: (2024)
por: Goshkoder, Daniil, et al.
Publicado: (2024)
The Impact of Globalization on the Development of Hiphop: How European and American Hip-Hop Mixed
por: Nikulin, Roman
Publicado: (2025)
por: Nikulin, Roman
Publicado: (2025)
Fonología segmental del chiquitano migueleño
por: Andrey Nikulin
Publicado: (2021)
por: Andrey Nikulin
Publicado: (2021)
Una metodología Sistémica y creativa para la gestión estratégica: Caso de Estudio Región de Atacama-Chile
por: Christopher Nikulin
Publicado: (2015)
por: Christopher Nikulin
Publicado: (2015)
A revised reconstruction of the Proto-Tupian vowel system
por: Andrey Nikulin
Publicado: (2022)
por: Andrey Nikulin
Publicado: (2022)
UNDECLARED WORK IN POLAND – CHARACTERISTICS AND PREVALENCE
por: Dagmara Nikulin
Publicado: (2016)
por: Dagmara Nikulin
Publicado: (2016)
A computational model of age‐dependent cardiomyocyte apoptosis
por: Elena Kutumova, et al.
Publicado: (2025)
por: Elena Kutumova, et al.
Publicado: (2025)
Ejemplares similares
-
Latent Action Learning Requires Supervision in the Presence of Distractors
por: Nikulin, Alexander, et al.
Publicado: (2025) -
Yes, Q-learning Helps Offline In-Context RL
por: Tarasov, Denis, et al.
Publicado: (2025) -
Vintix: Action Model via In-Context Reinforcement Learning
por: Polubarov, Andrey, et al.
Publicado: (2025) -
Object-Centric Latent Action Learning
por: Klepach, Albina, et al.
Publicado: (2025) -
Vision-Language Models Unlock Task-Centric Latent Actions
por: Nikulin, Alexander, et al.
Publicado: (2026)