Planning in a recurrent neural network that plays Sokoban
Fuente:
arXiv
Salvato in:
| Autori principali: | Taufeeque, Mohammad, Quirke, Philip, Li, Maximilian, Cundy, Chris, Tucker, Aaron David, Gleave, Adam, Garriga-Alonso, Adrià |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2025)
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2025)
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2026)
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2026)
Preference Learning with Lie Detectors can Induce Honesty or Evasion
di: Cundy, Chris, et al.
Pubblicazione: (2025)
di: Cundy, Chris, et al.
Pubblicazione: (2025)
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
di: Bowen, Dillon, et al.
Pubblicazione: (2025)
di: Bowen, Dillon, et al.
Pubblicazione: (2025)
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
di: Chanin, David, et al.
Pubblicazione: (2026)
di: Chanin, David, et al.
Pubblicazione: (2026)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2025)
di: Chanin, David, et al.
Pubblicazione: (2025)
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
di: Golechha, Satvik, et al.
Pubblicazione: (2025)
di: Golechha, Satvik, et al.
Pubblicazione: (2025)
Interpreting Emergent Planning in Model-Free Reinforcement Learning
di: Bush, Thomas, et al.
Pubblicazione: (2025)
di: Bush, Thomas, et al.
Pubblicazione: (2025)
Exploiting Novel GPT-4 APIs
di: Pelrine, Kellin, et al.
Pubblicazione: (2023)
di: Pelrine, Kellin, et al.
Pubblicazione: (2023)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
di: Chanin, David, et al.
Pubblicazione: (2025)
di: Chanin, David, et al.
Pubblicazione: (2025)
SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking
di: Cundy, Chris, et al.
Pubblicazione: (2023)
di: Cundy, Chris, et al.
Pubblicazione: (2023)
Biases in the Blind Spot: Detecting What LLMs Fail to Mention
di: Arcuschin, Iván, et al.
Pubblicazione: (2026)
di: Arcuschin, Iván, et al.
Pubblicazione: (2026)
Understanding Addition in Transformers
di: Quirke, Philip, et al.
Pubblicazione: (2023)
di: Quirke, Philip, et al.
Pubblicazione: (2023)
DiFR: Inference Verification Despite Nondeterminism
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
di: Karvonen, Adam, et al.
Pubblicazione: (2025)
On the dimension of pullback attractors in recurrent neural networks
di: Fadera, Muhammed
Pubblicazione: (2025)
di: Fadera, Muhammed
Pubblicazione: (2025)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
di: Kowal, Matthew, et al.
Pubblicazione: (2026)
di: Kowal, Matthew, et al.
Pubblicazione: (2026)
Hypothesis Testing the Circuit Hypothesis in LLMs
di: Shi, Claudia, et al.
Pubblicazione: (2024)
di: Shi, Claudia, et al.
Pubblicazione: (2024)
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models
di: Dombrowski, Ann-Kathrin, et al.
Pubblicazione: (2025)
di: Dombrowski, Ann-Kathrin, et al.
Pubblicazione: (2025)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
di: Struppek, Lukas, et al.
Pubblicazione: (2026)
di: Struppek, Lukas, et al.
Pubblicazione: (2026)
Can Go AIs be adversarially robust?
di: Tseng, Tom, et al.
Pubblicazione: (2024)
di: Tseng, Tom, et al.
Pubblicazione: (2024)
LayerCollapse: Adaptive compression of neural networks
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2023)
di: Shabgahi, Soheil Zibakhsh, et al.
Pubblicazione: (2023)
Deep neural networks have an inbuilt Occam's razor
di: Mingard, Chris, et al.
Pubblicazione: (2023)
di: Mingard, Chris, et al.
Pubblicazione: (2023)
Scaling Trends for Data Poisoning in LLMs
di: Bowen, Dillon, et al.
Pubblicazione: (2024)
di: Bowen, Dillon, et al.
Pubblicazione: (2024)
STARC: A General Framework For Quantifying Differences Between Reward Functions
di: Skalse, Joar, et al.
Pubblicazione: (2023)
di: Skalse, Joar, et al.
Pubblicazione: (2023)
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research
di: Harrasse, Abir, et al.
Pubblicazione: (2025)
di: Harrasse, Abir, et al.
Pubblicazione: (2025)
Advanced atom-level representations for protein flexibility prediction utilizing graph neural networks
di: Sarparast, Sina, et al.
Pubblicazione: (2024)
di: Sarparast, Sina, et al.
Pubblicazione: (2024)
Graph neural networks informed locally by thermodynamics
di: Tierz, Alicia, et al.
Pubblicazione: (2024)
di: Tierz, Alicia, et al.
Pubblicazione: (2024)
Geometry of naturalistic object representations in recurrent neural network models of working memory
di: Lei, Xiaoxuan, et al.
Pubblicazione: (2024)
di: Lei, Xiaoxuan, et al.
Pubblicazione: (2024)
Sharpe Ratio-Guided Active Learning for Preference Optimization in RLHF
di: Belakaria, Syrine, et al.
Pubblicazione: (2025)
di: Belakaria, Syrine, et al.
Pubblicazione: (2025)
Investigating the Indirect Object Identification circuit in Mamba
di: Ensign, Danielle, et al.
Pubblicazione: (2024)
di: Ensign, Danielle, et al.
Pubblicazione: (2024)
On permutation-invariant neural networks
di: Kimura, Masanari, et al.
Pubblicazione: (2024)
di: Kimura, Masanari, et al.
Pubblicazione: (2024)
Sobolev acceleration for neural networks
di: Oh, Jong Kwon, et al.
Pubblicazione: (2025)
di: Oh, Jong Kwon, et al.
Pubblicazione: (2025)
Attention mechanisms in neural networks
di: Hays, Hasi
Pubblicazione: (2026)
di: Hays, Hasi
Pubblicazione: (2026)
Sign language recognition from skeletal data using graph and recurrent neural networks
di: Mederos, B., et al.
Pubblicazione: (2025)
di: Mederos, B., et al.
Pubblicazione: (2025)
Principles of Lipschitz continuity in neural networks
di: Luo, Róisín
Pubblicazione: (2026)
di: Luo, Róisín
Pubblicazione: (2026)
Linearity-based neural network compression
di: Dobler, Silas, et al.
Pubblicazione: (2025)
di: Dobler, Silas, et al.
Pubblicazione: (2025)
Scaling Trends in Language Model Robustness
di: Howe, Nikolaus, et al.
Pubblicazione: (2024)
di: Howe, Nikolaus, et al.
Pubblicazione: (2024)
Applying graph neural network to SupplyGraph for supply chain network
di: Han, Kihwan
Pubblicazione: (2024)
di: Han, Kihwan
Pubblicazione: (2024)
Understanding the dynamics of the frequency bias in neural networks
di: Molina, Juan, et al.
Pubblicazione: (2024)
di: Molina, Juan, et al.
Pubblicazione: (2024)
Graph neural networks and non-commuting operators
di: Velasco, Mauricio, et al.
Pubblicazione: (2024)
di: Velasco, Mauricio, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2025) -
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
di: Taufeeque, Mohammad, et al.
Pubblicazione: (2026) -
Preference Learning with Lie Detectors can Induce Honesty or Evasion
di: Cundy, Chris, et al.
Pubblicazione: (2025) -
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
di: Bowen, Dillon, et al.
Pubblicazione: (2025) -
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
di: Chanin, David, et al.
Pubblicazione: (2026)