Filter Equivariant Functions: A symmetric account of length-general extrapolation on lists
Fuente:
arXiv
Salvato in:
| Autori principali: | Lewis, Owen, Ghani, Neil, Dudzik, Andrew, Perivolaropoulos, Christos, Pascanu, Razvan, Veličković, Petar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Round and Round We Go! What makes Rotary Positional Encodings useful?
di: Barbero, Federico, et al.
Pubblicazione: (2024)
di: Barbero, Federico, et al.
Pubblicazione: (2024)
Perplexity Cannot Always Tell Right from Wrong
di: Veličković, Petar, et al.
Pubblicazione: (2026)
di: Veličković, Petar, et al.
Pubblicazione: (2026)
Softmax is not Enough (for Sharp Size Generalisation)
di: Veličković, Petar, et al.
Pubblicazione: (2024)
di: Veličković, Petar, et al.
Pubblicazione: (2024)
Asynchronous Algorithmic Alignment with Cocycles
di: Dudzik, Andrew, et al.
Pubblicazione: (2023)
di: Dudzik, Andrew, et al.
Pubblicazione: (2023)
Transformers meet Neural Algorithmic Reasoners
di: Bounsi, Wilfried, et al.
Pubblicazione: (2024)
di: Bounsi, Wilfried, et al.
Pubblicazione: (2024)
Latent Space Representations of Neural Algorithmic Reasoners
di: Mirjanić, Vladimir V., et al.
Pubblicazione: (2023)
di: Mirjanić, Vladimir V., et al.
Pubblicazione: (2023)
The Illusion of Stochasticity in LLMs
di: Gu, Xiangming, et al.
Pubblicazione: (2026)
di: Gu, Xiangming, et al.
Pubblicazione: (2026)
Why do LLMs attend to the first token?
di: Barbero, Federico, et al.
Pubblicazione: (2025)
di: Barbero, Federico, et al.
Pubblicazione: (2025)
Mining Generalizable Activation Functions
di: Vitvitskyi, Alex, et al.
Pubblicazione: (2026)
di: Vitvitskyi, Alex, et al.
Pubblicazione: (2026)
Transformers need glasses! Information over-squashing in language tasks
di: Barbero, Federico, et al.
Pubblicazione: (2024)
di: Barbero, Federico, et al.
Pubblicazione: (2024)
Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
di: Gu, Xiangming, et al.
Pubblicazione: (2026)
di: Gu, Xiangming, et al.
Pubblicazione: (2026)
Amplifying human performance in combinatorial competitive programming
di: Veličković, Petar, et al.
Pubblicazione: (2024)
di: Veličković, Petar, et al.
Pubblicazione: (2024)
How do language models learn facts? Dynamics, curricula and hallucinations
di: Zucchet, Nicolas, et al.
Pubblicazione: (2025)
di: Zucchet, Nicolas, et al.
Pubblicazione: (2025)
Leveraging Classical Algorithms for Graph Neural Networks
di: Wu, Jason, et al.
Pubblicazione: (2025)
di: Wu, Jason, et al.
Pubblicazione: (2025)
Optimizers Qualitatively Alter Solutions And We Should Leverage This
di: Pascanu, Razvan, et al.
Pubblicazione: (2025)
di: Pascanu, Razvan, et al.
Pubblicazione: (2025)
Automatic Functional Differentiation in JAX
di: Lin, Min
Pubblicazione: (2023)
di: Lin, Min
Pubblicazione: (2023)
Language Models for Code Completion: A Practical Evaluation
di: Izadi, Maliheh, et al.
Pubblicazione: (2024)
di: Izadi, Maliheh, et al.
Pubblicazione: (2024)
On the generalization of language models from in-context learning and finetuning: a controlled study
di: Lampinen, Andrew K., et al.
Pubblicazione: (2025)
di: Lampinen, Andrew K., et al.
Pubblicazione: (2025)
Recurrent Aggregators in Neural Algorithmic Reasoning
di: Xu, Kaijia, et al.
Pubblicazione: (2024)
di: Xu, Kaijia, et al.
Pubblicazione: (2024)
Position: Categorical Deep Learning is an Algebraic Theory of All Architectures
di: Gavranović, Bruno, et al.
Pubblicazione: (2024)
di: Gavranović, Bruno, et al.
Pubblicazione: (2024)
How do LLMs Compute Verbal Confidence
di: Kumaran, Dharshan, et al.
Pubblicazione: (2026)
di: Kumaran, Dharshan, et al.
Pubblicazione: (2026)
Recursive Function Definitions in Static Dataflow Graphs and their Implementation in TensorFlow
di: Kostopoulou, Kelly, et al.
Pubblicazione: (2024)
di: Kostopoulou, Kelly, et al.
Pubblicazione: (2024)
Deep Grokking: Would Deep Neural Networks Generalize Better?
di: Fan, Simin, et al.
Pubblicazione: (2024)
di: Fan, Simin, et al.
Pubblicazione: (2024)
NaN-Propagation: A Novel Method for Sparsity Detection in Black-Box Computational Functions
di: Sharpe, Peter
Pubblicazione: (2025)
di: Sharpe, Peter
Pubblicazione: (2025)
Revisiting Dynamic Evaluation: Online Adaptation for Large Language Models
di: Rannen-Triki, Amal, et al.
Pubblicazione: (2024)
di: Rannen-Triki, Amal, et al.
Pubblicazione: (2024)
Parallel Algorithms Align with Neural Execution
di: Engelmayer, Valerie, et al.
Pubblicazione: (2023)
di: Engelmayer, Valerie, et al.
Pubblicazione: (2023)
Lattice: Learning to Efficiently Compress the Memory
di: Karami, Mahdi, et al.
Pubblicazione: (2025)
di: Karami, Mahdi, et al.
Pubblicazione: (2025)
NoProp: Training Neural Networks without Full Back-propagation or Full Forward-propagation
di: Li, Qinyu, et al.
Pubblicazione: (2025)
di: Li, Qinyu, et al.
Pubblicazione: (2025)
Automatically Testing Functional Properties of Code Translation Models
di: Eniser, Hasan Ferit, et al.
Pubblicazione: (2023)
di: Eniser, Hasan Ferit, et al.
Pubblicazione: (2023)
An LLM-Tool Compiler for Fused Parallel Function Calling
di: Singh, Simranjit, et al.
Pubblicazione: (2024)
di: Singh, Simranjit, et al.
Pubblicazione: (2024)
Learning logic programs by discovering higher-order abstractions
di: Hocquette, Céline, et al.
Pubblicazione: (2023)
di: Hocquette, Céline, et al.
Pubblicazione: (2023)
What Can Grokking Teach Us About Learning Under Nonstationarity?
di: Lyle, Clare, et al.
Pubblicazione: (2025)
di: Lyle, Clare, et al.
Pubblicazione: (2025)
Meta-learning how to Share Credit among Macro-Actions
di: Hosu, Ionel-Alexandru, et al.
Pubblicazione: (2025)
di: Hosu, Ionel-Alexandru, et al.
Pubblicazione: (2025)
Revisiting Adam for Streaming Reinforcement Learning
di: Gogianu, Florin, et al.
Pubblicazione: (2026)
di: Gogianu, Florin, et al.
Pubblicazione: (2026)
A Deep Dive into Function Inlining and its Security Implications for ML-based Binary Analysis
di: Abusabha, Omar, et al.
Pubblicazione: (2025)
di: Abusabha, Omar, et al.
Pubblicazione: (2025)
How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models
di: Kumaran, Dharshan, et al.
Pubblicazione: (2025)
di: Kumaran, Dharshan, et al.
Pubblicazione: (2025)
Commute-Time-Optimised Graphs for GNNs
di: Sterner, Igor, et al.
Pubblicazione: (2024)
di: Sterner, Igor, et al.
Pubblicazione: (2024)
From Memorization to Reasoning in the Spectrum of Loss Curvature
di: Merullo, Jack, et al.
Pubblicazione: (2025)
di: Merullo, Jack, et al.
Pubblicazione: (2025)
Kalman Filter for Online Classification of Non-Stationary Data
di: Titsias, Michalis K., et al.
Pubblicazione: (2023)
di: Titsias, Michalis K., et al.
Pubblicazione: (2023)
A Deep Learning Model for Predicting Transformation Legality
di: Tiwari, Avani, et al.
Pubblicazione: (2025)
di: Tiwari, Avani, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Round and Round We Go! What makes Rotary Positional Encodings useful?
di: Barbero, Federico, et al.
Pubblicazione: (2024) -
Perplexity Cannot Always Tell Right from Wrong
di: Veličković, Petar, et al.
Pubblicazione: (2026) -
Softmax is not Enough (for Sharp Size Generalisation)
di: Veličković, Petar, et al.
Pubblicazione: (2024) -
Asynchronous Algorithmic Alignment with Cocycles
di: Dudzik, Andrew, et al.
Pubblicazione: (2023) -
Transformers meet Neural Algorithmic Reasoners
di: Bounsi, Wilfried, et al.
Pubblicazione: (2024)