Softmax is not Enough (for Sharp Size Generalisation)
Fuente:
arXiv
Saved in:
| Main Authors: | Veličković, Petar, Perivolaropoulos, Christos, Barbero, Federico, Pascanu, Razvan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Perplexity Cannot Always Tell Right from Wrong
by: Veličković, Petar, et al.
Published: (2026)
by: Veličković, Petar, et al.
Published: (2026)
Round and Round We Go! What makes Rotary Positional Encodings useful?
by: Barbero, Federico, et al.
Published: (2024)
by: Barbero, Federico, et al.
Published: (2024)
Asynchronous Algorithmic Alignment with Cocycles
by: Dudzik, Andrew, et al.
Published: (2023)
by: Dudzik, Andrew, et al.
Published: (2023)
Mining Generalizable Activation Functions
by: Vitvitskyi, Alex, et al.
Published: (2026)
by: Vitvitskyi, Alex, et al.
Published: (2026)
Filter Equivariant Functions: A symmetric account of length-general extrapolation on lists
by: Lewis, Owen, et al.
Published: (2025)
by: Lewis, Owen, et al.
Published: (2025)
From Markov to Laplace: How Mamba In-Context Learns Markov Chains
by: Bondaschi, Marco, et al.
Published: (2025)
by: Bondaschi, Marco, et al.
Published: (2025)
Latent Space Representations of Neural Algorithmic Reasoners
by: Mirjanić, Vladimir V., et al.
Published: (2023)
by: Mirjanić, Vladimir V., et al.
Published: (2023)
Learning in Convolutional Neural Networks Accelerated by Transfer Entropy
by: Moldovan, Adrian, et al.
Published: (2024)
by: Moldovan, Adrian, et al.
Published: (2024)
Temporal Graph Rewiring with Expander Graphs
by: Petrović, Katarina, et al.
Published: (2024)
by: Petrović, Katarina, et al.
Published: (2024)
Bottlenecked Transformers: Periodic KV Cache Consolidation for Generalised Reasoning
by: Oomerjee, Adnan, et al.
Published: (2025)
by: Oomerjee, Adnan, et al.
Published: (2025)
Recurrent Aggregators in Neural Algorithmic Reasoning
by: Xu, Kaijia, et al.
Published: (2024)
by: Xu, Kaijia, et al.
Published: (2024)
Navigating Potholes with Geometry-Aware Sharpness Minimization
by: Dufort-Labbé, Simon, et al.
Published: (2026)
by: Dufort-Labbé, Simon, et al.
Published: (2026)
What makes a good feedforward computational graph?
by: Vitvitskyi, Alex, et al.
Published: (2025)
by: Vitvitskyi, Alex, et al.
Published: (2025)
Softmax is not Enough (for Adaptive Conformal Classification)
by: Attar, Navid Akhavan, et al.
Published: (2026)
by: Attar, Navid Akhavan, et al.
Published: (2026)
How do LLMs Compute Verbal Confidence
by: Kumaran, Dharshan, et al.
Published: (2026)
by: Kumaran, Dharshan, et al.
Published: (2026)
Lattice: Learning to Efficiently Compress the Memory
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Why do LLMs attend to the first token?
by: Barbero, Federico, et al.
Published: (2025)
by: Barbero, Federico, et al.
Published: (2025)
Information Plane Analysis Visualization in Deep Learning via Transfer Entropy
by: Moldovan, Adrian, et al.
Published: (2024)
by: Moldovan, Adrian, et al.
Published: (2024)
Meta-learning how to Share Credit among Macro-Actions
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
by: Hosu, Ionel-Alexandru, et al.
Published: (2025)
Revisiting Adam for Streaming Reinforcement Learning
by: Gogianu, Florin, et al.
Published: (2026)
by: Gogianu, Florin, et al.
Published: (2026)
How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models
by: Kumaran, Dharshan, et al.
Published: (2025)
by: Kumaran, Dharshan, et al.
Published: (2025)
Cayley Graph Propagation
by: Wilson, JJ, et al.
Published: (2024)
by: Wilson, JJ, et al.
Published: (2024)
PSL: Rethinking and Improving Softmax Loss from Pairwise Perspective for Recommendation
by: Yang, Weiqin, et al.
Published: (2024)
by: Yang, Weiqin, et al.
Published: (2024)
KNARsack: Teaching Neural Algorithmic Reasoners to Solve Pseudo-Polynomial Problems
by: Požgaj, Stjepan, et al.
Published: (2025)
by: Požgaj, Stjepan, et al.
Published: (2025)
Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory
by: Anwar, Usman, et al.
Published: (2026)
by: Anwar, Usman, et al.
Published: (2026)
Rotary Position Encodings for Graphs
by: Reid, Isaac, et al.
Published: (2025)
by: Reid, Isaac, et al.
Published: (2025)
Position: Categorical Deep Learning is an Algebraic Theory of All Architectures
by: Gavranović, Bruno, et al.
Published: (2024)
by: Gavranović, Bruno, et al.
Published: (2024)
State Soup: In-Context Skill Learning, Retrieval and Mixing
by: Pióro, Maciej, et al.
Published: (2024)
by: Pióro, Maciej, et al.
Published: (2024)
LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
by: Schmied, Thomas, et al.
Published: (2025)
by: Schmied, Thomas, et al.
Published: (2025)
Transformers need glasses! Information over-squashing in language tasks
by: Barbero, Federico, et al.
Published: (2024)
by: Barbero, Federico, et al.
Published: (2024)
The Gatekeeper Knows Enough
by: Abebayew, Fikresilase Wondmeneh
Published: (2025)
by: Abebayew, Fikresilase Wondmeneh
Published: (2025)
Intent-Aware DRL-Based NOMA Uplink Dynamic Scheduler for IIoT
by: Mostafa, Salwa, et al.
Published: (2024)
by: Mostafa, Salwa, et al.
Published: (2024)
Trustworthy Actionable Perturbations
by: Friedbaum, Jesse, et al.
Published: (2024)
by: Friedbaum, Jesse, et al.
Published: (2024)
Random Aggregate Beamforming for Over-the-Air Federated Learning in Large-Scale Networks
by: Xu, Chunmei, et al.
Published: (2024)
by: Xu, Chunmei, et al.
Published: (2024)
Partial Information Decomposition for Data Interpretability and Feature Selection
by: Westphal, Charles, et al.
Published: (2024)
by: Westphal, Charles, et al.
Published: (2024)
MambaJSCC: Adaptive Deep Joint Source-Channel Coding with Generalized State Space Model
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
Adaptive $k$-nearest neighbor classifier based on the local estimation of the shape operator
by: Levada, Alexandre Luís Magalhães, et al.
Published: (2024)
by: Levada, Alexandre Luís Magalhães, et al.
Published: (2024)
Combinatorial Multi-armed Bandits: Arm Selection via Group Testing
by: Mukherjee, Arpan, et al.
Published: (2024)
by: Mukherjee, Arpan, et al.
Published: (2024)
Hierarchical Over-the-Air Federated Learning with Awareness of Interference and Data Heterogeneity
by: Azimi-Abarghouyi, Seyed Mohammad, et al.
Published: (2024)
by: Azimi-Abarghouyi, Seyed Mohammad, et al.
Published: (2024)
Can Kernel Methods Explain How the Data Affects Neural Collapse?
by: Kothapalli, Vignesh, et al.
Published: (2024)
by: Kothapalli, Vignesh, et al.
Published: (2024)
Similar Items
-
Perplexity Cannot Always Tell Right from Wrong
by: Veličković, Petar, et al.
Published: (2026) -
Round and Round We Go! What makes Rotary Positional Encodings useful?
by: Barbero, Federico, et al.
Published: (2024) -
Asynchronous Algorithmic Alignment with Cocycles
by: Dudzik, Andrew, et al.
Published: (2023) -
Mining Generalizable Activation Functions
by: Vitvitskyi, Alex, et al.
Published: (2026) -
Filter Equivariant Functions: A symmetric account of length-general extrapolation on lists
by: Lewis, Owen, et al.
Published: (2025)