Beyond Position: the emergence of wavelet-like properties in Transformers
Fuente:
arXiv
Guardado en:
| Autores principales: | Ruscio, Valeria, Nanni, Umberto, Silvestri, Fabrizio |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
What are you sinking? A geometric approach on attention sink
por: Ruscio, Valeria, et al.
Publicado: (2025)
por: Ruscio, Valeria, et al.
Publicado: (2025)
The Phenomenology of Hallucinations
por: Ruscio, Valeria, et al.
Publicado: (2026)
por: Ruscio, Valeria, et al.
Publicado: (2026)
Where Pretraining writes and Alignment reads: the asymmetry of Transformer weight space
por: Ruscio, Valeria, et al.
Publicado: (2026)
por: Ruscio, Valeria, et al.
Publicado: (2026)
Titans Revisited: A Lightweight Reimplementation and Critical Analysis of a Test-Time Memory Model
por: Di Nepi, Gavriel, et al.
Publicado: (2025)
por: Di Nepi, Gavriel, et al.
Publicado: (2025)
Learning with Noisy Labels through Learnable Weighting and Centroid Similarity
por: Wani, Farooq Ahmad, et al.
Publicado: (2023)
por: Wani, Farooq Ahmad, et al.
Publicado: (2023)
Polynomial Neural Sheaf Diffusion: A Spectral Filtering Approach on Cellular Sheaves
por: Borgi, Alessio, et al.
Publicado: (2025)
por: Borgi, Alessio, et al.
Publicado: (2025)
Evading Community Detection via Counterfactual Neighborhood Search
por: Bernini, Andrea, et al.
Publicado: (2023)
por: Bernini, Andrea, et al.
Publicado: (2023)
Learning from Partial Chain-of-Thought via Truncated-Reasoning Self-Distillation
por: Silvestri, Gianluigi, et al.
Publicado: (2026)
por: Silvestri, Gianluigi, et al.
Publicado: (2026)
The Majority Vote Paradigm Shift: When Popular Meets Optimal
por: Purificato, Antonio, et al.
Publicado: (2025)
por: Purificato, Antonio, et al.
Publicado: (2025)
Understanding and Improving Laplacian Positional Encodings For Temporal GNNs
por: Galron, Yaniv, et al.
Publicado: (2025)
por: Galron, Yaniv, et al.
Publicado: (2025)
Graph Transformers without Positional Encodings
por: Garg, Ayush
Publicado: (2024)
por: Garg, Ayush
Publicado: (2024)
Position Paper: From Edge AI to Adaptive Edge AI
por: Pittorino, Fabrizio, et al.
Publicado: (2026)
por: Pittorino, Fabrizio, et al.
Publicado: (2026)
Benchmarking Positional Encodings for GNNs and Graph Transformers
por: Grötschla, Florian, et al.
Publicado: (2024)
por: Grötschla, Florian, et al.
Publicado: (2024)
On Task Vectors and Gradients
por: Zhou, Luca, et al.
Publicado: (2025)
por: Zhou, Luca, et al.
Publicado: (2025)
ARDNS-FN-Quantum: A Quantum-Enhanced Reinforcement Learning Framework with Cognitive-Inspired Adaptive Exploration for Dynamic Environments
por: de Sousa, Umberto Gonçalves
Publicado: (2025)
por: de Sousa, Umberto Gonçalves
Publicado: (2025)
Diffusion Transformers for Tabular Data Time Series Generation
por: Garuti, Fabrizio, et al.
Publicado: (2025)
por: Garuti, Fabrizio, et al.
Publicado: (2025)
Position: Beyond Euclidean -- Foundation Models Should Embrace Non-Euclidean Geometries
por: He, Neil, et al.
Publicado: (2025)
por: He, Neil, et al.
Publicado: (2025)
Beyond Myopia: Learning from Positive and Unlabeled Data through Holistic Predictive Trends
por: Wang, Xinrui, et al.
Publicado: (2023)
por: Wang, Xinrui, et al.
Publicado: (2023)
PruneGCRN: Minimizing and explaining spatio-temporal problems through node pruning
por: García-Sigüenza, Javier, et al.
Publicado: (2025)
por: García-Sigüenza, Javier, et al.
Publicado: (2025)
Beyond the Known: Decision Making with Counterfactual Reasoning Decision Transformer
por: Nguyen, Minh Hoang, et al.
Publicado: (2025)
por: Nguyen, Minh Hoang, et al.
Publicado: (2025)
One Transformer for All Time Series: Representing and Training with Time-Dependent Heterogeneous Tabular Data
por: Luetto, Simone, et al.
Publicado: (2023)
por: Luetto, Simone, et al.
Publicado: (2023)
Debiasing Machine Unlearning with Counterfactual Examples
por: Chen, Ziheng, et al.
Publicado: (2024)
por: Chen, Ziheng, et al.
Publicado: (2024)
New Statistical Framework for Extreme Error Probability in High-Stakes Domains for Reliable Machine Learning
por: Michelucci, Umberto, et al.
Publicado: (2025)
por: Michelucci, Umberto, et al.
Publicado: (2025)
MILP-SAT-GNN: Yet Another Neural SAT Solver
por: Cardillo, Franco Alberto, et al.
Publicado: (2025)
por: Cardillo, Franco Alberto, et al.
Publicado: (2025)
PFformer: A Position-Free Transformer Variant for Extreme-Adaptive Multivariate Time Series Forecasting
por: Li, Yanhong, et al.
Publicado: (2025)
por: Li, Yanhong, et al.
Publicado: (2025)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
por: Fan, Dongyang, et al.
Publicado: (2025)
por: Fan, Dongyang, et al.
Publicado: (2025)
Sensitivity-Positional Co-Localization in GQA Transformers
por: Rao, Manoj Chandrashekar
Publicado: (2026)
por: Rao, Manoj Chandrashekar
Publicado: (2026)
Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer
por: Wang, Yongyi, et al.
Publicado: (2026)
por: Wang, Yongyi, et al.
Publicado: (2026)
Beyond Labels: Aligning Large Language Models with Human-like Reasoning
por: Kabir, Muhammad Rafsan, et al.
Publicado: (2024)
por: Kabir, Muhammad Rafsan, et al.
Publicado: (2024)
Preference-Based Alignment of Discrete Diffusion Models
por: Borso, Umberto, et al.
Publicado: (2025)
por: Borso, Umberto, et al.
Publicado: (2025)
HOP to the Next Tasks and Domains for Continual Learning in NLP
por: Michieli, Umberto, et al.
Publicado: (2024)
por: Michieli, Umberto, et al.
Publicado: (2024)
The Infinite-Dimensional Nature of Spectroscopy and Why Models Succeed, Fail, and Mislead
por: Michelucci, Umberto, et al.
Publicado: (2026)
por: Michelucci, Umberto, et al.
Publicado: (2026)
LogGuardQ: A Cognitive-Enhanced Reinforcement Learning Framework for Cybersecurity Anomaly Detection in Security Logs
por: de Sousa, Umberto Gonçalves
Publicado: (2025)
por: de Sousa, Umberto Gonçalves
Publicado: (2025)
ATM: Improving Model Merging by Alternating Tuning and Merging
por: Zhou, Luca, et al.
Publicado: (2024)
por: Zhou, Luca, et al.
Publicado: (2024)
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond
por: Ke, Yekun, et al.
Publicado: (2024)
por: Ke, Yekun, et al.
Publicado: (2024)
Lost in the Middle at Birth: An Exact Theory of Transformer Position Bias
por: Chowdhury, Borun D
Publicado: (2026)
por: Chowdhury, Borun D
Publicado: (2026)
Context-Selective State Space Models: Feedback is All You Need
por: Zattra, Riccardo, et al.
Publicado: (2025)
por: Zattra, Riccardo, et al.
Publicado: (2025)
Position: The Turing-Completeness of Autoregressive Transformers Relies Heavily on Context Management
por: Cui, Guanyu, et al.
Publicado: (2026)
por: Cui, Guanyu, et al.
Publicado: (2026)
Algebraic Positional Encodings
por: Kogkalidis, Konstantinos, et al.
Publicado: (2023)
por: Kogkalidis, Konstantinos, et al.
Publicado: (2023)
$\nabla τ$: Gradient-based and Task-Agnostic machine Unlearning
por: Trippa, Daniel, et al.
Publicado: (2024)
por: Trippa, Daniel, et al.
Publicado: (2024)
Ejemplares similares
-
What are you sinking? A geometric approach on attention sink
por: Ruscio, Valeria, et al.
Publicado: (2025) -
The Phenomenology of Hallucinations
por: Ruscio, Valeria, et al.
Publicado: (2026) -
Where Pretraining writes and Alignment reads: the asymmetry of Transformer weight space
por: Ruscio, Valeria, et al.
Publicado: (2026) -
Titans Revisited: A Lightweight Reimplementation and Critical Analysis of a Test-Time Memory Model
por: Di Nepi, Gavriel, et al.
Publicado: (2025) -
Learning with Noisy Labels through Learnable Weighting and Centroid Similarity
por: Wani, Farooq Ahmad, et al.
Publicado: (2023)