TOAST: Transformer Optimization using Adaptive and Simple Transformations
Fuente:
arXiv
Saved in:
| Main Authors: | Cannistraci, Irene, Antonelli, Simone, Palumbo, Emanuele, Sutter, Thomas M., Rodolà, Emanuele, Rieck, Bastian, Vogt, Julia E. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
You Only Train Once: Differentiable Subset Selection for Omics Data
by: Chopard, Daphné, et al.
Published: (2025)
by: Chopard, Daphné, et al.
Published: (2025)
Domain Elastic Transform: Bayesian Function Registration for High-Dimensional Scientific Data
by: Hirose, Osamu, et al.
Published: (2026)
by: Hirose, Osamu, et al.
Published: (2026)
Communicating Sound Through Natural Language
by: Rossi, Emanuele, et al.
Published: (2026)
by: Rossi, Emanuele, et al.
Published: (2026)
From Logits to Hierarchies: Hierarchical Clustering made Simple
by: Palumbo, Emanuele, et al.
Published: (2024)
by: Palumbo, Emanuele, et al.
Published: (2024)
Metric Based Few-Shot Graph Classification
by: Crisostomi, Donato, et al.
Published: (2022)
by: Crisostomi, Donato, et al.
Published: (2022)
Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models
by: Marincione, Davide, et al.
Published: (2025)
by: Marincione, Davide, et al.
Published: (2025)
From Bricks to Bridges: Product of Invariances to Enhance Latent Space Communication
by: Cannistraci, Irene, et al.
Published: (2023)
by: Cannistraci, Irene, et al.
Published: (2023)
Mergenetic: a Simple Evolutionary Model Merging Library
by: Minut, Adrian Robert, et al.
Published: (2025)
by: Minut, Adrian Robert, et al.
Published: (2025)
Accelerating Transformer Inference for Translation via Parallel Decoding
by: Santilli, Andrea, et al.
Published: (2023)
by: Santilli, Andrea, et al.
Published: (2023)
Foundation Model for Cardiac Time Series via Masked Latent Attention
by: Vandenhirtz, Moritz, et al.
Published: (2026)
by: Vandenhirtz, Moritz, et al.
Published: (2026)
Two Is Better Than One: Aligned Representation Pairs for Anomaly Detection
by: Ryser, Alain, et al.
Published: (2024)
by: Ryser, Alain, et al.
Published: (2024)
CliquePH: Higher-Order Information for Graph Neural Networks through Persistent Homology on Clique Graphs
by: Buffelli, Davide, et al.
Published: (2024)
by: Buffelli, Davide, et al.
Published: (2024)
Language Models are Injective and Hence Invertible
by: Nikolaou, Giorgos, et al.
Published: (2025)
by: Nikolaou, Giorgos, et al.
Published: (2025)
Two-Scale Latent Dynamics for Recurrent-Depth Transformers
by: Pappone, Francesco, et al.
Published: (2025)
by: Pappone, Francesco, et al.
Published: (2025)
MASS: MoErging through Adaptive Subspace Selection
by: Crisostomi, Donato, et al.
Published: (2025)
by: Crisostomi, Donato, et al.
Published: (2025)
MERGE$^3$: Efficient Evolutionary Merging on Consumer-grade GPUs
by: Mencattini, Tommaso, et al.
Published: (2025)
by: Mencattini, Tommaso, et al.
Published: (2025)
Topology meets Machine Learning: An Introduction using the Euler Characteristic Transform
by: Rieck, Bastian
Published: (2024)
by: Rieck, Bastian
Published: (2024)
Graph and Simplicial Complex Prediction Gaussian Process via the Hodgelet Representations
by: Alain, Mathieu, et al.
Published: (2025)
by: Alain, Mathieu, et al.
Published: (2025)
Multi-Way Representation Alignment
by: Achara, Akshit, et al.
Published: (2026)
by: Achara, Akshit, et al.
Published: (2026)
On Task Vectors and Gradients
by: Zhou, Luca, et al.
Published: (2025)
by: Zhou, Luca, et al.
Published: (2025)
Unity by Diversity: Improved Representation Learning in Multimodal VAEs
by: Sutter, Thomas M., et al.
Published: (2024)
by: Sutter, Thomas M., et al.
Published: (2024)
Differentiable Euler Characteristic Transforms for Shape Classification
by: Roell, Ernst, et al.
Published: (2023)
by: Roell, Ernst, et al.
Published: (2023)
Transform then Explore: a Simple and Effective Technique for Exploratory Combinatorial Optimization with Reinforcement Learning
by: Pu, Tianle, et al.
Published: (2024)
by: Pu, Tianle, et al.
Published: (2024)
Mapping representations in Reinforcement Learning via Semantic Alignment for Zero-Shot Stitching
by: Ricciardi, Antonio Pio, et al.
Published: (2025)
by: Ricciardi, Antonio Pio, et al.
Published: (2025)
Not All Latent Spaces Are Flat: Hyperbolic Concept Control
by: Briglia, Maria Rosaria, et al.
Published: (2026)
by: Briglia, Maria Rosaria, et al.
Published: (2026)
Simple Path Structural Encoding for Graph Transformers
by: Airale, Louis, et al.
Published: (2025)
by: Airale, Louis, et al.
Published: (2025)
Projection Methods for Operator Learning and Universal Approximation
by: Zappala, Emanuele
Published: (2024)
by: Zappala, Emanuele
Published: (2024)
ATM: Improving Model Merging by Alternating Tuning and Merging
by: Zhou, Luca, et al.
Published: (2024)
by: Zhou, Luca, et al.
Published: (2024)
Zero-Shot Quantization via Weight-Space Arithmetic
by: Solombrino, Daniele, et al.
Published: (2026)
by: Solombrino, Daniele, et al.
Published: (2026)
Preserving Privacy in Large Language Models: A Survey on Current Threats and Solutions
by: Miranda, Michele, et al.
Published: (2024)
by: Miranda, Michele, et al.
Published: (2024)
Implicit Inversion turns CLIP into a Decoder
by: D'Orazio, Antonio, et al.
Published: (2025)
by: D'Orazio, Antonio, et al.
Published: (2025)
Test-Time Training Undermines Safety Guardrails
by: Antonelli, Simone, et al.
Published: (2026)
by: Antonelli, Simone, et al.
Published: (2026)
From Pixels to Components: Eigenvector Masking for Visual Representation Learning
by: Bizeul, Alice, et al.
Published: (2025)
by: Bizeul, Alice, et al.
Published: (2025)
Point Cloud Synthesis Using Inner Product Transforms
by: Röell, Ernst, et al.
Published: (2024)
by: Röell, Ernst, et al.
Published: (2024)
An Embarrassingly Simple Approach to Enhance Transformer Performance in Genomic Selection for Crop Breeding
by: Chen, Renqi, et al.
Published: (2024)
by: Chen, Renqi, et al.
Published: (2024)
Solving Probabilistic Verification Problems of Neural Networks using Branch and Bound
by: Boetius, David, et al.
Published: (2024)
by: Boetius, David, et al.
Published: (2024)
Hybrid Modeling of Photoplethysmography for Non-invasive Monitoring of Cardiovascular Parameters
by: Palumbo, Emanuele, et al.
Published: (2025)
by: Palumbo, Emanuele, et al.
Published: (2025)
Architecture Determines Observability of Transformers
by: Carmichael, Thomas
Published: (2026)
by: Carmichael, Thomas
Published: (2026)
R3L: Relative Representations for Reinforcement Learning
by: Ricciardi, Antonio Pio, et al.
Published: (2024)
by: Ricciardi, Antonio Pio, et al.
Published: (2024)
An Introduction to Transformers
by: Turner, Richard E.
Published: (2023)
by: Turner, Richard E.
Published: (2023)
Similar Items
-
You Only Train Once: Differentiable Subset Selection for Omics Data
by: Chopard, Daphné, et al.
Published: (2025) -
Domain Elastic Transform: Bayesian Function Registration for High-Dimensional Scientific Data
by: Hirose, Osamu, et al.
Published: (2026) -
Communicating Sound Through Natural Language
by: Rossi, Emanuele, et al.
Published: (2026) -
From Logits to Hierarchies: Hierarchical Clustering made Simple
by: Palumbo, Emanuele, et al.
Published: (2024) -
Metric Based Few-Shot Graph Classification
by: Crisostomi, Donato, et al.
Published: (2022)