Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems
Fuente:
arXiv
Saved in:
| Main Authors: | Hoffmann, David T., Schrodi, Simon, Bratulić, Jelena, Behrmann, Nadine, Fischer, Volker, Brox, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
by: Schrodi, Simon, et al.
Published: (2024)
by: Schrodi, Simon, et al.
Published: (2024)
On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models
by: Bratulić, Jelena, et al.
Published: (2025)
by: Bratulić, Jelena, et al.
Published: (2025)
Label-Efficient LiDAR Semantic Segmentation with 2D-3D Vision Transformer Adapters
by: Hindel, Julia, et al.
Published: (2025)
by: Hindel, Julia, et al.
Published: (2025)
When and How Does CLIP Enable Domain and Compositional Generalization?
by: Kempf, Elias, et al.
Published: (2025)
by: Kempf, Elias, et al.
Published: (2025)
Concept Bottleneck Models Without Predefined Concepts
by: Schrodi, Simon, et al.
Published: (2024)
by: Schrodi, Simon, et al.
Published: (2024)
Using Knowledge Graphs to harvest datasets for efficient CLIP model training
by: Ging, Simon, et al.
Published: (2025)
by: Ging, Simon, et al.
Published: (2025)
What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models
by: Farid, Karim, et al.
Published: (2025)
by: Farid, Karim, et al.
Published: (2025)
Common Data Properties Limit Object-Attribute Binding in CLIP
by: Gurung, Bijay, et al.
Published: (2025)
by: Gurung, Bijay, et al.
Published: (2025)
Assessing Multimodal Chronic Wound Embeddings with Expert Triplet Agreement
by: Kabus, Fabian, et al.
Published: (2026)
by: Kabus, Fabian, et al.
Published: (2026)
Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
DITTO: Demonstration Imitation by Trajectory Transformation
by: Heppert, Nick, et al.
Published: (2024)
by: Heppert, Nick, et al.
Published: (2024)
Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy
by: Ging, Simon, et al.
Published: (2024)
by: Ging, Simon, et al.
Published: (2024)
Softmax-free Linear Transformers
by: Lu, Jiachen, et al.
Published: (2022)
by: Lu, Jiachen, et al.
Published: (2022)
Anomaly Detection with Conditioned Denoising Diffusion Models
by: Mousakhan, Arian, et al.
Published: (2023)
by: Mousakhan, Arian, et al.
Published: (2023)
Detect, Classify, Act: Categorizing Industrial Anomalies with Multi-Modal Large Language Models
by: Mokhtar, Sassan, et al.
Published: (2025)
by: Mokhtar, Sassan, et al.
Published: (2025)
SimA: Simple Softmax-free Attention for Vision Transformers
by: Koohpayegani, Soroush Abbasi, et al.
Published: (2022)
by: Koohpayegani, Soroush Abbasi, et al.
Published: (2022)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
by: Schrodi, Simon, et al.
Published: (2025)
by: Schrodi, Simon, et al.
Published: (2025)
StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization
by: Gaur, Gopalji, et al.
Published: (2025)
by: Gaur, Gopalji, et al.
Published: (2025)
Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
by: Bhattacharyya, Apratim, et al.
Published: (2025)
by: Bhattacharyya, Apratim, et al.
Published: (2025)
TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
K$α$LOS finds Consensus: A Meta-Algorithm for Evaluating Inter-Annotator Agreement in Complex Vision Tasks
by: Tschirschwitz, David, et al.
Published: (2026)
by: Tschirschwitz, David, et al.
Published: (2026)
The Unified Balance Theory of Second-Moment Exponential Scaling Optimizers in Visual Tasks
by: Zhang, Gongyue, et al.
Published: (2024)
by: Zhang, Gongyue, et al.
Published: (2024)
Compute-Efficient Medical Image Classification with Softmax-Free Transformers and Sequence Normalization
by: Khader, Firas, et al.
Published: (2024)
by: Khader, Firas, et al.
Published: (2024)
sshELF: Single-Shot Hierarchical Extrapolation of Latent Features for 3D Reconstruction from Sparse-Views
by: Najafli, Eyvaz, et al.
Published: (2025)
by: Najafli, Eyvaz, et al.
Published: (2025)
Transparency Distortion Robustness for SOTA Image Segmentation Tasks
by: Knauthe, Volker, et al.
Published: (2024)
by: Knauthe, Volker, et al.
Published: (2024)
Unlocking In-Context Learning for Natural Datasets Beyond Language Modelling
by: Bratulić, Jelena, et al.
Published: (2025)
by: Bratulić, Jelena, et al.
Published: (2025)
Agent Attention: On the Integration of Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2023)
by: Han, Dongchen, et al.
Published: (2023)
Realigned Softmax Warping for Deep Metric Learning
by: DeMoor, Michael G., et al.
Published: (2024)
by: DeMoor, Michael G., et al.
Published: (2024)
Bridging the Divide: Reconsidering Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2024)
by: Han, Dongchen, et al.
Published: (2024)
Trio-ViT: Post-Training Quantization and Acceleration for Softmax-Free Efficient Vision Transformer
by: Shi, Huihong, et al.
Published: (2024)
by: Shi, Huihong, et al.
Published: (2024)
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
by: Yuan, Huaying, et al.
Published: (2025)
by: Yuan, Huaying, et al.
Published: (2025)
MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View Stereo
by: Cao, Chenjie, et al.
Published: (2024)
by: Cao, Chenjie, et al.
Published: (2024)
Neural Point Cloud Diffusion for Disentangled 3D Shape and Appearance Generation
by: Schröppel, Philipp, et al.
Published: (2023)
by: Schröppel, Philipp, et al.
Published: (2023)
TADFormer : Task-Adaptive Dynamic Transformer for Efficient Multi-Task Learning
by: Baek, Seungmin, et al.
Published: (2025)
by: Baek, Seungmin, et al.
Published: (2025)
MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
by: Meng, Fanqing, et al.
Published: (2025)
by: Meng, Fanqing, et al.
Published: (2025)
Transforming Vision Transformer: Towards Efficient Multi-Task Asynchronous Learning
by: Zhong, Hanwen, et al.
Published: (2025)
by: Zhong, Hanwen, et al.
Published: (2025)
Diffusion for Out-of-Distribution Detection on Road Scenes and Beyond
by: Galesso, Silvio, et al.
Published: (2024)
by: Galesso, Silvio, et al.
Published: (2024)
ViSIR: Vision Transformer Single Image Reconstruction Method for Earth System Models
by: Zeraatkar, Ehsan, et al.
Published: (2025)
by: Zeraatkar, Ehsan, et al.
Published: (2025)
Quantifying Task Priority for Multi-Task Optimization
by: Jeong, Wooseong, et al.
Published: (2024)
by: Jeong, Wooseong, et al.
Published: (2024)
When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
by: Cao, Zhuo, et al.
Published: (2025)
by: Cao, Zhuo, et al.
Published: (2025)
Similar Items
-
Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
by: Schrodi, Simon, et al.
Published: (2024) -
On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models
by: Bratulić, Jelena, et al.
Published: (2025) -
Label-Efficient LiDAR Semantic Segmentation with 2D-3D Vision Transformer Adapters
by: Hindel, Julia, et al.
Published: (2025) -
When and How Does CLIP Enable Domain and Compositional Generalization?
by: Kempf, Elias, et al.
Published: (2025) -
Concept Bottleneck Models Without Predefined Concepts
by: Schrodi, Simon, et al.
Published: (2024)