Beyond the final layer: Attentive multilayer fusion for vision transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Ciernik, Laure, Morik, Marco, Thede, Lukas, Eyring, Luca, Nakajima, Shinichi, Akata, Zeynep, Muttenthaler, Lukas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Objective drives the consistency of representational similarity across datasets
by: Ciernik, Laure, et al.
Published: (2024)
by: Ciernik, Laure, et al.
Published: (2024)
Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
by: Huang, Yiran, et al.
Published: (2025)
by: Huang, Yiran, et al.
Published: (2025)
Reflecting on the State of Rehearsal-free Continual Learning with Pretrained Models
by: Thede, Lukas, et al.
Published: (2024)
by: Thede, Lukas, et al.
Published: (2024)
ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization
by: Eyring, Luca, et al.
Published: (2024)
by: Eyring, Luca, et al.
Published: (2024)
Fantastic Gains and Where to Find Them: On the Existence and Prospect of General Knowledge Transfer between Any Pretrained Model
by: Roth, Karsten, et al.
Published: (2023)
by: Roth, Karsten, et al.
Published: (2023)
Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
by: Eyring, Luca, et al.
Published: (2025)
by: Eyring, Luca, et al.
Published: (2025)
Disentangled Representation Learning with the Gromov-Monge Gap
by: Uscidda, Théo, et al.
Published: (2024)
by: Uscidda, Théo, et al.
Published: (2024)
Unbalancedness in Neural Monge Maps Improves Unpaired Domain Translation
by: Eyring, Luca, et al.
Published: (2023)
by: Eyring, Luca, et al.
Published: (2023)
Sparse Autoencoders are Topic Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
Vision-by-Language for Training-Free Compositional Image Retrieval
by: Karthik, Shyamgopal, et al.
Published: (2023)
by: Karthik, Shyamgopal, et al.
Published: (2023)
DeLoRA: Decoupling Angles and Strength in Low-rank Adaptation
by: Bini, Massimo, et al.
Published: (2025)
by: Bini, Massimo, et al.
Published: (2025)
CAST: Cross-Attentive Spatio-Temporal feature fusion for deepfake detection
by: Thakre, Aryan, et al.
Published: (2025)
by: Thakre, Aryan, et al.
Published: (2025)
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
by: Hummel, Thomas, et al.
Published: (2024)
by: Hummel, Thomas, et al.
Published: (2024)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
by: Kim, Jae Myung, et al.
Published: (2025)
by: Kim, Jae Myung, et al.
Published: (2025)
A Large Scale Analysis of Gender Biases in Text-to-Image Generative Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
by: Kim, Sanghwan, et al.
Published: (2025)
by: Kim, Sanghwan, et al.
Published: (2025)
Audio-Visual Generalized Zero-Shot Learning using Pre-Trained Large Multi-Modal Models
by: Kurzendörfer, David, et al.
Published: (2024)
by: Kurzendörfer, David, et al.
Published: (2024)
Rethinking Concept Bottleneck Models: From Pitfalls to Solutions
by: Tapli, Merve, et al.
Published: (2026)
by: Tapli, Merve, et al.
Published: (2026)
Explaining CLIP Zero-shot Predictions Through Concepts
by: Ozdemir, Onat, et al.
Published: (2026)
by: Ozdemir, Onat, et al.
Published: (2026)
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
by: Wu, Boyong, et al.
Published: (2026)
by: Wu, Boyong, et al.
Published: (2026)
Learning to segment anatomy and lesions from disparately labeled sources in brain MRI
by: Himmetoglu, Meva, et al.
Published: (2025)
by: Himmetoglu, Meva, et al.
Published: (2025)
ETHER: Efficient Finetuning of Large-Scale Models with Hyperplane Reflections
by: Bini, Massimo, et al.
Published: (2024)
by: Bini, Massimo, et al.
Published: (2024)
SAFER: Sharpness Aware layer-selective Finetuning for Enhanced Robustness in vision transformers
by: Gopal, Bhavna, et al.
Published: (2025)
by: Gopal, Bhavna, et al.
Published: (2025)
The Manifold Hypothesis for Gradient-Based Explanations
by: Bordt, Sebastian, et al.
Published: (2022)
by: Bordt, Sebastian, et al.
Published: (2022)
Scalable Ranked Preference Optimization for Text-to-Image Generation
by: Karthik, Shyamgopal, et al.
Published: (2024)
by: Karthik, Shyamgopal, et al.
Published: (2024)
Beyond conventional vision: RGB-event fusion for robust object detection in dynamic traffic scenarios
by: Liu, Zhanwen, et al.
Published: (2025)
by: Liu, Zhanwen, et al.
Published: (2025)
Dimensions underlying the representational alignment of deep neural networks with humans
by: Mahner, Florian P., et al.
Published: (2024)
by: Mahner, Florian P., et al.
Published: (2024)
Set Learning for Accurate and Calibrated Models
by: Muttenthaler, Lukas, et al.
Published: (2023)
by: Muttenthaler, Lukas, et al.
Published: (2023)
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
by: Xiao, Rui, et al.
Published: (2026)
by: Xiao, Rui, et al.
Published: (2026)
SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions
by: Bader, Jessica, et al.
Published: (2025)
by: Bader, Jessica, et al.
Published: (2025)
Road Obstacle Video Segmentation
by: Rai, Shyam Nandan, et al.
Published: (2025)
by: Rai, Shyam Nandan, et al.
Published: (2025)
FLAIR: VLM with Fine-grained Language-informed Image Representations
by: Xiao, Rui, et al.
Published: (2024)
by: Xiao, Rui, et al.
Published: (2024)
DataDream: Few-shot Guided Dataset Generation
by: Kim, Jae Myung, et al.
Published: (2024)
by: Kim, Jae Myung, et al.
Published: (2024)
Human alignment of neural network representations
by: Muttenthaler, Lukas, et al.
Published: (2022)
by: Muttenthaler, Lukas, et al.
Published: (2022)
Beyond flattening: a geometrically principled positional encoding for vision transformers with Weierstrass elliptic functions
by: Xin, Zhihang, et al.
Published: (2025)
by: Xin, Zhihang, et al.
Published: (2025)
HSFusion: A high-level vision task-driven infrared and visible image fusion network via semantic and geometric domain transformation
by: Jiang, Chengjie, et al.
Published: (2024)
by: Jiang, Chengjie, et al.
Published: (2024)
Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models
by: Singhi, Nishad, et al.
Published: (2024)
by: Singhi, Nishad, et al.
Published: (2024)
When Does Perceptual Alignment Benefit Vision Representations?
by: Sundaram, Shobhita, et al.
Published: (2024)
by: Sundaram, Shobhita, et al.
Published: (2024)
Context-Aware Multimodal Pretraining
by: Roth, Karsten, et al.
Published: (2024)
by: Roth, Karsten, et al.
Published: (2024)
Self-Supervised Training with Autoencoders for Visual Anomaly Detection
by: Bauer, Alexander, et al.
Published: (2022)
by: Bauer, Alexander, et al.
Published: (2022)
Similar Items
-
Objective drives the consistency of representational similarity across datasets
by: Ciernik, Laure, et al.
Published: (2024) -
Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
by: Huang, Yiran, et al.
Published: (2025) -
Reflecting on the State of Rehearsal-free Continual Learning with Pretrained Models
by: Thede, Lukas, et al.
Published: (2024) -
ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization
by: Eyring, Luca, et al.
Published: (2024) -
Fantastic Gains and Where to Find Them: On the Existence and Prospect of General Knowledge Transfer between Any Pretrained Model
by: Roth, Karsten, et al.
Published: (2023)