Guardado en:
| Autores principales: | Su, Qin, Luo, Tie |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2603.01605 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
What You See is What You Classify: Black Box Attributions
por: Stalder, Steven, et al.
Publicado: (2022)
por: Stalder, Steven, et al.
Publicado: (2022)
LINE: LLM-based Iterative Neuron Explanations for Vision Models
por: Zaigrajew, Vladimir, et al.
Publicado: (2026)
por: Zaigrajew, Vladimir, et al.
Publicado: (2026)
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
por: Bendikas, Rokas, et al.
Publicado: (2025)
por: Bendikas, Rokas, et al.
Publicado: (2025)
Probabilistic Conceptual Explainers: Trustworthy Conceptual Explanations for Vision Foundation Models
por: Wang, Hengyi, et al.
Publicado: (2024)
por: Wang, Hengyi, et al.
Publicado: (2024)
When Shared Knowledge Hurts: Spectral Over-Accumulation in Model Merging
por: Li, Yayuan, et al.
Publicado: (2026)
por: Li, Yayuan, et al.
Publicado: (2026)
HER2 Expression Prediction with Flexible Multi-Modal Inputs via Dynamic Bidirectional Reconstruction
por: Qin, Jie, et al.
Publicado: (2025)
por: Qin, Jie, et al.
Publicado: (2025)
What Matters in Practical Learned Image Compression
por: Tatwawadi, Kedar, et al.
Publicado: (2026)
por: Tatwawadi, Kedar, et al.
Publicado: (2026)
What augmentations are sensitive to hyper-parameters and why?
por: Awais, Ch Muhammad, et al.
Publicado: (2021)
por: Awais, Ch Muhammad, et al.
Publicado: (2021)
Don't Collapse Your Features: Why CenterLoss Hurts OOD Detection and Multi-Scale Mahalanobis Wins
por: Ray, Rahul D
Publicado: (2026)
por: Ray, Rahul D
Publicado: (2026)
When Generative Augmentation Hurts: A Benchmark Study of GAN and Diffusion Models for Bias Correction in AI Classification Systems
por: Gupta, Shesh Narayan, et al.
Publicado: (2026)
por: Gupta, Shesh Narayan, et al.
Publicado: (2026)
Tell What You Hear From What You See -- Video to Audio Generation Through Text
por: Liu, Xiulong, et al.
Publicado: (2024)
por: Liu, Xiulong, et al.
Publicado: (2024)
What Matters in Range View 3D Object Detection
por: Wilson, Benjamin, et al.
Publicado: (2024)
por: Wilson, Benjamin, et al.
Publicado: (2024)
What Variables Affect Out-of-Distribution Generalization in Pretrained Models?
por: Harun, Md Yousuf, et al.
Publicado: (2024)
por: Harun, Md Yousuf, et al.
Publicado: (2024)
ViGText: Deepfake Image Detection with Vision-Language Model Explanations and Graph Neural Networks
por: ALBarqawi, Ahmad, et al.
Publicado: (2025)
por: ALBarqawi, Ahmad, et al.
Publicado: (2025)
What to align in multimodal contrastive learning?
por: Dufumier, Benoit, et al.
Publicado: (2024)
por: Dufumier, Benoit, et al.
Publicado: (2024)
What to Do Next? Memorizing skills from Egocentric Instructional Video
por: Bi, Jing, et al.
Publicado: (2025)
por: Bi, Jing, et al.
Publicado: (2025)
What Happens Next? Anticipating Future Motion by Generating Point Trajectories
por: Boduljak, Gabrijel, et al.
Publicado: (2025)
por: Boduljak, Gabrijel, et al.
Publicado: (2025)
What You See is What You GAN: Rendering Every Pixel for High-Fidelity Geometry in 3D GANs
por: Trevithick, Alex, et al.
Publicado: (2024)
por: Trevithick, Alex, et al.
Publicado: (2024)
Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off
por: Lee, Seungyong, et al.
Publicado: (2025)
por: Lee, Seungyong, et al.
Publicado: (2025)
Improving Interpretation Faithfulness for Vision Transformers
por: Hu, Lijie, et al.
Publicado: (2023)
por: Hu, Lijie, et al.
Publicado: (2023)
Block-Recurrent Dynamics in Vision Transformers
por: Jacobs, Mozes, et al.
Publicado: (2025)
por: Jacobs, Mozes, et al.
Publicado: (2025)
What's Holding Back Latent Visual Reasoning?
por: Viveiros, André G., et al.
Publicado: (2026)
por: Viveiros, André G., et al.
Publicado: (2026)
Learning What Matters: Steering Diffusion via Spectrally Anisotropic Forward Noise
por: Scimeca, Luca, et al.
Publicado: (2025)
por: Scimeca, Luca, et al.
Publicado: (2025)
On What Depends the Robustness of Multi-source Models to Missing Data in Earth Observation?
por: Mena, Francisco, et al.
Publicado: (2025)
por: Mena, Francisco, et al.
Publicado: (2025)
Explanation Bottleneck Models
por: Yamaguchi, Shin'ya, et al.
Publicado: (2024)
por: Yamaguchi, Shin'ya, et al.
Publicado: (2024)
Continual Adaptation of Vision Transformers for Federated Learning
por: Halbe, Shaunak, et al.
Publicado: (2023)
por: Halbe, Shaunak, et al.
Publicado: (2023)
Mechanisms of Non-Monotonic Scaling in Vision Transformers
por: Kumar, Anantha Padmanaban Krishna
Publicado: (2025)
por: Kumar, Anantha Padmanaban Krishna
Publicado: (2025)
Discovering Influential Neuron Path in Vision Transformers
por: Wang, Yifan, et al.
Publicado: (2025)
por: Wang, Yifan, et al.
Publicado: (2025)
DiffiT: Diffusion Vision Transformers for Image Generation
por: Hatamizadeh, Ali, et al.
Publicado: (2023)
por: Hatamizadeh, Ali, et al.
Publicado: (2023)
ADAPT to Robustify Prompt Tuning Vision Transformers
por: Eskandar, Masih, et al.
Publicado: (2024)
por: Eskandar, Masih, et al.
Publicado: (2024)
Accelerating Vision Transformers with Adaptive Patch Sizes
por: Choudhury, Rohan, et al.
Publicado: (2025)
por: Choudhury, Rohan, et al.
Publicado: (2025)
Class-Discriminative Attention Maps for Vision Transformers
por: Brocki, Lennart, et al.
Publicado: (2023)
por: Brocki, Lennart, et al.
Publicado: (2023)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
por: Lagos, Maximiliano Hormazábal, et al.
Publicado: (2025)
por: Lagos, Maximiliano Hormazábal, et al.
Publicado: (2025)
What Can We Learn from Inter-Annotator Variability in Skin Lesion Segmentation?
por: Abhishek, Kumar, et al.
Publicado: (2025)
por: Abhishek, Kumar, et al.
Publicado: (2025)
Two Complementary Perspectives to Continual Learning: Ask Not Only What to Optimize, But Also How
por: Hess, Timm, et al.
Publicado: (2023)
por: Hess, Timm, et al.
Publicado: (2023)
What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models
por: Farid, Karim, et al.
Publicado: (2025)
por: Farid, Karim, et al.
Publicado: (2025)
Guess What I Think: Streamlined EEG-to-Image Generation with Latent Diffusion Models
por: Lopez, Eleonora, et al.
Publicado: (2024)
por: Lopez, Eleonora, et al.
Publicado: (2024)
Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models
por: Lian, Chenyu, et al.
Publicado: (2025)
por: Lian, Chenyu, et al.
Publicado: (2025)
Intriguing Equivalence Structures of the Embedding Space of Vision Transformers
por: Salman, Shaeke, et al.
Publicado: (2024)
por: Salman, Shaeke, et al.
Publicado: (2024)
Oscillation-Reduced MXFP4 Training for Vision Transformers
por: Chen, Yuxiang, et al.
Publicado: (2025)
por: Chen, Yuxiang, et al.
Publicado: (2025)
Ejemplares similares
-
What You See is What You Classify: Black Box Attributions
por: Stalder, Steven, et al.
Publicado: (2022) -
LINE: LLM-based Iterative Neuron Explanations for Vision Models
por: Zaigrajew, Vladimir, et al.
Publicado: (2026) -
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
por: Bendikas, Rokas, et al.
Publicado: (2025) -
Probabilistic Conceptual Explainers: Trustworthy Conceptual Explanations for Vision Foundation Models
por: Wang, Hengyi, et al.
Publicado: (2024) -
When Shared Knowledge Hurts: Spectral Over-Accumulation in Model Merging
por: Li, Yayuan, et al.
Publicado: (2026)