Scratching Visual Transformer's Back with Uniform Attention
Fuente:
arXiv
Salvato in:
| Autori principali: | Hyeon-Woo, Nam, Yu-Ji, Kim, Heo, Byeongho, Han, Dongyoon, Oh, Seong Joon, Oh, Tae-Hyun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2022
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Morphing Tokens Draw Strong Masked Image Models
di: Kim, Taekyung, et al.
Pubblicazione: (2023)
di: Kim, Taekyung, et al.
Pubblicazione: (2023)
DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs
di: Kim, Donghyun, et al.
Pubblicazione: (2024)
di: Kim, Donghyun, et al.
Pubblicazione: (2024)
Learning with Unmasked Tokens Drives Stronger Vision Learners
di: Kim, Taekyung, et al.
Pubblicazione: (2023)
di: Kim, Taekyung, et al.
Pubblicazione: (2023)
Rotary Position Embedding for Vision Transformer
di: Heo, Byeongho, et al.
Pubblicazione: (2024)
di: Heo, Byeongho, et al.
Pubblicazione: (2024)
Learning Correlation-aware Aleatoric Uncertainty for 3D Hand Pose Estimation
di: Chae-Yeon, Lee, et al.
Pubblicazione: (2025)
di: Chae-Yeon, Lee, et al.
Pubblicazione: (2025)
Masking meets Supervision: A Strong Learning Alliance
di: Heo, Byeongho, et al.
Pubblicazione: (2023)
di: Heo, Byeongho, et al.
Pubblicazione: (2023)
VLM's Eye Examination: Instruct and Inspect Visual Competency of Vision Language Models
di: Hyeon-Woo, Nam, et al.
Pubblicazione: (2024)
di: Hyeon-Woo, Nam, et al.
Pubblicazione: (2024)
Exploring Conditions for Diffusion models in Robotic Control
di: Shin, Heeseong, et al.
Pubblicazione: (2025)
di: Shin, Heeseong, et al.
Pubblicazione: (2025)
Token Bottleneck: One Token to Remember Dynamics
di: Kim, Taekyung, et al.
Pubblicazione: (2025)
di: Kim, Taekyung, et al.
Pubblicazione: (2025)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
di: Song, Junha, et al.
Pubblicazione: (2026)
di: Song, Junha, et al.
Pubblicazione: (2026)
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
di: Kim, Jiwon, et al.
Pubblicazione: (2023)
di: Kim, Jiwon, et al.
Pubblicazione: (2023)
DNNs May Determine Major Properties of Their Outputs Early, with Timing Possibly Driven by Bias
di: Park, Song, et al.
Pubblicazione: (2025)
di: Park, Song, et al.
Pubblicazione: (2025)
BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-language Models
di: Ye-Bin, Moon, et al.
Pubblicazione: (2024)
di: Ye-Bin, Moon, et al.
Pubblicazione: (2024)
SeiT++: Masked Token Modeling Improves Storage-efficient Training
di: Lee, Minhyun, et al.
Pubblicazione: (2023)
di: Lee, Minhyun, et al.
Pubblicazione: (2023)
GaussExplorer: 3D Gaussian Splatting for Embodied Exploration and Reasoning
di: Yu-Ji, Kim, et al.
Pubblicazione: (2026)
di: Yu-Ji, Kim, et al.
Pubblicazione: (2026)
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
di: Sung-Bin, Kim, et al.
Pubblicazione: (2024)
di: Sung-Bin, Kim, et al.
Pubblicazione: (2024)
Early Failure Detection and Intervention in Video Diffusion Models
di: Byung-Ki, Kwon, et al.
Pubblicazione: (2026)
di: Byung-Ki, Kwon, et al.
Pubblicazione: (2026)
RL makes MLLMs see better than SFT
di: Song, Junha, et al.
Pubblicazione: (2025)
di: Song, Junha, et al.
Pubblicazione: (2025)
MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation
di: Lee, Minhyun, et al.
Pubblicazione: (2024)
di: Lee, Minhyun, et al.
Pubblicazione: (2024)
SYNAuG: Exploiting Synthetic Data for Data Imbalance Problems
di: Ye-Bin, Moon, et al.
Pubblicazione: (2023)
di: Ye-Bin, Moon, et al.
Pubblicazione: (2023)
Enhancing Speech-Driven 3D Facial Animation with Audio-Visual Guidance from Lip Reading Expert
di: EunGi, Han, et al.
Pubblicazione: (2024)
di: EunGi, Han, et al.
Pubblicazione: (2024)
Balancing Efficiency and Quality: MoEISR for Arbitrary-Scale Image Super-Resolution
di: Oh, Young Jae, et al.
Pubblicazione: (2023)
di: Oh, Young Jae, et al.
Pubblicazione: (2023)
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
di: Choi, Jeongsoo, et al.
Pubblicazione: (2025)
di: Choi, Jeongsoo, et al.
Pubblicazione: (2025)
Lipsum-FT: Robust Fine-Tuning of Zero-Shot Models Using Random Text Guidance
di: Nam, Giung, et al.
Pubblicazione: (2024)
di: Nam, Giung, et al.
Pubblicazione: (2024)
VSC: Visual Search Compositional Text-to-Image Diffusion Model
di: Dat, Do Huu, et al.
Pubblicazione: (2025)
di: Dat, Do Huu, et al.
Pubblicazione: (2025)
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering
di: Lim, Su Hyeon, et al.
Pubblicazione: (2024)
di: Lim, Su Hyeon, et al.
Pubblicazione: (2024)
WWW: A Unified Framework for Explaining What, Where and Why of Neural Networks by Interpretation of Neuron Concepts
di: Ahn, Yong Hyun, et al.
Pubblicazione: (2024)
di: Ahn, Yong Hyun, et al.
Pubblicazione: (2024)
Mask-Free Neuron Concept Annotation for Interpreting Neural Networks in Medical Domain
di: Kim, Hyeon Bae, et al.
Pubblicazione: (2024)
di: Kim, Hyeon Bae, et al.
Pubblicazione: (2024)
Perceptually Accurate 3D Talking Head Generation: New Definitions, Speech-Mesh Representation, and Evaluation Metrics
di: Chae-Yeon, Lee, et al.
Pubblicazione: (2025)
di: Chae-Yeon, Lee, et al.
Pubblicazione: (2025)
SA-ResGS: Self-Augmented Residual 3D Gaussian Splatting for Next Best View Selection
di: Jun-Seong, Kim, et al.
Pubblicazione: (2026)
di: Jun-Seong, Kim, et al.
Pubblicazione: (2026)
Learning-based Axial Video Motion Magnification
di: Byung-Ki, Kwon, et al.
Pubblicazione: (2023)
di: Byung-Ki, Kwon, et al.
Pubblicazione: (2023)
Similarity of Neural Architectures using Adversarial Attack Transferability
di: Hwang, Jaehui, et al.
Pubblicazione: (2022)
di: Hwang, Jaehui, et al.
Pubblicazione: (2022)
Dr. Splat: Directly Referring 3D Gaussian Splatting via Direct Language Embedding Registration
di: Jun-Seong, Kim, et al.
Pubblicazione: (2025)
di: Jun-Seong, Kim, et al.
Pubblicazione: (2025)
GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
di: Kim, Nayeong, et al.
Pubblicazione: (2025)
di: Kim, Nayeong, et al.
Pubblicazione: (2025)
Pretrained Visual Uncertainties
di: Kirchhof, Michael, et al.
Pubblicazione: (2024)
di: Kirchhof, Michael, et al.
Pubblicazione: (2024)
HDR-NSFF: High Dynamic Range Neural Scene Flow Fields
di: Dong-Yeon, Shin, et al.
Pubblicazione: (2026)
di: Dong-Yeon, Shin, et al.
Pubblicazione: (2026)
Factorized Multi-Resolution HashGrid for Efficient Neural Radiance Fields: Execution on Edge-Devices
di: Jun-Seong, Kim, et al.
Pubblicazione: (2026)
di: Jun-Seong, Kim, et al.
Pubblicazione: (2026)
Half-Truths Break Similarity-Based Retrieval
di: Kargi, Bora, et al.
Pubblicazione: (2026)
di: Kargi, Bora, et al.
Pubblicazione: (2026)
On the rankability of visual embeddings
di: Sonthalia, Ankit, et al.
Pubblicazione: (2025)
di: Sonthalia, Ankit, et al.
Pubblicazione: (2025)
SoundBrush: Sound as a Brush for Visual Scene Editing
di: Sung-Bin, Kim, et al.
Pubblicazione: (2024)
di: Sung-Bin, Kim, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Morphing Tokens Draw Strong Masked Image Models
di: Kim, Taekyung, et al.
Pubblicazione: (2023) -
DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs
di: Kim, Donghyun, et al.
Pubblicazione: (2024) -
Learning with Unmasked Tokens Drives Stronger Vision Learners
di: Kim, Taekyung, et al.
Pubblicazione: (2023) -
Rotary Position Embedding for Vision Transformer
di: Heo, Byeongho, et al.
Pubblicazione: (2024) -
Learning Correlation-aware Aleatoric Uncertainty for 3D Hand Pose Estimation
di: Chae-Yeon, Lee, et al.
Pubblicazione: (2025)