Interpreting vision transformers via residual replacement model
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jinyeong, Kim, Junhyeok, Shim, Yumin, Kim, Joohyeok, Jung, Sunyoung, Hwang, Seong Jae |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FEAST: Fully Connected Expressive Attention for Spatial Transcriptomics
by: Jeong, Taejin, et al.
Published: (2026)
by: Jeong, Taejin, et al.
Published: (2026)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Real-Time Visual Attribution Streaming in Thinking Model
by: Kang, Seil, et al.
Published: (2026)
by: Kang, Seil, et al.
Published: (2026)
Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
by: Choi, Tae Eun, et al.
Published: (2026)
by: Choi, Tae Eun, et al.
Published: (2026)
Rethinking Glaucoma Calibration: Voting-Based Binocular and Metadata Integration
by: Jeong, Taejin, et al.
Published: (2025)
by: Jeong, Taejin, et al.
Published: (2025)
Advancing Text-Driven Chest X-Ray Generation with Policy-Based Reinforcement Learning
by: Han, Woojung, et al.
Published: (2024)
by: Han, Woojung, et al.
Published: (2024)
PLATYPUS: Progressive Local Surface Estimator for Arbitrary-Scale Point Cloud Upsampling
by: Kim, Donghyun, et al.
Published: (2024)
by: Kim, Donghyun, et al.
Published: (2024)
Is it safe to cross? Interpretable Risk Assessment with GPT-4V for Safety-Aware Street Crossing
by: Hwang, Hochul, et al.
Published: (2024)
by: Hwang, Hochul, et al.
Published: (2024)
Classifier-guided CLIP Distillation for Unsupervised Multi-label Classification
by: Kim, Dongseob, et al.
Published: (2025)
by: Kim, Dongseob, et al.
Published: (2025)
Solution for SMART-101 Challenge of CVPR Multi-modal Algorithmic Reasoning Task 2024
by: Ahn, Jinwoo, et al.
Published: (2024)
by: Ahn, Jinwoo, et al.
Published: (2024)
Parameter Efficient Fine Tuning for Multi-scanner PET to PET Reconstruction
by: Kim, Yumin, et al.
Published: (2024)
by: Kim, Yumin, et al.
Published: (2024)
Pathology-Aware Adaptive Watermarking for Text-Driven Medical Image Synthesis
by: Kim, Chanyoung, et al.
Published: (2025)
by: Kim, Chanyoung, et al.
Published: (2025)
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
by: Kim, Junhyeok, et al.
Published: (2025)
by: Kim, Junhyeok, et al.
Published: (2025)
Learning Primitive Relations for Compositional Zero-Shot Learning
by: Lee, Insu, et al.
Published: (2025)
by: Lee, Insu, et al.
Published: (2025)
Parallel Rescaling: Rebalancing Consistency Guidance for Personalized Diffusion Models
by: Chae, JungWoo, et al.
Published: (2025)
by: Chae, JungWoo, et al.
Published: (2025)
Towards Continuous Sign Language Conversation from Isolated Signs
by: Kim, Youngmin, et al.
Published: (2026)
by: Kim, Youngmin, et al.
Published: (2026)
Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
by: Jun, Youngjun, et al.
Published: (2026)
by: Jun, Youngjun, et al.
Published: (2026)
Naïve Exposure of Generative AI Capabilities Undermines Deepfake Detection
by: Kim, Sunpill, et al.
Published: (2026)
by: Kim, Sunpill, et al.
Published: (2026)
Rethinking Data Augmentation for Robust LiDAR Semantic Segmentation in Adverse Weather
by: Park, Junsung, et al.
Published: (2024)
by: Park, Junsung, et al.
Published: (2024)
Re-Scoring Using Image-Language Similarity for Few-Shot Object Detection
by: Jung, Min Jae, et al.
Published: (2023)
by: Jung, Min Jae, et al.
Published: (2023)
Clustering-based Image-Text Graph Matching for Domain Generalization
by: Park, Nokyung, et al.
Published: (2023)
by: Park, Nokyung, et al.
Published: (2023)
GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
by: Kim, Nayeong, et al.
Published: (2025)
by: Kim, Nayeong, et al.
Published: (2025)
Localized Concept Erasure in Text-to-Image Diffusion Models via High-Level Representation Misdirection
by: Lee, Uichan, et al.
Published: (2026)
by: Lee, Uichan, et al.
Published: (2026)
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
by: Kim, Donghoon, et al.
Published: (2025)
by: Kim, Donghoon, et al.
Published: (2025)
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
by: Kim, Jeonghyeon, et al.
Published: (2025)
by: Kim, Jeonghyeon, et al.
Published: (2025)
Synergistic Integration of Coordinate Network and Tensorial Feature for Improving Neural Radiance Fields from Sparse Inputs
by: Kim, Mingyu, et al.
Published: (2024)
by: Kim, Mingyu, et al.
Published: (2024)
Leveraging Textual Compositional Reasoning for Robust Change Captioning
by: Park, Kyu Ri, et al.
Published: (2025)
by: Park, Kyu Ri, et al.
Published: (2025)
Two Birds, One Projection: Harmonizing Safety and Utility in LVLMs via Inference-time Feature Projection
by: Han, Yewon, et al.
Published: (2026)
by: Han, Yewon, et al.
Published: (2026)
Reflexive Guidance: Improving OoDD in Vision-Language Models via Self-Guided Image-Adaptive Concept Generation
by: Kim, Jihyo, et al.
Published: (2024)
by: Kim, Jihyo, et al.
Published: (2024)
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
by: Park, Geon, et al.
Published: (2025)
by: Park, Geon, et al.
Published: (2025)
FALCON: False-Negative Aware Learning of Contrastive Negatives in Vision-Language Alignment
by: Kim, Myunsoo, et al.
Published: (2025)
by: Kim, Myunsoo, et al.
Published: (2025)
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
by: Kim, Ji-Hyeon, et al.
Published: (2026)
by: Kim, Ji-Hyeon, et al.
Published: (2026)
GradMix: Gradient-based Selective Mixup for Robust Data Augmentation in Class-Incremental Learning
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
by: Kim, Youngmin, et al.
Published: (2025)
by: Kim, Youngmin, et al.
Published: (2025)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
by: Lim, Byeonggeuk, et al.
Published: (2026)
by: Lim, Byeonggeuk, et al.
Published: (2026)
ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approximation
by: Kim, Suyoung, et al.
Published: (2026)
by: Kim, Suyoung, et al.
Published: (2026)
DragText: Rethinking Text Embedding in Point-based Image Editing
by: Choi, Gayoon, et al.
Published: (2024)
by: Choi, Gayoon, et al.
Published: (2024)
Real-Time Long Horizon Air Quality Forecasting via Group-Relative Policy Optimization
by: Kang, Inha, et al.
Published: (2025)
by: Kang, Inha, et al.
Published: (2025)
Similar Items
-
FEAST: Fully Connected Expressive Attention for Spatial Transcriptomics
by: Jeong, Taejin, et al.
Published: (2026) -
See What You Are Told: Visual Attention Sink in Large Multimodal Models
by: Kang, Seil, et al.
Published: (2025) -
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
by: Kang, Seil, et al.
Published: (2025) -
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025) -
Real-Time Visual Attribution Streaming in Thinking Model
by: Kang, Seil, et al.
Published: (2026)