SiNGER: A Clearer Voice Distills Vision Transformers Further
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Geunhyeok, Jeong, Sunjae, Choi, Yoonyoung, Kim, Jaeseung, Hwang, Hyoseok |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A2XP: Towards Private Domain Generalization
by: Yu, Geunhyeok, et al.
Published: (2023)
by: Yu, Geunhyeok, et al.
Published: (2023)
Scale-Consistent State-Space Dynamics via Fractal of Stationary Transformations
by: Yu, Geunhyeok, et al.
Published: (2026)
by: Yu, Geunhyeok, et al.
Published: (2026)
Is it safe to cross? Interpretable Risk Assessment with GPT-4V for Safety-Aware Street Crossing
by: Hwang, Hochul, et al.
Published: (2024)
by: Hwang, Hochul, et al.
Published: (2024)
Multimodal Distribution Matching for Vision-Language Dataset Distillation
by: Jeong, Jongoh, et al.
Published: (2026)
by: Jeong, Jongoh, et al.
Published: (2026)
CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space
by: Lim, Sohwi, et al.
Published: (2026)
by: Lim, Sohwi, et al.
Published: (2026)
Frequency-Aware Token Reduction for Efficient Vision Transformer
by: Lee, Dong-Jae, et al.
Published: (2025)
by: Lee, Dong-Jae, et al.
Published: (2025)
ESC: Erasing Space Concept for Knowledge Deletion
by: Lee, Tae-Young, et al.
Published: (2025)
by: Lee, Tae-Young, et al.
Published: (2025)
JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers
by: Byung-Ki, Kwon, et al.
Published: (2025)
by: Byung-Ki, Kwon, et al.
Published: (2025)
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
by: Lee, Seonho, et al.
Published: (2025)
by: Lee, Seonho, et al.
Published: (2025)
Continual Vision-and-Language Navigation
by: Jeong, Seongjun, et al.
Published: (2024)
by: Jeong, Seongjun, et al.
Published: (2024)
RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2024)
by: Woo, Sangmin, et al.
Published: (2024)
DragText: Rethinking Text Embedding in Point-based Image Editing
by: Choi, Gayoon, et al.
Published: (2024)
by: Choi, Gayoon, et al.
Published: (2024)
Towards Optimal Trade-offs in Knowledge Distillation for CNNs and Vision Transformers at the Edge
by: Violos, John, et al.
Published: (2024)
by: Violos, John, et al.
Published: (2024)
Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
by: Lee, Dohun, et al.
Published: (2025)
by: Lee, Dohun, et al.
Published: (2025)
Image Clustering Conditioned on Text Criteria
by: Kwon, Sehyun, et al.
Published: (2023)
by: Kwon, Sehyun, et al.
Published: (2023)
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
by: Jung, Woojun, et al.
Published: (2025)
by: Jung, Woojun, et al.
Published: (2025)
Patch Rebirth: Toward Fast and Transferable Model Inversion of Vision Transformers
by: Heo, Seongsoo, et al.
Published: (2025)
by: Heo, Seongsoo, et al.
Published: (2025)
Adversarial Prompt Distillation for Vision-Language Models
by: Luo, Lin, et al.
Published: (2024)
by: Luo, Lin, et al.
Published: (2024)
FEAST: Fully Connected Expressive Attention for Spatial Transcriptomics
by: Jeong, Taejin, et al.
Published: (2026)
by: Jeong, Taejin, et al.
Published: (2026)
Mixed Non-linear Quantization for Vision Transformers
by: Kim, Gihwan, et al.
Published: (2024)
by: Kim, Gihwan, et al.
Published: (2024)
Reflexive Guidance: Improving OoDD in Vision-Language Models via Self-Guided Image-Adaptive Concept Generation
by: Kim, Jihyo, et al.
Published: (2024)
by: Kim, Jihyo, et al.
Published: (2024)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2024)
by: Woo, Sangmin, et al.
Published: (2024)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
X-Distill: Cross-Architecture Vision Distillation for Visuomotor Learning
by: Shao, Maanping, et al.
Published: (2026)
by: Shao, Maanping, et al.
Published: (2026)
SIA: Enhancing Safety via Intent Awareness for Vision-Language Models
by: Na, Youngjin, et al.
Published: (2025)
by: Na, Youngjin, et al.
Published: (2025)
Exploiting Style Latent Flows for Generalizing Deepfake Video Detection
by: Choi, Jongwook, et al.
Published: (2024)
by: Choi, Jongwook, et al.
Published: (2024)
SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining
by: Song, Chull Hwan, et al.
Published: (2024)
by: Song, Chull Hwan, et al.
Published: (2024)
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
by: Kim, Jiwan, et al.
Published: (2025)
by: Kim, Jiwan, et al.
Published: (2025)
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
by: Jeong, Suchae, et al.
Published: (2025)
by: Jeong, Suchae, et al.
Published: (2025)
IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision Transformers
by: Kim, Gihwan, et al.
Published: (2025)
by: Kim, Gihwan, et al.
Published: (2025)
Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
by: Choi, Tae Eun, et al.
Published: (2026)
by: Choi, Tae Eun, et al.
Published: (2026)
Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision
by: Kim, Jiyeong, et al.
Published: (2026)
by: Kim, Jiyeong, et al.
Published: (2026)
MoST: Motion Style Transformer between Diverse Action Contents
by: Kim, Boeun, et al.
Published: (2024)
by: Kim, Boeun, et al.
Published: (2024)
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment
by: Jeong, Seongjun, et al.
Published: (2024)
by: Jeong, Seongjun, et al.
Published: (2024)
Representation Separation for Semantic Segmentation with Vision Transformers
by: Hong, Yuanduo, et al.
Published: (2022)
by: Hong, Yuanduo, et al.
Published: (2022)
Dynamic Weight Adjustment for Knowledge Distillation: Leveraging Vision Transformer for High-Accuracy Lung Cancer Detection and Real-Time Deployment
by: Khan, Saif Ur Rehman, et al.
Published: (2025)
by: Khan, Saif Ur Rehman, et al.
Published: (2025)
LMLT: Low-to-high Multi-Level Vision Transformer for Image Super-Resolution
by: Kim, Jeongsoo, et al.
Published: (2024)
by: Kim, Jeongsoo, et al.
Published: (2024)
Spiking Vision Transformer with Saccadic Attention
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
CaddieSet: A Golf Swing Dataset with Human Joint Features and Ball Information
by: Jung, Seunghyeon, et al.
Published: (2025)
by: Jung, Seunghyeon, et al.
Published: (2025)
Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation
by: Yang, Yang, et al.
Published: (2024)
by: Yang, Yang, et al.
Published: (2024)
Similar Items
-
A2XP: Towards Private Domain Generalization
by: Yu, Geunhyeok, et al.
Published: (2023) -
Scale-Consistent State-Space Dynamics via Fractal of Stationary Transformations
by: Yu, Geunhyeok, et al.
Published: (2026) -
Is it safe to cross? Interpretable Risk Assessment with GPT-4V for Safety-Aware Street Crossing
by: Hwang, Hochul, et al.
Published: (2024) -
Multimodal Distribution Matching for Vision-Language Dataset Distillation
by: Jeong, Jongoh, et al.
Published: (2026) -
CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space
by: Lim, Sohwi, et al.
Published: (2026)