SiNGER: A Clearer Voice Distills Vision Transformers Further
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Geunhyeok, Jeong, Sunjae, Choi, Yoonyoung, Kim, Jaeseung, Hwang, Hyoseok |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A2XP: Towards Private Domain Generalization
von: Yu, Geunhyeok, et al.
Veröffentlicht: (2023)
von: Yu, Geunhyeok, et al.
Veröffentlicht: (2023)
Scale-Consistent State-Space Dynamics via Fractal of Stationary Transformations
von: Yu, Geunhyeok, et al.
Veröffentlicht: (2026)
von: Yu, Geunhyeok, et al.
Veröffentlicht: (2026)
Is it safe to cross? Interpretable Risk Assessment with GPT-4V for Safety-Aware Street Crossing
von: Hwang, Hochul, et al.
Veröffentlicht: (2024)
von: Hwang, Hochul, et al.
Veröffentlicht: (2024)
Multimodal Distribution Matching for Vision-Language Dataset Distillation
von: Jeong, Jongoh, et al.
Veröffentlicht: (2026)
von: Jeong, Jongoh, et al.
Veröffentlicht: (2026)
CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space
von: Lim, Sohwi, et al.
Veröffentlicht: (2026)
von: Lim, Sohwi, et al.
Veröffentlicht: (2026)
Frequency-Aware Token Reduction for Efficient Vision Transformer
von: Lee, Dong-Jae, et al.
Veröffentlicht: (2025)
von: Lee, Dong-Jae, et al.
Veröffentlicht: (2025)
ESC: Erasing Space Concept for Knowledge Deletion
von: Lee, Tae-Young, et al.
Veröffentlicht: (2025)
von: Lee, Tae-Young, et al.
Veröffentlicht: (2025)
JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers
von: Byung-Ki, Kwon, et al.
Veröffentlicht: (2025)
von: Byung-Ki, Kwon, et al.
Veröffentlicht: (2025)
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
Continual Vision-and-Language Navigation
von: Jeong, Seongjun, et al.
Veröffentlicht: (2024)
von: Jeong, Seongjun, et al.
Veröffentlicht: (2024)
RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
DragText: Rethinking Text Embedding in Point-based Image Editing
von: Choi, Gayoon, et al.
Veröffentlicht: (2024)
von: Choi, Gayoon, et al.
Veröffentlicht: (2024)
Towards Optimal Trade-offs in Knowledge Distillation for CNNs and Vision Transformers at the Edge
von: Violos, John, et al.
Veröffentlicht: (2024)
von: Violos, John, et al.
Veröffentlicht: (2024)
Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
von: Lee, Dohun, et al.
Veröffentlicht: (2025)
von: Lee, Dohun, et al.
Veröffentlicht: (2025)
Image Clustering Conditioned on Text Criteria
von: Kwon, Sehyun, et al.
Veröffentlicht: (2023)
von: Kwon, Sehyun, et al.
Veröffentlicht: (2023)
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
von: Jung, Woojun, et al.
Veröffentlicht: (2025)
von: Jung, Woojun, et al.
Veröffentlicht: (2025)
Patch Rebirth: Toward Fast and Transferable Model Inversion of Vision Transformers
von: Heo, Seongsoo, et al.
Veröffentlicht: (2025)
von: Heo, Seongsoo, et al.
Veröffentlicht: (2025)
Adversarial Prompt Distillation for Vision-Language Models
von: Luo, Lin, et al.
Veröffentlicht: (2024)
von: Luo, Lin, et al.
Veröffentlicht: (2024)
FEAST: Fully Connected Expressive Attention for Spatial Transcriptomics
von: Jeong, Taejin, et al.
Veröffentlicht: (2026)
von: Jeong, Taejin, et al.
Veröffentlicht: (2026)
Mixed Non-linear Quantization for Vision Transformers
von: Kim, Gihwan, et al.
Veröffentlicht: (2024)
von: Kim, Gihwan, et al.
Veröffentlicht: (2024)
Reflexive Guidance: Improving OoDD in Vision-Language Models via Self-Guided Image-Adaptive Concept Generation
von: Kim, Jihyo, et al.
Veröffentlicht: (2024)
von: Kim, Jihyo, et al.
Veröffentlicht: (2024)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
von: Kang, Seil, et al.
Veröffentlicht: (2025)
von: Kang, Seil, et al.
Veröffentlicht: (2025)
X-Distill: Cross-Architecture Vision Distillation for Visuomotor Learning
von: Shao, Maanping, et al.
Veröffentlicht: (2026)
von: Shao, Maanping, et al.
Veröffentlicht: (2026)
SIA: Enhancing Safety via Intent Awareness for Vision-Language Models
von: Na, Youngjin, et al.
Veröffentlicht: (2025)
von: Na, Youngjin, et al.
Veröffentlicht: (2025)
Exploiting Style Latent Flows for Generalizing Deepfake Video Detection
von: Choi, Jongwook, et al.
Veröffentlicht: (2024)
von: Choi, Jongwook, et al.
Veröffentlicht: (2024)
SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining
von: Song, Chull Hwan, et al.
Veröffentlicht: (2024)
von: Song, Chull Hwan, et al.
Veröffentlicht: (2024)
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
von: Jeong, Suchae, et al.
Veröffentlicht: (2025)
von: Jeong, Suchae, et al.
Veröffentlicht: (2025)
IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision Transformers
von: Kim, Gihwan, et al.
Veröffentlicht: (2025)
von: Kim, Gihwan, et al.
Veröffentlicht: (2025)
Anchoring and Rescaling Attention for Semantically Coherent Inbetweening
von: Choi, Tae Eun, et al.
Veröffentlicht: (2026)
von: Choi, Tae Eun, et al.
Veröffentlicht: (2026)
Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision
von: Kim, Jiyeong, et al.
Veröffentlicht: (2026)
von: Kim, Jiyeong, et al.
Veröffentlicht: (2026)
MoST: Motion Style Transformer between Diverse Action Contents
von: Kim, Boeun, et al.
Veröffentlicht: (2024)
von: Kim, Boeun, et al.
Veröffentlicht: (2024)
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment
von: Jeong, Seongjun, et al.
Veröffentlicht: (2024)
von: Jeong, Seongjun, et al.
Veröffentlicht: (2024)
Representation Separation for Semantic Segmentation with Vision Transformers
von: Hong, Yuanduo, et al.
Veröffentlicht: (2022)
von: Hong, Yuanduo, et al.
Veröffentlicht: (2022)
Dynamic Weight Adjustment for Knowledge Distillation: Leveraging Vision Transformer for High-Accuracy Lung Cancer Detection and Real-Time Deployment
von: Khan, Saif Ur Rehman, et al.
Veröffentlicht: (2025)
von: Khan, Saif Ur Rehman, et al.
Veröffentlicht: (2025)
LMLT: Low-to-high Multi-Level Vision Transformer for Image Super-Resolution
von: Kim, Jeongsoo, et al.
Veröffentlicht: (2024)
von: Kim, Jeongsoo, et al.
Veröffentlicht: (2024)
Spiking Vision Transformer with Saccadic Attention
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
CaddieSet: A Golf Swing Dataset with Human Joint Features and Ball Information
von: Jung, Seunghyeon, et al.
Veröffentlicht: (2025)
von: Jung, Seunghyeon, et al.
Veröffentlicht: (2025)
Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation
von: Yang, Yang, et al.
Veröffentlicht: (2024)
von: Yang, Yang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A2XP: Towards Private Domain Generalization
von: Yu, Geunhyeok, et al.
Veröffentlicht: (2023) -
Scale-Consistent State-Space Dynamics via Fractal of Stationary Transformations
von: Yu, Geunhyeok, et al.
Veröffentlicht: (2026) -
Is it safe to cross? Interpretable Risk Assessment with GPT-4V for Safety-Aware Street Crossing
von: Hwang, Hochul, et al.
Veröffentlicht: (2024) -
Multimodal Distribution Matching for Vision-Language Dataset Distillation
von: Jeong, Jongoh, et al.
Veröffentlicht: (2026) -
CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space
von: Lim, Sohwi, et al.
Veröffentlicht: (2026)