HeSS: Head Sensitivity Score for Sparsity Redistribution in VGGT
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Yongsung, Song, Wooseok, Lew, Jaihyun, Hwangbo, Hun, Lee, Jaehoon, Yoon, Sungroh |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unsupervised Homography Estimation on Multimodal Image Pair via Alternating Optimization
by: Song, Sanghyeob, et al.
Published: (2024)
by: Song, Sanghyeob, et al.
Published: (2024)
Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
by: Baek, Kanghyun, et al.
Published: (2026)
by: Baek, Kanghyun, et al.
Published: (2026)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
Disentangled Motion Modeling for Video Frame Interpolation
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
Improving Geometry in Sparse-View 3DGS via Reprojection-based DoF Separation
by: Kim, Yongsung, et al.
Published: (2024)
by: Kim, Yongsung, et al.
Published: (2024)
ControlDreamer: Blending Geometry and Style in Text-to-3D
by: Oh, Yeongtak, et al.
Published: (2023)
by: Oh, Yeongtak, et al.
Published: (2023)
Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs
by: Lee, Jaehoon, et al.
Published: (2026)
by: Lee, Jaehoon, et al.
Published: (2026)
CKNN: Cleansed k-Nearest Neighbor for Unsupervised Video Anomaly Detection
by: Yi, Jihun, et al.
Published: (2024)
by: Yi, Jihun, et al.
Published: (2024)
HTTM: Head-wise Temporal Token Merging for Faster VGGT
by: Wang, Weitian, et al.
Published: (2025)
by: Wang, Weitian, et al.
Published: (2025)
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
by: Sun, Xiangyu, et al.
Published: (2026)
by: Sun, Xiangyu, et al.
Published: (2026)
VGGT-HPE: Reframing Head Pose Estimation as Relative Pose Prediction
by: Vasileiou, Vasiliki, et al.
Published: (2026)
by: Vasileiou, Vasiliki, et al.
Published: (2026)
Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization
by: Oh, Yeongtak, et al.
Published: (2026)
by: Oh, Yeongtak, et al.
Published: (2026)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)
by: Jung, Mingi, et al.
Published: (2025)
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
by: Lee, Saehyung, et al.
Published: (2024)
by: Lee, Saehyung, et al.
Published: (2024)
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
by: Shin, Chaehun, et al.
Published: (2024)
by: Shin, Chaehun, et al.
Published: (2024)
On mitigating stability-plasticity dilemma in CLIP-guided image morphing via geodesic distillation loss
by: Oh, Yeongtak, et al.
Published: (2024)
by: Oh, Yeongtak, et al.
Published: (2024)
Normality Addition via Normality Detection in Industrial Image Anomaly Detection Models
by: Yi, Jihun, et al.
Published: (2024)
by: Yi, Jihun, et al.
Published: (2024)
VGGT-$Ω$
by: Wang, Jianyuan, et al.
Published: (2026)
by: Wang, Jianyuan, et al.
Published: (2026)
GaussianTalker: Real-Time High-Fidelity Talking Head Synthesis with Audio-Driven 3D Gaussian Splatting
by: Cho, Kyusun, et al.
Published: (2024)
by: Cho, Kyusun, et al.
Published: (2024)
VGGT-X: When VGGT Meets Dense Novel View Synthesis
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
Style-Friendly SNR Sampler for Style-Driven Generation
by: Choi, Jooyoung, et al.
Published: (2024)
by: Choi, Jooyoung, et al.
Published: (2024)
VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments
by: Xu, Jingyi, et al.
Published: (2026)
by: Xu, Jingyi, et al.
Published: (2026)
FrameVGGT: Geometry-Aligned Frame-Level Memory for Bounded Streaming VGGT
by: Xu, Zhisong, et al.
Published: (2026)
by: Xu, Zhisong, et al.
Published: (2026)
VGGT-SLAM++
by: Mandal, Avilasha, et al.
Published: (2026)
by: Mandal, Avilasha, et al.
Published: (2026)
Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach
by: Lee, Saehyung, et al.
Published: (2024)
by: Lee, Saehyung, et al.
Published: (2024)
StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression
by: Liu, Xuanyi, et al.
Published: (2026)
by: Liu, Xuanyi, et al.
Published: (2026)
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
by: Shu, Zhijian, et al.
Published: (2025)
by: Shu, Zhijian, et al.
Published: (2025)
Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation
by: Seo, Junyoung, et al.
Published: (2023)
by: Seo, Junyoung, et al.
Published: (2023)
Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
by: Park, Junsung, et al.
Published: (2025)
by: Park, Junsung, et al.
Published: (2025)
Alignment Scores: Robust Metrics for Multiview Pose Accuracy Evaluation
by: Lee, Seong Hun, et al.
Published: (2024)
by: Lee, Seong Hun, et al.
Published: (2024)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
by: Deng, Kai, et al.
Published: (2025)
by: Deng, Kai, et al.
Published: (2025)
Retrieval-Augmented Score Distillation for Text-to-3D Generation
by: Seo, Junyoung, et al.
Published: (2024)
by: Seo, Junyoung, et al.
Published: (2024)
VGGT: Visual Geometry Grounded Transformer
by: Wang, Jianyuan, et al.
Published: (2025)
by: Wang, Jianyuan, et al.
Published: (2025)
Dense Semantic Matching with VGGT Prior
by: Yang, Songlin, et al.
Published: (2025)
by: Yang, Songlin, et al.
Published: (2025)
Efficient Diffusion-Driven Corruption Editor for Test-Time Adaptation
by: Oh, Yeongtak, et al.
Published: (2024)
by: Oh, Yeongtak, et al.
Published: (2024)
Textual Training for the Hassle-Free Removal of Unwanted Visual Data: Case Studies on OOD and Hateful Image Detection
by: Lee, Saehyung, et al.
Published: (2024)
by: Lee, Saehyung, et al.
Published: (2024)
Talk3D: High-Fidelity Talking Portrait Synthesis via Personalized 3D Generative Prior
by: Ko, Jaehoon, et al.
Published: (2024)
by: Ko, Jaehoon, et al.
Published: (2024)
AVGGT: Rethinking Global Attention for Accelerating VGGT
by: Sun, Xianbing, et al.
Published: (2025)
by: Sun, Xianbing, et al.
Published: (2025)
Similar Items
-
Unsupervised Homography Estimation on Multimodal Image Pair via Alternating Optimization
by: Song, Sanghyeob, et al.
Published: (2024) -
Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
by: Baek, Kanghyun, et al.
Published: (2026) -
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
by: Lew, Jaihyun, et al.
Published: (2024) -
Disentangled Motion Modeling for Video Frame Interpolation
by: Lew, Jaihyun, et al.
Published: (2024) -
Improving Geometry in Sparse-View 3DGS via Reprojection-based DoF Separation
by: Kim, Yongsung, et al.
Published: (2024)