Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Kuzucu, Selim, Naeem, Muhammad Ferjad, Kukleva, Anna, Tombari, Federico, Schiele, Bernt |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
di: Kuzucu, Selim, et al.
Pubblicazione: (2026)
di: Kuzucu, Selim, et al.
Pubblicazione: (2026)
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
di: Kukleva, Anna, et al.
Pubblicazione: (2025)
di: Kukleva, Anna, et al.
Pubblicazione: (2025)
Toward a Diffusion-Based Generalist for Dense Vision Tasks
di: Fan, Yue, et al.
Pubblicazione: (2024)
di: Fan, Yue, et al.
Pubblicazione: (2024)
Learning to Prompt with Text Only Supervision for Vision-Language Models
di: Khattak, Muhammad Uzair, et al.
Pubblicazione: (2024)
di: Khattak, Muhammad Uzair, et al.
Pubblicazione: (2024)
OrCo: Towards Better Generalization via Orthogonality and Contrast for Few-Shot Class-Incremental Learning
di: Ahmed, Noor, et al.
Pubblicazione: (2024)
di: Ahmed, Noor, et al.
Pubblicazione: (2024)
GiT: Towards Generalist Vision Transformer through Universal Language Interface
di: Wang, Haiyang, et al.
Pubblicazione: (2024)
di: Wang, Haiyang, et al.
Pubblicazione: (2024)
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models
di: Xie, Jiahao, et al.
Pubblicazione: (2026)
di: Xie, Jiahao, et al.
Pubblicazione: (2026)
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
di: Shvetsova, Nina, et al.
Pubblicazione: (2023)
di: Shvetsova, Nina, et al.
Pubblicazione: (2023)
Do Instance Priors Help Weakly Supervised Semantic Segmentation?
di: Das, Anurag, et al.
Pubblicazione: (2026)
di: Das, Anurag, et al.
Pubblicazione: (2026)
VAT: Vision Action Transformer by Unlocking Full Representation of ViT
di: Li, Wenhao, et al.
Pubblicazione: (2025)
di: Li, Wenhao, et al.
Pubblicazione: (2025)
MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data
di: Sheludzko, Siarhei, et al.
Pubblicazione: (2026)
di: Sheludzko, Siarhei, et al.
Pubblicazione: (2026)
DAVE: Distribution-aware Attribution via ViT Gradient Decomposition
di: Wróbel, Adam, et al.
Pubblicazione: (2026)
di: Wróbel, Adam, et al.
Pubblicazione: (2026)
ViT$^3$: Unlocking Test-Time Training in Vision
di: Han, Dongchen, et al.
Pubblicazione: (2025)
di: Han, Dongchen, et al.
Pubblicazione: (2025)
R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs
di: Xie, Jiahao, et al.
Pubblicazione: (2026)
di: Xie, Jiahao, et al.
Pubblicazione: (2026)
Human Pose Descriptions and Subject-Focused Attention for Improved Zero-Shot Transfer in Human-Centric Classification Tasks
di: Khan, Muhammad Saif Ullah, et al.
Pubblicazione: (2024)
di: Khan, Muhammad Saif Ullah, et al.
Pubblicazione: (2024)
Test-Time Visual In-Context Tuning
di: Xie, Jiahao, et al.
Pubblicazione: (2025)
di: Xie, Jiahao, et al.
Pubblicazione: (2025)
How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
di: Khattak, Muhammad Uzair, et al.
Pubblicazione: (2024)
di: Khattak, Muhammad Uzair, et al.
Pubblicazione: (2024)
ViT-5: Vision Transformers for The Mid-2020s
di: Wang, Feng, et al.
Pubblicazione: (2026)
di: Wang, Feng, et al.
Pubblicazione: (2026)
ViT-DD: Multi-Task Vision Transformer for Semi-Supervised Driver Distraction Detection
di: Ma, Yunsheng, et al.
Pubblicazione: (2022)
di: Ma, Yunsheng, et al.
Pubblicazione: (2022)
X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization
di: Kukleva, Anna, et al.
Pubblicazione: (2024)
di: Kukleva, Anna, et al.
Pubblicazione: (2024)
AIM: Amending Inherent Interpretability via Self-Supervised Masking
di: Alshami, Eyad, et al.
Pubblicazione: (2025)
di: Alshami, Eyad, et al.
Pubblicazione: (2025)
ACC-ViT : Atrous Convolution's Comeback in Vision Transformers
di: Ibtehaz, Nabil, et al.
Pubblicazione: (2024)
di: Ibtehaz, Nabil, et al.
Pubblicazione: (2024)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
di: Zhu, Chen, et al.
Pubblicazione: (2025)
di: Zhu, Chen, et al.
Pubblicazione: (2025)
B-cos Alignment for Inherently Interpretable CNNs and Vision Transformers
di: Böhle, Moritz, et al.
Pubblicazione: (2023)
di: Böhle, Moritz, et al.
Pubblicazione: (2023)
Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
di: Shah, Arya, et al.
Pubblicazione: (2025)
di: Shah, Arya, et al.
Pubblicazione: (2025)
CAS-ViT: Convolutional Additive Self-attention Vision Transformers for Efficient Mobile Applications
di: Zhang, Tianfang, et al.
Pubblicazione: (2024)
di: Zhang, Tianfang, et al.
Pubblicazione: (2024)
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
di: Li, Yifan, et al.
Pubblicazione: (2026)
di: Li, Yifan, et al.
Pubblicazione: (2026)
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
di: Ma, Xiaochen, et al.
Pubblicazione: (2023)
di: Ma, Xiaochen, et al.
Pubblicazione: (2023)
NT-ViT: Neural Transcoding Vision Transformers for EEG-to-fMRI Synthesis
di: Lanzino, Romeo, et al.
Pubblicazione: (2024)
di: Lanzino, Romeo, et al.
Pubblicazione: (2024)
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs
di: Yao, Ting, et al.
Pubblicazione: (2024)
di: Yao, Ting, et al.
Pubblicazione: (2024)
ViT-Explainer: An Interactive Walkthrough of the Vision Transformer Pipeline
di: Hernandez, Juan Manuel, et al.
Pubblicazione: (2026)
di: Hernandez, Juan Manuel, et al.
Pubblicazione: (2026)
ViT-FIQA: Assessing Face Image Quality using Vision Transformers
di: Atzori, Andrea, et al.
Pubblicazione: (2025)
di: Atzori, Andrea, et al.
Pubblicazione: (2025)
MPTQ-ViT: Mixed-Precision Post-Training Quantization for Vision Transformer
di: Tai, Yu-Shan, et al.
Pubblicazione: (2024)
di: Tai, Yu-Shan, et al.
Pubblicazione: (2024)
Unsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey
di: Siméoni, Oriane, et al.
Pubblicazione: (2023)
di: Siméoni, Oriane, et al.
Pubblicazione: (2023)
ViT-2SPN: Vision Transformer-based Dual-Stream Self-Supervised Pretraining Networks for Retinal OCT Classification
di: Saraei, Mohammadreza, et al.
Pubblicazione: (2025)
di: Saraei, Mohammadreza, et al.
Pubblicazione: (2025)
TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters
di: Wang, Haiyang, et al.
Pubblicazione: (2024)
di: Wang, Haiyang, et al.
Pubblicazione: (2024)
ADFQ-ViT: Activation-Distribution-Friendly Post-Training Quantization for Vision Transformers
di: Jiang, Yanfeng, et al.
Pubblicazione: (2024)
di: Jiang, Yanfeng, et al.
Pubblicazione: (2024)
Hyb-KAN ViT: Hybrid Kolmogorov-Arnold Networks Augmented Vision Transformer
di: Dey, Sainath, et al.
Pubblicazione: (2025)
di: Dey, Sainath, et al.
Pubblicazione: (2025)
SVD-ViT: Does SVD Make Vision Transformers Attend More to the Foreground?
di: Murata, Haruhiko, et al.
Pubblicazione: (2026)
di: Murata, Haruhiko, et al.
Pubblicazione: (2026)
ViT-1.58b: Mobile Vision Transformers in the 1-bit Era
di: Yuan, Zhengqing, et al.
Pubblicazione: (2024)
di: Yuan, Zhengqing, et al.
Pubblicazione: (2024)
Documenti analoghi
-
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
di: Kuzucu, Selim, et al.
Pubblicazione: (2026) -
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
di: Kukleva, Anna, et al.
Pubblicazione: (2025) -
Toward a Diffusion-Based Generalist for Dense Vision Tasks
di: Fan, Yue, et al.
Pubblicazione: (2024) -
Learning to Prompt with Text Only Supervision for Vision-Language Models
di: Khattak, Muhammad Uzair, et al.
Pubblicazione: (2024) -
OrCo: Towards Better Generalization via Orthogonality and Contrast for Few-Shot Class-Incremental Learning
di: Ahmed, Noor, et al.
Pubblicazione: (2024)