GiT: Towards Generalist Vision Transformer through Universal Language Interface
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Haiyang, Tang, Hao, Jiang, Li, Shi, Shaoshuai, Naeem, Muhammad Ferjad, Li, Hongsheng, Schiele, Bernt, Wang, Liwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Toward a Diffusion-Based Generalist for Dense Vision Tasks
von: Fan, Yue, et al.
Veröffentlicht: (2024)
von: Fan, Yue, et al.
Veröffentlicht: (2024)
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
von: Kuzucu, Selim, et al.
Veröffentlicht: (2025)
von: Kuzucu, Selim, et al.
Veröffentlicht: (2025)
TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters
von: Wang, Haiyang, et al.
Veröffentlicht: (2024)
von: Wang, Haiyang, et al.
Veröffentlicht: (2024)
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
von: Kuzucu, Selim, et al.
Veröffentlicht: (2026)
von: Kuzucu, Selim, et al.
Veröffentlicht: (2026)
MTR++: Multi-Agent Motion Prediction with Symmetric Scene Modeling and Guided Intention Querying
von: Shi, Shaoshuai, et al.
Veröffentlicht: (2023)
von: Shi, Shaoshuai, et al.
Veröffentlicht: (2023)
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
von: Kukleva, Anna, et al.
Veröffentlicht: (2025)
von: Kukleva, Anna, et al.
Veröffentlicht: (2025)
Learning to Prompt with Text Only Supervision for Vision-Language Models
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
B-cos Alignment for Inherently Interpretable CNNs and Vision Transformers
von: Böhle, Moritz, et al.
Veröffentlicht: (2023)
von: Böhle, Moritz, et al.
Veröffentlicht: (2023)
Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
von: Żukowska, Nina, et al.
Veröffentlicht: (2026)
von: Żukowska, Nina, et al.
Veröffentlicht: (2026)
MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
von: Das, Anurag, et al.
Veröffentlicht: (2024)
von: Das, Anurag, et al.
Veröffentlicht: (2024)
Exposing the Copycat Problem of Imitation-based Planner: A Novel Closed-Loop Simulator, Causal Benchmark and Joint IL-RL Baseline
von: Zhou, Hui, et al.
Veröffentlicht: (2025)
von: Zhou, Hui, et al.
Veröffentlicht: (2025)
SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
von: Chen, Xuesong, et al.
Veröffentlicht: (2025)
von: Chen, Xuesong, et al.
Veröffentlicht: (2025)
Towards Better Understanding Attribution Methods
von: Rao, Sukrut, et al.
Veröffentlicht: (2022)
von: Rao, Sukrut, et al.
Veröffentlicht: (2022)
OmniFashion: Towards Generalist Fashion Intelligence via Multi-Task Vision-Language Learning
von: Yang, Zhengwei, et al.
Veröffentlicht: (2026)
von: Yang, Zhengwei, et al.
Veröffentlicht: (2026)
VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow
von: Gorgun, Ada, et al.
Veröffentlicht: (2025)
von: Gorgun, Ada, et al.
Veröffentlicht: (2025)
OrCo: Towards Better Generalization via Orthogonality and Contrast for Few-Shot Class-Incremental Learning
von: Ahmed, Noor, et al.
Veröffentlicht: (2024)
von: Ahmed, Noor, et al.
Veröffentlicht: (2024)
UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
CFM: Language-aligned Concept Foundation Model for Vision
von: Wittenmayer, Kai, et al.
Veröffentlicht: (2026)
von: Wittenmayer, Kai, et al.
Veröffentlicht: (2026)
GLID: Pre-training a Generalist Encoder-Decoder Vision Model
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable
von: Arya, Shreyash, et al.
Veröffentlicht: (2024)
von: Arya, Shreyash, et al.
Veröffentlicht: (2024)
Towards Learning a Generalist Model for Embodied Navigation
von: Zheng, Duo, et al.
Veröffentlicht: (2023)
von: Zheng, Duo, et al.
Veröffentlicht: (2023)
Convolutional Dynamic Alignment Networks for Interpretable Classifications
von: Böhle, Moritz, et al.
Veröffentlicht: (2021)
von: Böhle, Moritz, et al.
Veröffentlicht: (2021)
Better Understanding Differences in Attribution Methods via Systematic Evaluations
von: Rao, Sukrut, et al.
Veröffentlicht: (2023)
von: Rao, Sukrut, et al.
Veröffentlicht: (2023)
DWDN: Deep Wiener Deconvolution Network for Non-Blind Image Deblurring
von: Dong, Jiangxin, et al.
Veröffentlicht: (2021)
von: Dong, Jiangxin, et al.
Veröffentlicht: (2021)
Optimising for Interpretability: Convolutional Dynamic Alignment Networks
von: Böhle, Moritz, et al.
Veröffentlicht: (2021)
von: Böhle, Moritz, et al.
Veröffentlicht: (2021)
Adversarial Training against Location-Optimized Adversarial Patches
von: Rao, Sukrut, et al.
Veröffentlicht: (2020)
von: Rao, Sukrut, et al.
Veröffentlicht: (2020)
GiMeFive: Towards Interpretable Facial Emotion Classification
von: Wang, Jiawen, et al.
Veröffentlicht: (2024)
von: Wang, Jiawen, et al.
Veröffentlicht: (2024)
GiLT: Augmenting Transformer Language Models with Dependency Graphs
von: Huang, Tianyu, et al.
Veröffentlicht: (2026)
von: Huang, Tianyu, et al.
Veröffentlicht: (2026)
ColaVLA: Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning in Autonomous Driving
von: Peng, Qihang, et al.
Veröffentlicht: (2025)
von: Peng, Qihang, et al.
Veröffentlicht: (2025)
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
von: Shvetsova, Nina, et al.
Veröffentlicht: (2025)
von: Shvetsova, Nina, et al.
Veröffentlicht: (2025)
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
PersonaHOI: Effortlessly Improving Personalized Face with Human-Object Interaction Generation
von: Hu, Xinting, et al.
Veröffentlicht: (2025)
von: Hu, Xinting, et al.
Veröffentlicht: (2025)
GiVE: Guiding Visual Encoder to Perceive Overlooked Information
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
DF-LoGiT: Data-Free Logic-Gated Backdoor Attacks in Vision Transformers
von: Shen, Xiaozuo, et al.
Veröffentlicht: (2026)
von: Shen, Xiaozuo, et al.
Veröffentlicht: (2026)
Human Pose Descriptions and Subject-Focused Attention for Improved Zero-Shot Transfer in Human-Centric Classification Tasks
von: Khan, Muhammad Saif Ullah, et al.
Veröffentlicht: (2024)
von: Khan, Muhammad Saif Ullah, et al.
Veröffentlicht: (2024)
What Matters in Building Vision-Language-Action Models for Generalist Robots
von: Li, Xinghang, et al.
Veröffentlicht: (2024)
von: Li, Xinghang, et al.
Veröffentlicht: (2024)
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
von: Shvetsova, Nina, et al.
Veröffentlicht: (2023)
von: Shvetsova, Nina, et al.
Veröffentlicht: (2023)
SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models
von: Xie, Jiahao, et al.
Veröffentlicht: (2026)
von: Xie, Jiahao, et al.
Veröffentlicht: (2026)
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
von: Hou, Zhi, et al.
Veröffentlicht: (2025)
von: Hou, Zhi, et al.
Veröffentlicht: (2025)
Towards Generalist Intelligence in Dentistry: Vision Foundation Models for Oral and Maxillofacial Radiology
von: Huang, Xinrui, et al.
Veröffentlicht: (2025)
von: Huang, Xinrui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Toward a Diffusion-Based Generalist for Dense Vision Tasks
von: Fan, Yue, et al.
Veröffentlicht: (2024) -
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
von: Kuzucu, Selim, et al.
Veröffentlicht: (2025) -
TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters
von: Wang, Haiyang, et al.
Veröffentlicht: (2024) -
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
von: Kuzucu, Selim, et al.
Veröffentlicht: (2026) -
MTR++: Multi-Agent Motion Prediction with Symmetric Scene Modeling and Guided Intention Querying
von: Shi, Shaoshuai, et al.
Veröffentlicht: (2023)