Let ViT Speak: Generative Language-Image Pre-training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fang, Yan, Lan, Mengcheng, Huang, Zilong, Lei, Weixian, Zhao, Yunqing, Zhong, Yujie, Yu, Yingchen, She, Qi, Zhao, Yao, Wei, Yunchao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ThinkGen: Generalized Thinking for Visual Generation
von: Jiao, Siyu, et al.
Veröffentlicht: (2025)
von: Jiao, Siyu, et al.
Veröffentlicht: (2025)
Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
von: Lan, Mengcheng, et al.
Veröffentlicht: (2025)
von: Lan, Mengcheng, et al.
Veröffentlicht: (2025)
ViT-Lens: Towards Omni-modal Representations
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
von: Song, Qi, et al.
Veröffentlicht: (2025)
von: Song, Qi, et al.
Veröffentlicht: (2025)
ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
ViT Registers and Fractal ViT
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2026)
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2026)
Pneumonia Image Classification Based on Lightweight Mobile ViT Networks
von: Zhiqiang Zheng, et al.
Veröffentlicht: (2025)
von: Zhiqiang Zheng, et al.
Veröffentlicht: (2025)
How to train your ViT for OOD Detection
von: Mueller, Maximilian, et al.
Veröffentlicht: (2024)
von: Mueller, Maximilian, et al.
Veröffentlicht: (2024)
Refining Datapath for Microscaling ViTs
von: Xiao, Can, et al.
Veröffentlicht: (2025)
von: Xiao, Can, et al.
Veröffentlicht: (2025)
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
von: Zhong, Yunshan, et al.
Veröffentlicht: (2023)
von: Zhong, Yunshan, et al.
Veröffentlicht: (2023)
IPSeg: Image Posterior Mitigates Semantic Drift in Class-Incremental Segmentation
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
PreFM: Online Audio-Visual Event Parsing via Predictive Future Modeling
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
ODE-ViT: Plug & Play Attention Layer from the Generalization of the ViT as an Ordinary Differential Equation
von: Riera, Carlos Boned, et al.
Veröffentlicht: (2025)
von: Riera, Carlos Boned, et al.
Veröffentlicht: (2025)
UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register
von: Qiu, Congpei, et al.
Veröffentlicht: (2026)
von: Qiu, Congpei, et al.
Veröffentlicht: (2026)
MIPHEI-ViT: Multiplex Immunofluorescence Prediction from H&E Images using ViT Foundation Models
von: Balezo, Guillaume, et al.
Veröffentlicht: (2025)
von: Balezo, Guillaume, et al.
Veröffentlicht: (2025)
TextSculptor: Training and Benchmarking Scene Text Editing
von: Lin, Yiheng, et al.
Veröffentlicht: (2026)
von: Lin, Yiheng, et al.
Veröffentlicht: (2026)
Your ViT is Secretly an Image Segmentation Model
von: Kerssies, Tommie, et al.
Veröffentlicht: (2025)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2025)
Let Confucian Philosophy Speak on Its Own Terms
von: Jing Liu, et al.
Veröffentlicht: (2024)
von: Jing Liu, et al.
Veröffentlicht: (2024)
TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
von: Yang, Zuhao, et al.
Veröffentlicht: (2025)
von: Yang, Zuhao, et al.
Veröffentlicht: (2025)
MobilePlantViT: A Mobile-friendly Hybrid ViT for Generalized Plant Disease Image Classification
von: Tonmoy, Moshiur Rahman, et al.
Veröffentlicht: (2025)
von: Tonmoy, Moshiur Rahman, et al.
Veröffentlicht: (2025)
Deeper Inside Deep ViT
von: Hong, Sungrae
Veröffentlicht: (2025)
von: Hong, Sungrae
Veröffentlicht: (2025)
CLAMP-ViT: Contrastive Data-Free Learning for Adaptive Post-Training Quantization of ViTs
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2024)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2024)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
von: Chattopadhyay, Nandish, et al.
Veröffentlicht: (2026)
von: Chattopadhyay, Nandish, et al.
Veröffentlicht: (2026)
Let the Results Speak: A Replication-First Paradigm for LLM Behavioral Benchmarking
von: Yuming, et al.
Veröffentlicht: (2026)
von: Yuming, et al.
Veröffentlicht: (2026)
Case-Enhanced Vision Transformer: Improving Explanations of Image Similarity with a ViT-based Similarity Metric
von: Zhao, Ziwei, et al.
Veröffentlicht: (2024)
von: Zhao, Ziwei, et al.
Veröffentlicht: (2024)
ViTCAE: ViT-based Class-conditioned Autoencoder
von: Jebraeeli, Vahid, et al.
Veröffentlicht: (2025)
von: Jebraeeli, Vahid, et al.
Veröffentlicht: (2025)
HydraViT: Stacking Heads for a Scalable ViT
von: Haberer, Janek, et al.
Veröffentlicht: (2024)
von: Haberer, Janek, et al.
Veröffentlicht: (2024)
Applying ViT in Generalized Few-shot Semantic Segmentation
von: Geng, Liyuan, et al.
Veröffentlicht: (2024)
von: Geng, Liyuan, et al.
Veröffentlicht: (2024)
One-Shot Multilingual Font Generation Via ViT
von: Wang, Zhiheng, et al.
Veröffentlicht: (2024)
von: Wang, Zhiheng, et al.
Veröffentlicht: (2024)
DreamLCM: Towards High-Quality Text-to-3D Generation via Latent Consistency Model
von: Zhong, Yiming, et al.
Veröffentlicht: (2024)
von: Zhong, Yiming, et al.
Veröffentlicht: (2024)
ACC-ViT : Atrous Convolution's Comeback in Vision Transformers
von: Ibtehaz, Nabil, et al.
Veröffentlicht: (2024)
von: Ibtehaz, Nabil, et al.
Veröffentlicht: (2024)
Pretrained ViTs Yield Versatile Representations For Medical Images
von: Matsoukas, Christos, et al.
Veröffentlicht: (2023)
von: Matsoukas, Christos, et al.
Veröffentlicht: (2023)
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
von: Ma, Xiaochen, et al.
Veröffentlicht: (2023)
von: Ma, Xiaochen, et al.
Veröffentlicht: (2023)
Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentation
von: Tang, Fenghe, et al.
Veröffentlicht: (2025)
von: Tang, Fenghe, et al.
Veröffentlicht: (2025)
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs
von: Yao, Ting, et al.
Veröffentlicht: (2024)
von: Yao, Ting, et al.
Veröffentlicht: (2024)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
von: Zhu, Chen, et al.
Veröffentlicht: (2025)
von: Zhu, Chen, et al.
Veröffentlicht: (2025)
Lightweight ViT with Multiscale Feature Fusion for Driving Risk Rating Warning System
von: Hao Tang, et al.
Veröffentlicht: (2024)
von: Hao Tang, et al.
Veröffentlicht: (2024)
ERVD: An Efficient and Robust ViT-Based Distillation Framework for Remote Sensing Image Retrieval
von: Dong, Le, et al.
Veröffentlicht: (2024)
von: Dong, Le, et al.
Veröffentlicht: (2024)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024)
von: Zhao, Wangbo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ThinkGen: Generalized Thinking for Visual Generation
von: Jiao, Siyu, et al.
Veröffentlicht: (2025) -
Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
von: Lan, Mengcheng, et al.
Veröffentlicht: (2025) -
ViT-Lens: Towards Omni-modal Representations
von: Lei, Weixian, et al.
Veröffentlicht: (2023) -
CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
von: Song, Qi, et al.
Veröffentlicht: (2025) -
ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights
von: Lei, Weixian, et al.
Veröffentlicht: (2023)