Saved in:
| Main Authors: | Zhu, Duowang, Huang, Xiaohu, Huang, Haiyan, Shao, Zhenfeng, Cheng, Qimin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2406.12847 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Change3D: Revisiting Change Detection and Captioning from A Video Modeling Perspective
by: Zhu, Duowang, et al.
Published: (2025)
by: Zhu, Duowang, et al.
Published: (2025)
ViTOC: Vision Transformer and Object-aware Captioner
by: Huang, Feiyang
Published: (2024)
by: Huang, Feiyang
Published: (2024)
ViTAR: Vision Transformer with Any Resolution
by: Fan, Qihang, et al.
Published: (2024)
by: Fan, Qihang, et al.
Published: (2024)
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Unleashing Vision-Language Semantics for Deepfake Video Detection
by: Zhu, Jiawen, et al.
Published: (2026)
by: Zhu, Jiawen, et al.
Published: (2026)
VPNeXt -- Rethinking Dense Decoding for Plain Vision Transformer
by: Tang, Xikai, et al.
Published: (2025)
by: Tang, Xikai, et al.
Published: (2025)
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
by: Wang, Shijing, et al.
Published: (2025)
by: Wang, Shijing, et al.
Published: (2025)
IDET: Iterative Difference-Enhanced Transformers for High-Quality Change Detection
by: Guo, Qing, et al.
Published: (2022)
by: Guo, Qing, et al.
Published: (2022)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
by: Zhu, Chen, et al.
Published: (2025)
by: Zhu, Chen, et al.
Published: (2025)
ViT-5: Vision Transformers for The Mid-2020s
by: Wang, Feng, et al.
Published: (2026)
by: Wang, Feng, et al.
Published: (2026)
LocalViT: Analyzing Locality in Vision Transformers
by: Li, Yawei, et al.
Published: (2021)
by: Li, Yawei, et al.
Published: (2021)
FTerViT: Fully Ternary Vision Transformer
by: Ruciński, Szymon, et al.
Published: (2026)
by: Ruciński, Szymon, et al.
Published: (2026)
ViT$^3$: Unlocking Test-Time Training in Vision
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
ViLaCD-R1: A Vision-Language Framework for Semantic Change Detection in Remote Sensing
by: Ma, Xingwei, et al.
Published: (2025)
by: Ma, Xingwei, et al.
Published: (2025)
Retina Vision Transformer (RetinaViT): Introducing Scaled Patches into Vision Transformers
by: Shu, Yuyang, et al.
Published: (2024)
by: Shu, Yuyang, et al.
Published: (2024)
Exploring Plain ViT Reconstruction for Multi-class Unsupervised Anomaly Detection
by: Zhang, Jiangning, et al.
Published: (2023)
by: Zhang, Jiangning, et al.
Published: (2023)
APHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision Transformers
by: Wu, Zhuguanyu, et al.
Published: (2025)
by: Wu, Zhuguanyu, et al.
Published: (2025)
ViT-DD: Multi-Task Vision Transformer for Semi-Supervised Driver Distraction Detection
by: Ma, Yunsheng, et al.
Published: (2022)
by: Ma, Yunsheng, et al.
Published: (2022)
GenConViT: Deepfake Video Detection Using Generative Convolutional Vision Transformer
by: Deressa, Deressa Wodajo, et al.
Published: (2023)
by: Deressa, Deressa Wodajo, et al.
Published: (2023)
ACC-ViT : Atrous Convolution's Comeback in Vision Transformers
by: Ibtehaz, Nabil, et al.
Published: (2024)
by: Ibtehaz, Nabil, et al.
Published: (2024)
ViTCN: Vision Transformer Contrastive Network For Reasoning
by: Song, Bo, et al.
Published: (2024)
by: Song, Bo, et al.
Published: (2024)
Trio-ViT: Post-Training Quantization and Acceleration for Softmax-Free Efficient Vision Transformer
by: Shi, Huihong, et al.
Published: (2024)
by: Shi, Huihong, et al.
Published: (2024)
HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
by: Tao, Chenxin, et al.
Published: (2024)
by: Tao, Chenxin, et al.
Published: (2024)
CEBSNet: Change-Excited and Background-Suppressed Network with Temporal Dependency Modeling for Bitemporal Change Detection
by: Xu, Qi'ao, et al.
Published: (2025)
by: Xu, Qi'ao, et al.
Published: (2025)
Scene Change Detection with Vision-Language Representation Learning
by: Sheng, Diwei, et al.
Published: (2026)
by: Sheng, Diwei, et al.
Published: (2026)
DenSe-AdViT: A novel Vision Transformer for Dense SAR Object Detection
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
LAMM-ViT: AI Face Detection via Layer-Aware Modulation of Region-Guided Attention
by: Zhang, Jiangling, et al.
Published: (2025)
by: Zhang, Jiangling, et al.
Published: (2025)
ViTGaze: Gaze Following with Interaction Features in Vision Transformers
by: Song, Yuehao, et al.
Published: (2024)
by: Song, Yuehao, et al.
Published: (2024)
ViTALS: Vision Transformer for Action Localization in Surgical Nephrectomy
by: Chandra, Soumyadeep, et al.
Published: (2024)
by: Chandra, Soumyadeep, et al.
Published: (2024)
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
UniViTAR: Unified Vision Transformer with Native Resolution
by: Qiao, Limeng, et al.
Published: (2025)
by: Qiao, Limeng, et al.
Published: (2025)
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
by: Ma, Xiaochen, et al.
Published: (2023)
by: Ma, Xiaochen, et al.
Published: (2023)
ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inference
by: Hojjat, Ali, et al.
Published: (2025)
by: Hojjat, Ali, et al.
Published: (2025)
SAC-ViT: Semantic-Aware Clustering Vision Transformer with Early Exit
by: Hu, Youbing, et al.
Published: (2025)
by: Hu, Youbing, et al.
Published: (2025)
Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding
by: Cheng, Zhiheng, et al.
Published: (2024)
by: Cheng, Zhiheng, et al.
Published: (2024)
LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
by: Yue, Tongtian, et al.
Published: (2025)
by: Yue, Tongtian, et al.
Published: (2025)
Advanced Feature Manipulation for Enhanced Change Detection Leveraging Natural Language Models
by: Li, Zhenglin, et al.
Published: (2024)
by: Li, Zhenglin, et al.
Published: (2024)
ViTGuard: Attention-aware Detection against Adversarial Examples for Vision Transformer
by: Sun, Shihua, et al.
Published: (2024)
by: Sun, Shihua, et al.
Published: (2024)
ChangeDINO: DINOv3-Driven Building Change Detection in Optical Remote Sensing Imagery
by: Cheng, Ching-Heng, et al.
Published: (2025)
by: Cheng, Ching-Heng, et al.
Published: (2025)
Beyond Segmentation: An Oil Spill Change Detection Framework Using Synthetic SAR Imagery
by: Lai, Chenyang, et al.
Published: (2026)
by: Lai, Chenyang, et al.
Published: (2026)
Similar Items
-
Change3D: Revisiting Change Detection and Captioning from A Video Modeling Perspective
by: Zhu, Duowang, et al.
Published: (2025) -
ViTOC: Vision Transformer and Object-aware Captioner
by: Huang, Feiyang
Published: (2024) -
ViTAR: Vision Transformer with Any Resolution
by: Fan, Qihang, et al.
Published: (2024) -
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
by: Li, Yifan, et al.
Published: (2025) -
Unleashing Vision-Language Semantics for Deepfake Video Detection
by: Zhu, Jiawen, et al.
Published: (2026)