Vision Augmentation Prediction Autoencoder with Attention Design (VAPAAD)
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Yin, Yiqiao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Causal Interpretation of Sparse Autoencoder Features in Vision
von: Han, Sangyu, et al.
Veröffentlicht: (2025)
von: Han, Sangyu, et al.
Veröffentlicht: (2025)
EA: An Event Autoencoder for High-Speed Vision Sensing
von: Islam, Riadul, et al.
Veröffentlicht: (2025)
von: Islam, Riadul, et al.
Veröffentlicht: (2025)
SAUCE: Selective Concept Unlearning in Vision-Language Models with Sparse Autoencoders
von: Li, Qing, et al.
Veröffentlicht: (2025)
von: Li, Qing, et al.
Veröffentlicht: (2025)
The Linear Attention Resurrection in Vision Transformer
von: Zheng, Chuanyang
Veröffentlicht: (2025)
von: Zheng, Chuanyang
Veröffentlicht: (2025)
Improved Belief-Attention in Vision Task
von: Zhang, Guoqiang
Veröffentlicht: (2026)
von: Zhang, Guoqiang
Veröffentlicht: (2026)
Spiking Vision Transformer with Saccadic Attention
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models
von: Rao, Zhifeng, et al.
Veröffentlicht: (2026)
von: Rao, Zhifeng, et al.
Veröffentlicht: (2026)
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
von: Tao, Hongyuan, et al.
Veröffentlicht: (2025)
von: Tao, Hongyuan, et al.
Veröffentlicht: (2025)
Attention Guided CAM: Visual Explanations of Vision Transformer Guided by Self-Attention
von: Leem, Saebom, et al.
Veröffentlicht: (2024)
von: Leem, Saebom, et al.
Veröffentlicht: (2024)
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
von: Böhle, Moritz, et al.
Veröffentlicht: (2025)
von: Böhle, Moritz, et al.
Veröffentlicht: (2025)
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
von: Zou, Dongyun, et al.
Veröffentlicht: (2026)
von: Zou, Dongyun, et al.
Veröffentlicht: (2026)
Revisiting the Integration of Convolution and Attention for Vision Backbone
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
Attention Retention for Continual Learning with Vision Transformers
von: Lu, Yue, et al.
Veröffentlicht: (2026)
von: Lu, Yue, et al.
Veröffentlicht: (2026)
DGSAN: Dual-Graph Spatiotemporal Attention Network for Pulmonary Nodule Malignancy Prediction
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
Attention Prompting on Image for Large Vision-Language Models
von: Yu, Runpeng, et al.
Veröffentlicht: (2024)
von: Yu, Runpeng, et al.
Veröffentlicht: (2024)
Large Vision-Language Models Get Lost in Attention
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
Rethinking Causal Mask Attention for Vision-Language Inference
von: Pei, Xiaohuan, et al.
Veröffentlicht: (2025)
von: Pei, Xiaohuan, et al.
Veröffentlicht: (2025)
VISTA: A Vision and Intent-Aware Social Attention Framework for Multi-Agent Trajectory Prediction
von: Martins, Stephane Da Silva, et al.
Veröffentlicht: (2025)
von: Martins, Stephane Da Silva, et al.
Veröffentlicht: (2025)
Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models
von: Apedo, Yvon, et al.
Veröffentlicht: (2026)
von: Apedo, Yvon, et al.
Veröffentlicht: (2026)
An Exploratory Study on Human-Centric Video Anomaly Detection through Variational Autoencoders and Trajectory Prediction
von: Noghre, Ghazal Alinezhad, et al.
Veröffentlicht: (2024)
von: Noghre, Ghazal Alinezhad, et al.
Veröffentlicht: (2024)
Gaussian Masked Autoencoders
von: Rajasegaran, Jathushan, et al.
Veröffentlicht: (2025)
von: Rajasegaran, Jathushan, et al.
Veröffentlicht: (2025)
Tighnari: Multi-modal Plant Species Prediction Based on Hierarchical Cross-Attention Using Graph-Based and Vision Backbone-Extracted Features
von: Liu, Haixu, et al.
Veröffentlicht: (2025)
von: Liu, Haixu, et al.
Veröffentlicht: (2025)
A-VL: Adaptive Attention for Large Vision-Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
SPARO: Selective Attention for Robust and Compositional Transformer Encodings for Vision
von: Vani, Ankit, et al.
Veröffentlicht: (2024)
von: Vani, Ankit, et al.
Veröffentlicht: (2024)
Leveraging Vision-Language Models to Detect Attention in Educational Videos
von: Becquet, Gabriel, et al.
Veröffentlicht: (2026)
von: Becquet, Gabriel, et al.
Veröffentlicht: (2026)
Learning to Look: Cognitive Attention Alignment with Vision-Language Models
von: Yang, Ryan L., et al.
Veröffentlicht: (2025)
von: Yang, Ryan L., et al.
Veröffentlicht: (2025)
PolaFormer: Polarity-aware Linear Attention for Vision Transformers
von: Meng, Weikang, et al.
Veröffentlicht: (2025)
von: Meng, Weikang, et al.
Veröffentlicht: (2025)
Computer Vision and Deep Learning for 4D Augmented Reality
von: Shivashankar, Karthik
Veröffentlicht: (2025)
von: Shivashankar, Karthik
Veröffentlicht: (2025)
Frequency-Dynamic Attention Modulation for Dense Prediction
von: Chen, Linwei, et al.
Veröffentlicht: (2025)
von: Chen, Linwei, et al.
Veröffentlicht: (2025)
Delta-K: Boosting Multi-Instance Generation via Cross-Attention Augmentation
von: Wang, Zitong, et al.
Veröffentlicht: (2026)
von: Wang, Zitong, et al.
Veröffentlicht: (2026)
SCE-MAE: Selective Correspondence Enhancement with Masked Autoencoder for Self-Supervised Landmark Estimation
von: Yin, Kejia, et al.
Veröffentlicht: (2024)
von: Yin, Kejia, et al.
Veröffentlicht: (2024)
FasterViT: Fast Vision Transformers with Hierarchical Attention
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2023)
Vision Eagle Attention: a new lens for advancing image classification
von: Hasan, Mahmudul
Veröffentlicht: (2024)
von: Hasan, Mahmudul
Veröffentlicht: (2024)
Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models
von: Wang, Zhiqiang, et al.
Veröffentlicht: (2026)
von: Wang, Zhiqiang, et al.
Veröffentlicht: (2026)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
von: Liu, Hanqing, et al.
Veröffentlicht: (2026)
von: Liu, Hanqing, et al.
Veröffentlicht: (2026)
\textsc{NaVIDA}: Vision-Language Navigation with Inverse Dynamics Augmentation
von: Zhu, Weiye, et al.
Veröffentlicht: (2026)
von: Zhu, Weiye, et al.
Veröffentlicht: (2026)
Genetic Learning for Designing Sim-to-Real Data Augmentations
von: Vanherle, Bram, et al.
Veröffentlicht: (2024)
von: Vanherle, Bram, et al.
Veröffentlicht: (2024)
NutriScreener: Retrieval-Augmented Multi-Pose Graph Attention Network for Malnourishment Screening
von: Khan, Misaal, et al.
Veröffentlicht: (2025)
von: Khan, Misaal, et al.
Veröffentlicht: (2025)
Towards Robust Unsupervised Attention Prediction in Autonomous Driving
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Causal Interpretation of Sparse Autoencoder Features in Vision
von: Han, Sangyu, et al.
Veröffentlicht: (2025) -
EA: An Event Autoencoder for High-Speed Vision Sensing
von: Islam, Riadul, et al.
Veröffentlicht: (2025) -
SAUCE: Selective Concept Unlearning in Vision-Language Models with Sparse Autoencoders
von: Li, Qing, et al.
Veröffentlicht: (2025) -
The Linear Attention Resurrection in Vision Transformer
von: Zheng, Chuanyang
Veröffentlicht: (2025) -
Improved Belief-Attention in Vision Task
von: Zhang, Guoqiang
Veröffentlicht: (2026)