Face Pyramid Vision Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Islam, Khawar, Zaheer, Muhammad Zaigham, Mahmood, Arif |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiffuseMix: Label-Preserving Data Augmentation with Diffusion Models
by: Islam, Khawar, et al.
Published: (2024)
by: Islam, Khawar, et al.
Published: (2024)
GenMix: Effective Data Augmentation with Generative Diffusion Model Image Editing
by: Islam, Khawar, et al.
Published: (2024)
by: Islam, Khawar, et al.
Published: (2024)
Clustering Aided Weakly Supervised Training to Detect Anomalous Events in Surveillance Videos
by: Zaheer, Muhammad Zaigham, et al.
Published: (2022)
by: Zaheer, Muhammad Zaigham, et al.
Published: (2022)
Stabilizing Adversarially Learned One-Class Novelty Detection Using Pseudo Anomalies
by: Zaheer, Muhammad Zaigham, et al.
Published: (2022)
by: Zaheer, Muhammad Zaigham, et al.
Published: (2022)
Constricting Normal Latent Space for Anomaly Detection with Normal-only Training Data
by: Astrid, Marcella, et al.
Published: (2024)
by: Astrid, Marcella, et al.
Published: (2024)
Face-Voice Association with Inductive Bias for Maximum Class Separation
by: Moscati, Marta, et al.
Published: (2026)
by: Moscati, Marta, et al.
Published: (2026)
Deep Learning for Video-based Person Re-Identification: A Survey
by: Islam, Khawar
Published: (2023)
by: Islam, Khawar
Published: (2023)
Linking Faces and Voices Across Languages: Insights from the FAME 2026 Challenge
by: Moscati, Marta, et al.
Published: (2025)
by: Moscati, Marta, et al.
Published: (2025)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
by: Demidov, Dmitry, et al.
Published: (2025)
by: Demidov, Dmitry, et al.
Published: (2025)
Face-voice Association in Multilingual Environments (FAME) 2026 Challenge Evaluation Plan
by: Moscati, Marta, et al.
Published: (2025)
by: Moscati, Marta, et al.
Published: (2025)
Context-guided Responsible Data Augmentation with Diffusion Models
by: Islam, Khawar, et al.
Published: (2025)
by: Islam, Khawar, et al.
Published: (2025)
Collaborative Learning of Anomalies with Privacy (CLAP) for Unsupervised Video Anomaly Detection: A New Baseline
by: Al-lahham, Anas, et al.
Published: (2024)
by: Al-lahham, Anas, et al.
Published: (2024)
Exploiting Autoencoder's Weakness to Generate Pseudo Anomalies
by: Astrid, Marcella, et al.
Published: (2024)
by: Astrid, Marcella, et al.
Published: (2024)
Enhancing 3D Human Pose Estimation Amidst Severe Occlusion with Dual Transformer Fusion
by: Ghafoor, Mehwish, et al.
Published: (2024)
by: Ghafoor, Mehwish, et al.
Published: (2024)
Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs
by: Demidov, Dmitry, et al.
Published: (2025)
by: Demidov, Dmitry, et al.
Published: (2025)
Pyramid Hierarchical Transformer for Hyperspectral Image Classification
by: Ahmad, Muhammad, et al.
Published: (2024)
by: Ahmad, Muhammad, et al.
Published: (2024)
SB-BEVFusion: Enhancing the Robustness against Sensor Malfunction and Corruptions
by: Essl, Markus, et al.
Published: (2026)
by: Essl, Markus, et al.
Published: (2026)
AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
by: Nawaz, Umair, et al.
Published: (2025)
by: Nawaz, Umair, et al.
Published: (2025)
Face-voice Association in Multilingual Environments (FAME) Challenge 2024 Evaluation Plan
by: Saeed, Muhammad Saad, et al.
Published: (2024)
by: Saeed, Muhammad Saad, et al.
Published: (2024)
How Good is my Histopathology Vision-Language Foundation Model? A Holistic Benchmark
by: Majzoub, Roba Al, et al.
Published: (2025)
by: Majzoub, Roba Al, et al.
Published: (2025)
Chameleon: Images Are What You Need For Multimodal Learning Robust To Missing Modalities
by: Liaqat, Muhammad Irzam, et al.
Published: (2024)
by: Liaqat, Muhammad Irzam, et al.
Published: (2024)
Modality Invariant Multimodal Learning to Handle Missing Modalities: A Single-Branch Approach
by: Saeed, Muhammad Saad, et al.
Published: (2024)
by: Saeed, Muhammad Saad, et al.
Published: (2024)
EdgeDAM: Real-time Object Tracking for Mobile Devices
by: Raza, Syed Muhammad, et al.
Published: (2026)
by: Raza, Syed Muhammad, et al.
Published: (2026)
Pose-Guided Self-Training with Two-Stage Clustering for Unsupervised Landmark Discovery
by: Tourani, Siddharth, et al.
Published: (2024)
by: Tourani, Siddharth, et al.
Published: (2024)
NT-VOT211: A Large-Scale Benchmark for Night-time Visual Object Tracking
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
Polyp-PVT: Polyp Segmentation with Pyramid Vision Transformers
by: Dong, Bo, et al.
Published: (2021)
by: Dong, Bo, et al.
Published: (2021)
A Survey of the Self Supervised Learning Mechanisms for Vision Transformers
by: Khan, Asifullah, et al.
Published: (2024)
by: Khan, Asifullah, et al.
Published: (2024)
HRPVT: High-Resolution Pyramid Vision Transformer for medium and small-scale human pose estimation
by: Xu, Zhoujie
Published: (2024)
by: Xu, Zhoujie
Published: (2024)
OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents
by: Shabbir, Akashah, et al.
Published: (2026)
by: Shabbir, Akashah, et al.
Published: (2026)
Depth Attention for Robust RGB Tracking
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
GVTNet: Graph Vision Transformer For Face Super-Resolution
by: Yang, Chao, et al.
Published: (2025)
by: Yang, Chao, et al.
Published: (2025)
HSFM: Hard-Set-Guided Feature-Space Meta-Learning for Robust Classification under Spurious Correlations
by: Parast, Aryan Yazdan, et al.
Published: (2026)
by: Parast, Aryan Yazdan, et al.
Published: (2026)
Leveraging Intermediate Features of Vision Transformer for Face Anti-Spoofing
by: Feng, Mika, et al.
Published: (2025)
by: Feng, Mika, et al.
Published: (2025)
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
by: Xing, Long, et al.
Published: (2024)
by: Xing, Long, et al.
Published: (2024)
FaceLiVT: Face Recognition using Linear Vision Transformer with Structural Reparameterization For Mobile Device
by: Setyawan, Novendra, et al.
Published: (2025)
by: Setyawan, Novendra, et al.
Published: (2025)
Low-Rank Continual Pyramid Vision Transformer: Incrementally Segment Whole-Body Organs in CT with Light-Weighted Adaptation
by: Zhu, Vince, et al.
Published: (2024)
by: Zhu, Vince, et al.
Published: (2024)
Multi-view Pyramid Transformer: Look Coarser to See Broader
by: Kang, Gyeongjin, et al.
Published: (2025)
by: Kang, Gyeongjin, et al.
Published: (2025)
S-Adapter: Generalizing Vision Transformer for Face Anti-Spoofing with Statistical Tokens
by: Cai, Rizhao, et al.
Published: (2023)
by: Cai, Rizhao, et al.
Published: (2023)
ViT-FIQA: Assessing Face Image Quality using Vision Transformers
by: Atzori, Andrea, et al.
Published: (2025)
by: Atzori, Andrea, et al.
Published: (2025)
Enhancing Learnable Descriptive Convolutional Vision Transformer for Face Anti-Spoofing
by: Huanga, Pei-Kai, et al.
Published: (2025)
by: Huanga, Pei-Kai, et al.
Published: (2025)
Similar Items
-
DiffuseMix: Label-Preserving Data Augmentation with Diffusion Models
by: Islam, Khawar, et al.
Published: (2024) -
GenMix: Effective Data Augmentation with Generative Diffusion Model Image Editing
by: Islam, Khawar, et al.
Published: (2024) -
Clustering Aided Weakly Supervised Training to Detect Anomalous Events in Surveillance Videos
by: Zaheer, Muhammad Zaigham, et al.
Published: (2022) -
Stabilizing Adversarially Learned One-Class Novelty Detection Using Pseudo Anomalies
by: Zaheer, Muhammad Zaigham, et al.
Published: (2022) -
Constricting Normal Latent Space for Anomaly Detection with Normal-only Training Data
by: Astrid, Marcella, et al.
Published: (2024)