DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Xiaoya, Zhang, Bodong, Ho, Man Minh, Knudsen, Beatrice S., Tasdizen, Tolga |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DuoFormer: Leveraging Hierarchical Visual Representations by Local and Global Attention
by: Tang, Xiaoya, et al.
Published: (2024)
by: Tang, Xiaoya, et al.
Published: (2024)
SRA: A Novel Method to Improve Feature Embedding in Self-supervised Learning for Histopathological Images
by: Manoochehri, Hamid, et al.
Published: (2024)
by: Manoochehri, Hamid, et al.
Published: (2024)
Weakly Supervised Contrastive Learning for Histopathology Patch Embeddings
by: Zhang, Bodong, et al.
Published: (2026)
by: Zhang, Bodong, et al.
Published: (2026)
WeakSupCon: Weakly Supervised Contrastive Learning for Encoder Pre-training
by: Zhang, Bodong, et al.
Published: (2025)
by: Zhang, Bodong, et al.
Published: (2025)
CLASS-M: Adaptive stain separation-based contrastive learning with pseudo-labeling for histopathological image classification
by: Zhang, Bodong, et al.
Published: (2023)
by: Zhang, Bodong, et al.
Published: (2023)
HAVT-IVD: Heterogeneity-Aware Cross-Modal Network for Audio-Visual Surveillance: Idling Vehicles Detection With Multichannel Audio and Multiscale Visual Cues
by: Li, Xiwen, et al.
Published: (2025)
by: Li, Xiwen, et al.
Published: (2025)
VISTA: A Visual and Textual Attention Dataset for Interpreting Multimodal Models
by: Harshit, et al.
Published: (2024)
by: Harshit, et al.
Published: (2024)
F2FLDM: Latent Diffusion Models with Histopathology Pre-Trained Embeddings for Unpaired Frozen Section to FFPE Translation
by: Ho, Man M., et al.
Published: (2024)
by: Ho, Man M., et al.
Published: (2024)
DISC: Latent Diffusion Models with Self-Distillation from Separated Conditions for Prostate Cancer Grading
by: Ho, Man M., et al.
Published: (2024)
by: Ho, Man M., et al.
Published: (2024)
A Comparison of Object Detection and Phrase Grounding Models in Chest X-ray Abnormality Localization using Eye-tracking Data
by: Ghelichkhan, Elham, et al.
Published: (2025)
by: Ghelichkhan, Elham, et al.
Published: (2025)
How to Build Robust, Scalable Models for GSV-Based Indicators in Neighborhood Research
by: Tang, Xiaoya, et al.
Published: (2026)
by: Tang, Xiaoya, et al.
Published: (2026)
MedFormer: Hierarchical Medical Vision Transformer with Content-Aware Dual Sparse Selection Attention
by: Xia, Zunhui, et al.
Published: (2025)
by: Xia, Zunhui, et al.
Published: (2025)
AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation
by: Wu, Ruipu, et al.
Published: (2025)
by: Wu, Ruipu, et al.
Published: (2025)
Vision Transformers with Hierarchical Attention
by: Liu, Yun, et al.
Published: (2021)
by: Liu, Yun, et al.
Published: (2021)
HGTS-Former: Hierarchical HyperGraph Transformer for Multivariate Time Series Analysis
by: Si, Hao, et al.
Published: (2025)
by: Si, Hao, et al.
Published: (2025)
PolaFormer: Polarity-aware Linear Attention for Vision Transformers
by: Meng, Weikang, et al.
Published: (2025)
by: Meng, Weikang, et al.
Published: (2025)
AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer
by: Shan, Jiquan, et al.
Published: (2025)
by: Shan, Jiquan, et al.
Published: (2025)
DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention
by: Long, Nguyen Huu Bao, et al.
Published: (2024)
by: Long, Nguyen Huu Bao, et al.
Published: (2024)
DuoCast: Duo-Probabilistic Diffusion for Precipitation Nowcasting
by: Wen, Penghui, et al.
Published: (2024)
by: Wen, Penghui, et al.
Published: (2024)
ScatterFormer: Efficient Voxel Transformer with Scattered Linear Attention
by: He, Chenhang, et al.
Published: (2024)
by: He, Chenhang, et al.
Published: (2024)
PartFormer: Awakening Latent Diverse Representation from Vision Transformer for Object Re-Identification
by: Tan, Lei, et al.
Published: (2024)
by: Tan, Lei, et al.
Published: (2024)
Calibration Attention: Learning Reliability-Aware Representations for Vision Transformers
by: Liang, Wenhao, et al.
Published: (2025)
by: Liang, Wenhao, et al.
Published: (2025)
CarcassFormer: An End-to-end Transformer-based Framework for Simultaneous Localization, Segmentation and Classification of Poultry Carcass Defect
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
DuoSpaceNet: Leveraging Both Bird's-Eye-View and Perspective View Representations for 3D Object Detection
by: Huang, Zhe, et al.
Published: (2024)
by: Huang, Zhe, et al.
Published: (2024)
RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
by: Huang, Wen, et al.
Published: (2025)
by: Huang, Wen, et al.
Published: (2025)
Global2Local: A Joint-Hierarchical Attention for Video Captioning
by: Dai, Chengpeng, et al.
Published: (2022)
by: Dai, Chengpeng, et al.
Published: (2022)
LoFormer: Local Frequency Transformer for Image Deblurring
by: Mao, Xintian, et al.
Published: (2024)
by: Mao, Xintian, et al.
Published: (2024)
Hierarchical Transformer for Electrocardiogram Diagnosis
by: Tang, Xiaoya, et al.
Published: (2024)
by: Tang, Xiaoya, et al.
Published: (2024)
Analyzing Local Representations of Self-supervised Vision Transformers
by: Vanyan, Ani, et al.
Published: (2023)
by: Vanyan, Ani, et al.
Published: (2023)
RoadFormer : Local-Global Feature Fusion for Road Surface Classification in Autonomous Driving
by: Wang, Tianze, et al.
Published: (2025)
by: Wang, Tianze, et al.
Published: (2025)
Joint Audio-Visual Idling Vehicle Detection with Streamlined Input Dependencies
by: Li, Xiwen, et al.
Published: (2024)
by: Li, Xiwen, et al.
Published: (2024)
Distilled Transformers with Locally Enhanced Global Representations for Face Forgery Detection
by: Zhang, Yaning, et al.
Published: (2024)
by: Zhang, Yaning, et al.
Published: (2024)
HexFormer: Hyperbolic Vision Transformer with Exponential Map Aggregation
by: Alyoussef, Haya, et al.
Published: (2026)
by: Alyoussef, Haya, et al.
Published: (2026)
TreeFormers -- An Exploration of Vision Transformers for Deforestation Driver Classification
by: Ochuba, Uche
Published: (2024)
by: Ochuba, Uche
Published: (2024)
ParFormer: A Vision Transformer with Parallel Mixer and Sparse Channel Attention Patch Embedding
by: Setyawan, Novendra, et al.
Published: (2024)
by: Setyawan, Novendra, et al.
Published: (2024)
Unifying Global-Local Representations in Salient Object Detection with Transformer
by: Ren, Sucheng, et al.
Published: (2021)
by: Ren, Sucheng, et al.
Published: (2021)
ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance Segmentation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformers
by: Xiao, Chaodong, et al.
Published: (2026)
by: Xiao, Chaodong, et al.
Published: (2026)
UniForm: A Reuse Attention Mechanism Optimized for Efficient Vision Transformers on Edge Devices
by: Yeom, Seul-Ki, et al.
Published: (2024)
by: Yeom, Seul-Ki, et al.
Published: (2024)
Technical Report for ICRA 2025 GOOSE 3D Semantic Segmentation Challenge: Adaptive Point Cloud Understanding for Heterogeneous Robotic Systems
by: Zhang, Xiaoya
Published: (2025)
by: Zhang, Xiaoya
Published: (2025)
Similar Items
-
DuoFormer: Leveraging Hierarchical Visual Representations by Local and Global Attention
by: Tang, Xiaoya, et al.
Published: (2024) -
SRA: A Novel Method to Improve Feature Embedding in Self-supervised Learning for Histopathological Images
by: Manoochehri, Hamid, et al.
Published: (2024) -
Weakly Supervised Contrastive Learning for Histopathology Patch Embeddings
by: Zhang, Bodong, et al.
Published: (2026) -
WeakSupCon: Weakly Supervised Contrastive Learning for Encoder Pre-training
by: Zhang, Bodong, et al.
Published: (2025) -
CLASS-M: Adaptive stain separation-based contrastive learning with pseudo-labeling for histopathological image classification
by: Zhang, Bodong, et al.
Published: (2023)