DuoFormer: Leveraging Hierarchical Visual Representations by Local and Global Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Xiaoya, Zhang, Bodong, Knudsen, Beatrice S., Tasdizen, Tolga |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer
by: Tang, Xiaoya, et al.
Published: (2025)
by: Tang, Xiaoya, et al.
Published: (2025)
SRA: A Novel Method to Improve Feature Embedding in Self-supervised Learning for Histopathological Images
by: Manoochehri, Hamid, et al.
Published: (2024)
by: Manoochehri, Hamid, et al.
Published: (2024)
Weakly Supervised Contrastive Learning for Histopathology Patch Embeddings
by: Zhang, Bodong, et al.
Published: (2026)
by: Zhang, Bodong, et al.
Published: (2026)
WeakSupCon: Weakly Supervised Contrastive Learning for Encoder Pre-training
by: Zhang, Bodong, et al.
Published: (2025)
by: Zhang, Bodong, et al.
Published: (2025)
HAVT-IVD: Heterogeneity-Aware Cross-Modal Network for Audio-Visual Surveillance: Idling Vehicles Detection With Multichannel Audio and Multiscale Visual Cues
by: Li, Xiwen, et al.
Published: (2025)
by: Li, Xiwen, et al.
Published: (2025)
VISTA: A Visual and Textual Attention Dataset for Interpreting Multimodal Models
by: Harshit, et al.
Published: (2024)
by: Harshit, et al.
Published: (2024)
CLASS-M: Adaptive stain separation-based contrastive learning with pseudo-labeling for histopathological image classification
by: Zhang, Bodong, et al.
Published: (2023)
by: Zhang, Bodong, et al.
Published: (2023)
HGTS-Former: Hierarchical HyperGraph Transformer for Multivariate Time Series Analysis
by: Si, Hao, et al.
Published: (2025)
by: Si, Hao, et al.
Published: (2025)
RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
by: Huang, Wen, et al.
Published: (2025)
by: Huang, Wen, et al.
Published: (2025)
A Comparison of Object Detection and Phrase Grounding Models in Chest X-ray Abnormality Localization using Eye-tracking Data
by: Ghelichkhan, Elham, et al.
Published: (2025)
by: Ghelichkhan, Elham, et al.
Published: (2025)
How to Build Robust, Scalable Models for GSV-Based Indicators in Neighborhood Research
by: Tang, Xiaoya, et al.
Published: (2026)
by: Tang, Xiaoya, et al.
Published: (2026)
HCVP: Leveraging Hierarchical Contrastive Visual Prompt for Domain Generalization
by: Zhou, Guanglin, et al.
Published: (2024)
by: Zhou, Guanglin, et al.
Published: (2024)
LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel
by: Feng, Zhe, et al.
Published: (2026)
by: Feng, Zhe, et al.
Published: (2026)
PolaFormer: Polarity-aware Linear Attention for Vision Transformers
by: Meng, Weikang, et al.
Published: (2025)
by: Meng, Weikang, et al.
Published: (2025)
Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion
by: Zhang, Junkai, et al.
Published: (2025)
by: Zhang, Junkai, et al.
Published: (2025)
F2FLDM: Latent Diffusion Models with Histopathology Pre-Trained Embeddings for Unpaired Frozen Section to FFPE Translation
by: Ho, Man M., et al.
Published: (2024)
by: Ho, Man M., et al.
Published: (2024)
H3Former: Hypergraph-based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification
by: Zhang, Yongji, et al.
Published: (2025)
by: Zhang, Yongji, et al.
Published: (2025)
Representation Learning of Point Cloud Upsampling in Global and Local Inputs
by: Zhang, Tongxu, et al.
Published: (2025)
by: Zhang, Tongxu, et al.
Published: (2025)
Leveraging Swin Transformer for Local-to-Global Weakly Supervised Semantic Segmentation
by: Ahmadi, Rozhan, et al.
Published: (2024)
by: Ahmadi, Rozhan, et al.
Published: (2024)
Re-Attentional Controllable Video Diffusion Editing
by: Wang, Yuanzhi, et al.
Published: (2024)
by: Wang, Yuanzhi, et al.
Published: (2024)
Local and Global Feature Attention Fusion Network for Face Recognition
by: Yu, Wang, et al.
Published: (2024)
by: Yu, Wang, et al.
Published: (2024)
SERNet-Former: Semantic Segmentation by Efficient Residual Network with Attention-Boosting Gates and Attention-Fusion Networks
by: Erisen, Serdar
Published: (2024)
by: Erisen, Serdar
Published: (2024)
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models
by: Dong, Xinpeng, et al.
Published: (2026)
by: Dong, Xinpeng, et al.
Published: (2026)
GRAD-Former: Gated Robust Attention-based Differential Transformer for Change Detection
by: Ameta, Durgesh, et al.
Published: (2026)
by: Ameta, Durgesh, et al.
Published: (2026)
Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models
by: He, Hulingxiao, et al.
Published: (2026)
by: He, Hulingxiao, et al.
Published: (2026)
Local-Global Attention: An Adaptive Mechanism for Multi-Scale Feature Integration
by: Shao, Yifan
Published: (2024)
by: Shao, Yifan
Published: (2024)
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
by: Jiang, Yubo, et al.
Published: (2026)
by: Jiang, Yubo, et al.
Published: (2026)
Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection
by: Zhang, Daichi, et al.
Published: (2025)
by: Zhang, Daichi, et al.
Published: (2025)
MetaCOG: A Hierarchical Probabilistic Model for Learning Meta-Cognitive Visual Representations
by: Berke, Marlene D., et al.
Published: (2021)
by: Berke, Marlene D., et al.
Published: (2021)
Unified Attention Modeling for Efficient Free-Viewing and Visual Search via Shared Representations
by: Mohammed, Fatma Youssef, et al.
Published: (2025)
by: Mohammed, Fatma Youssef, et al.
Published: (2025)
CarFormer: Self-Driving with Learned Object-Centric Representations
by: Hamdan, Shadi, et al.
Published: (2024)
by: Hamdan, Shadi, et al.
Published: (2024)
Learning and Leveraging World Models in Visual Representation Learning
by: Garrido, Quentin, et al.
Published: (2024)
by: Garrido, Quentin, et al.
Published: (2024)
Global-Local Medical SAM Adaptor Based on Full Adaption
by: Wang, Meng, et al.
Published: (2024)
by: Wang, Meng, et al.
Published: (2024)
Adversarial Attacks Leverage Interference Between Features in Superposition
by: Stevinson, Edward, et al.
Published: (2025)
by: Stevinson, Edward, et al.
Published: (2025)
Bridge the Gap: From Weak to Full Supervision for Temporal Action Localization with PseudoFormer
by: Liu, Ziyi, et al.
Published: (2025)
by: Liu, Ziyi, et al.
Published: (2025)
EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
by: Su, Qile, et al.
Published: (2025)
by: Su, Qile, et al.
Published: (2025)
MetaSeg: MetaFormer-based Global Contexts-aware Network for Efficient Semantic Segmentation
by: Kang, Beoungwoo, et al.
Published: (2024)
by: Kang, Beoungwoo, et al.
Published: (2024)
Adaptive Global-Local Representation Learning and Selection for Cross-Domain Facial Expression Recognition
by: Gao, Yuefang, et al.
Published: (2024)
by: Gao, Yuefang, et al.
Published: (2024)
A benchmark dataset for deep learning-based airplane detection: HRPlanes
by: Bakirman, Tolga, et al.
Published: (2022)
by: Bakirman, Tolga, et al.
Published: (2022)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
by: An, Wenbin, et al.
Published: (2024)
by: An, Wenbin, et al.
Published: (2024)
Similar Items
-
DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer
by: Tang, Xiaoya, et al.
Published: (2025) -
SRA: A Novel Method to Improve Feature Embedding in Self-supervised Learning for Histopathological Images
by: Manoochehri, Hamid, et al.
Published: (2024) -
Weakly Supervised Contrastive Learning for Histopathology Patch Embeddings
by: Zhang, Bodong, et al.
Published: (2026) -
WeakSupCon: Weakly Supervised Contrastive Learning for Encoder Pre-training
by: Zhang, Bodong, et al.
Published: (2025) -
HAVT-IVD: Heterogeneity-Aware Cross-Modal Network for Audio-Visual Surveillance: Idling Vehicles Detection With Multichannel Audio and Multiscale Visual Cues
by: Li, Xiwen, et al.
Published: (2025)