Contrast: A Hybrid Architecture of Transformers and State Space Models for Low-Level Vision
Fuente:
arXiv
Saved in:
| Main Authors: | Urumbekov, Aman, Chen, Zheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Sensorimotor Vision Transformer
by: Gadzicki, Konrad, et al.
Published: (2025)
by: Gadzicki, Konrad, et al.
Published: (2025)
MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
by: Galella, Santiago, et al.
Published: (2026)
by: Galella, Santiago, et al.
Published: (2026)
LISTA-Transformer Model Based on Sparse Coding and Attention Mechanism and Its Application in Fault Diagnosis
by: Liu, Shuang, et al.
Published: (2026)
by: Liu, Shuang, et al.
Published: (2026)
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
by: Chaybouti, Sofian, et al.
Published: (2025)
by: Chaybouti, Sofian, et al.
Published: (2025)
ViTNF: Leveraging Neural Fields to Boost Vision Transformers in Generalized Category Discovery
by: Su, Jiayi, et al.
Published: (2025)
by: Su, Jiayi, et al.
Published: (2025)
Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and Generation
by: Estepa, Imanol G., et al.
Published: (2026)
by: Estepa, Imanol G., et al.
Published: (2026)
All4One: Symbiotic Neighbour Contrastive Learning via Self-Attention and Redundancy Reduction
by: Estepa, Imanol G., et al.
Published: (2023)
by: Estepa, Imanol G., et al.
Published: (2023)
FM-AE: Frequency-masked Multimodal Autoencoder for Zinc Electrolysis Plate Contact Abnormality Detection
by: Zhou, Canzong, et al.
Published: (2024)
by: Zhou, Canzong, et al.
Published: (2024)
Transformer-Based Vector Font Classification Using Different Font Formats: TrueType versus PostScript
by: Fujioka, Takumu, et al.
Published: (2025)
by: Fujioka, Takumu, et al.
Published: (2025)
Predicting When to Trust Vision-Language Models for Spatial Reasoning
by: Imran, Muhammad, et al.
Published: (2026)
by: Imran, Muhammad, et al.
Published: (2026)
A General Ambiguity Model for Binary Edge Images with Edge Tracing and its Implementation
by: Hennig, Markus, et al.
Published: (2024)
by: Hennig, Markus, et al.
Published: (2024)
OUS: Scene-Guided Dynamic Facial Expression Recognition
by: Mai, Xinji, et al.
Published: (2024)
by: Mai, Xinji, et al.
Published: (2024)
Evaluation Before Generation: A Paradigm for Robust Multimodal Sentiment Analysis with Missing Modalities
by: Chen, Rongfei, et al.
Published: (2026)
by: Chen, Rongfei, et al.
Published: (2026)
GDDS: A Single Domain Generalized Defect Detection Frame of Open World Scenario using Gather and Distribute Domain-shift Suppression Network
by: Chen, Haiyong, et al.
Published: (2024)
by: Chen, Haiyong, et al.
Published: (2024)
A Persistent Homology Design Space for 3D Point Cloud Deep Learning
by: Kudeshia, Prachi, et al.
Published: (2026)
by: Kudeshia, Prachi, et al.
Published: (2026)
SpATr: MoCap 3D Human Action Recognition based on Spiral Auto-encoder and Transformer Network
by: Bouzid, Hamza, et al.
Published: (2023)
by: Bouzid, Hamza, et al.
Published: (2023)
High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery
by: Peng, Hongxing, et al.
Published: (2025)
by: Peng, Hongxing, et al.
Published: (2025)
CASE: Contrastive Activation for Saliency Estimation
by: Williamson, Dane, et al.
Published: (2025)
by: Williamson, Dane, et al.
Published: (2025)
Characterizing Model Robustness via Natural Input Gradients
by: Rodríguez-Muñoz, Adrián, et al.
Published: (2024)
by: Rodríguez-Muñoz, Adrián, et al.
Published: (2024)
When Does Margin Clamping Affect Training Variance? Dataset-Dependent Effects in Contrastive Forward-Forward Learning
by: Steier, Joshua
Published: (2026)
by: Steier, Joshua
Published: (2026)
Co-Training with Active Contrastive Learning and Meta-Pseudo-Labeling on 2D Projections for Deep Semi-Supervised Learning
by: Aparco-Cardenas, David, et al.
Published: (2025)
by: Aparco-Cardenas, David, et al.
Published: (2025)
Gaze-Guided Learning: Avoiding Shortcut Bias in Visual Classification
by: Li, Jiahang, et al.
Published: (2025)
by: Li, Jiahang, et al.
Published: (2025)
SRPL-SFDA: SAM-Guided Reliable Pseudo-Labels for Source-Free Domain Adaptation in Medical Image Segmentation
by: Liu, Xinya, et al.
Published: (2025)
by: Liu, Xinya, et al.
Published: (2025)
Synergizing Deep Learning and Biological Heuristics for Extreme Long-Tail White Blood Cell Classification
by: Nguyen, Duc T., et al.
Published: (2026)
by: Nguyen, Duc T., et al.
Published: (2026)
Multi-Scale Spatial-Temporal Self-Attention Graph Convolutional Networks for Skeleton-based Action Recognition
by: Nakamura, Ikuo
Published: (2024)
by: Nakamura, Ikuo
Published: (2024)
Robust Confidence Intervals in Stereo Matching using Possibility Theory
by: Malinowski, Roman, et al.
Published: (2024)
by: Malinowski, Roman, et al.
Published: (2024)
Canonical Space Representation for 4D Panoptic Segmentation of Articulated Objects
by: Gomes, Manuel, et al.
Published: (2025)
by: Gomes, Manuel, et al.
Published: (2025)
TailorMe: Self-Supervised Learning of an Anatomically Constrained Volumetric Human Shape Model
by: Wenninger, Stephan, et al.
Published: (2023)
by: Wenninger, Stephan, et al.
Published: (2023)
Attention-Aware Transformer-Based Aggregation Network for Video Periocular Recognition
by: Carreira, Luiz G F, et al.
Published: (2026)
by: Carreira, Luiz G F, et al.
Published: (2026)
DMFourLLIE: Dual-Stage and Multi-Branch Fourier Network for Low-Light Image Enhancement
by: Zhang, Tongshun, et al.
Published: (2024)
by: Zhang, Tongshun, et al.
Published: (2024)
SAM Fewshot Finetuning for Anatomical Segmentation in Medical Images
by: Xie, Weiyi, et al.
Published: (2024)
by: Xie, Weiyi, et al.
Published: (2024)
Adaptive Cascading Network for Continual Test-Time Adaptation
by: Nguyen, Kien X., et al.
Published: (2024)
by: Nguyen, Kien X., et al.
Published: (2024)
Data Augmentation with Diffusion Models for Colon Polyp Localization on the Low Data Regime: How much real data is enough?
by: Tormos, Adrian, et al.
Published: (2024)
by: Tormos, Adrian, et al.
Published: (2024)
RoNFA: Robust Neural Field-based Approach for Few-Shot Image Classification with Noisy Labels
by: Xiang, Nan, et al.
Published: (2025)
by: Xiang, Nan, et al.
Published: (2025)
Separating Knowledge and Perception with Procedural Data
by: Rodríguez-Muñoz, Adrián, et al.
Published: (2025)
by: Rodríguez-Muñoz, Adrián, et al.
Published: (2025)
Segmentation and Characterization of Macerated Fibers and Vessels Using Deep Learning
by: Qamar, Saqib, et al.
Published: (2024)
by: Qamar, Saqib, et al.
Published: (2024)
Supervised deep learning of elastic SRV distances on the shape space of curves
by: Hartman, Emmanuel, et al.
Published: (2021)
by: Hartman, Emmanuel, et al.
Published: (2021)
Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIP
by: Balasubramanian, Sriram, et al.
Published: (2024)
by: Balasubramanian, Sriram, et al.
Published: (2024)
Kolmogorov-Arnold Attention: Is Learnable Attention Better For Vision Transformers?
by: Maity, Subhajit, et al.
Published: (2025)
by: Maity, Subhajit, et al.
Published: (2025)
Prompt-Driven Building Footprint Extraction in Aerial Images with Offset-Building Model
by: Li, Kai, et al.
Published: (2023)
by: Li, Kai, et al.
Published: (2023)
Similar Items
-
A Sensorimotor Vision Transformer
by: Gadzicki, Konrad, et al.
Published: (2025) -
MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
by: Galella, Santiago, et al.
Published: (2026) -
LISTA-Transformer Model Based on Sparse Coding and Attention Mechanism and Its Application in Fault Diagnosis
by: Liu, Shuang, et al.
Published: (2026) -
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
by: Chaybouti, Sofian, et al.
Published: (2025) -
ViTNF: Leveraging Neural Fields to Boost Vision Transformers in Generalized Category Discovery
by: Su, Jiayi, et al.
Published: (2025)