A Light-weight Transformer-based Self-supervised Matching Network for Heterogeneous Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Wang, Li, Tingting, Zhang, Yuntian, Pei, Gensheng, Jiang, Xiruo, Yao, Yazhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semi-supervised Semantic Segmentation with Multi-Constraint Consistency Learning
von: Yin, Jianjian, et al.
Veröffentlicht: (2025)
von: Yin, Jianjian, et al.
Veröffentlicht: (2025)
Relating CNN-Transformer Fusion Network for Change Detection
von: Gao, Yuhao, et al.
Veröffentlicht: (2024)
von: Gao, Yuhao, et al.
Veröffentlicht: (2024)
SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer
von: Zhu, Rui, et al.
Veröffentlicht: (2024)
von: Zhu, Rui, et al.
Veröffentlicht: (2024)
Efficient Token Compression for Vision Transformer with Spatial Information Preserved
von: Mao, Junzhu, et al.
Veröffentlicht: (2025)
von: Mao, Junzhu, et al.
Veröffentlicht: (2025)
Self-supervised Photographic Image Layout Representation Learning
von: Zhao, Zhaoran, et al.
Veröffentlicht: (2024)
von: Zhao, Zhaoran, et al.
Veröffentlicht: (2024)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
von: Song, Zijie, et al.
Veröffentlicht: (2023)
von: Song, Zijie, et al.
Veröffentlicht: (2023)
PEARL: Geometry Aligns Semantics for Training-Free Open-Vocabulary Semantic Segmentation
von: Pei, Gensheng, et al.
Veröffentlicht: (2026)
von: Pei, Gensheng, et al.
Veröffentlicht: (2026)
Taming SAM3 in the Wild: A Concept Bank for Open-Vocabulary Segmentation
von: Pei, Gensheng, et al.
Veröffentlicht: (2026)
von: Pei, Gensheng, et al.
Veröffentlicht: (2026)
VideoMAC: Video Masked Autoencoders Meet ConvNets
von: Pei, Gensheng, et al.
Veröffentlicht: (2024)
von: Pei, Gensheng, et al.
Veröffentlicht: (2024)
Towards Realistic Low-Light Image Enhancement via ISP Driven Data Modeling
von: Wang, Zhihua, et al.
Veröffentlicht: (2025)
von: Wang, Zhihua, et al.
Veröffentlicht: (2025)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
von: Yao, Lei, et al.
Veröffentlicht: (2025)
von: Yao, Lei, et al.
Veröffentlicht: (2025)
Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval
von: Zhang, Bolin, et al.
Veröffentlicht: (2026)
von: Zhang, Bolin, et al.
Veröffentlicht: (2026)
Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and Alignment
von: Zhang, Tong, et al.
Veröffentlicht: (2025)
von: Zhang, Tong, et al.
Veröffentlicht: (2025)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
von: Wang, Xiang, et al.
Veröffentlicht: (2025)
von: Wang, Xiang, et al.
Veröffentlicht: (2025)
SFFNet: Synergistic Feature Fusion Network With Dual-Domain Edge Enhancement for UAV Image Object Detection
von: Zhang, Wenfeng, et al.
Veröffentlicht: (2026)
von: Zhang, Wenfeng, et al.
Veröffentlicht: (2026)
Visual Semantic Description Generation with MLLMs for Image-Text Matching
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
Semi-supervised Chinese Poem-to-Painting Generation via Cycle-consistent Adversarial Networks
von: Lu, Zhengyang, et al.
Veröffentlicht: (2024)
von: Lu, Zhengyang, et al.
Veröffentlicht: (2024)
Deep Boosting Learning: A Brand-new Cooperative Approach for Image-Text Matching
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
Dual-Branch Network for Portrait Image Quality Assessment
von: Sun, Wei, et al.
Veröffentlicht: (2024)
von: Sun, Wei, et al.
Veröffentlicht: (2024)
Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration
von: Jiang, Xun, et al.
Veröffentlicht: (2026)
von: Jiang, Xun, et al.
Veröffentlicht: (2026)
Self-distilled Dynamic Fusion Network for Language-based Fashion Retrieval
von: Wu, Yiming, et al.
Veröffentlicht: (2024)
von: Wu, Yiming, et al.
Veröffentlicht: (2024)
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
von: Cai, Qi, et al.
Veröffentlicht: (2026)
von: Cai, Qi, et al.
Veröffentlicht: (2026)
Serial Low-rank Adaptation of Vision Transformer
von: Zhong, Houqiang, et al.
Veröffentlicht: (2025)
von: Zhong, Houqiang, et al.
Veröffentlicht: (2025)
Learning Efficient Unsupervised Satellite Image-based Building Damage Detection
von: Zhang, Yiyun, et al.
Veröffentlicht: (2023)
von: Zhang, Yiyun, et al.
Veröffentlicht: (2023)
Mitigating Image Captioning Hallucinations in Vision-Language Models
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
von: Cai, Qi, et al.
Veröffentlicht: (2025)
von: Cai, Qi, et al.
Veröffentlicht: (2025)
PanMatch: Unleashing the Potential of Large Vision Models for Unified Matching Models
von: Zhang, Yongjian, et al.
Veröffentlicht: (2025)
von: Zhang, Yongjian, et al.
Veröffentlicht: (2025)
FreeEnhance: Tuning-Free Image Enhancement via Content-Consistent Noising-and-Denoising Process
von: Luo, Yang, et al.
Veröffentlicht: (2024)
von: Luo, Yang, et al.
Veröffentlicht: (2024)
HOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generation
von: Cheng, Hongye, et al.
Veröffentlicht: (2025)
von: Cheng, Hongye, et al.
Veröffentlicht: (2025)
MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and Generalizability
von: Liu, Buyu, et al.
Veröffentlicht: (2024)
von: Liu, Buyu, et al.
Veröffentlicht: (2024)
Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval
von: Zhang, Deyu, et al.
Veröffentlicht: (2025)
von: Zhang, Deyu, et al.
Veröffentlicht: (2025)
Universal Organizer of SAM for Unsupervised Semantic Segmentation
von: Li, Tingting, et al.
Veröffentlicht: (2024)
von: Li, Tingting, et al.
Veröffentlicht: (2024)
Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image Matching
von: Zhang, Yafei, et al.
Veröffentlicht: (2025)
von: Zhang, Yafei, et al.
Veröffentlicht: (2025)
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs
von: Yao, Ting, et al.
Veröffentlicht: (2024)
von: Yao, Ting, et al.
Veröffentlicht: (2024)
HDCompression: Hybrid-Diffusion Image Compression for Ultra-Low Bitrates
von: Lu, Lei, et al.
Veröffentlicht: (2025)
von: Lu, Lei, et al.
Veröffentlicht: (2025)
TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffusion Models
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2024)
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2024)
HeGraphAdapter: Tuning Multi-Modal Vision-Language Models with Heterogeneous Graph Adapter
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
SSTFormer: Bridging Spiking Neural Network and Memory Support Transformer for Frame-Event based Recognition
von: Wang, Xiao, et al.
Veröffentlicht: (2023)
von: Wang, Xiao, et al.
Veröffentlicht: (2023)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
SmartSplat: Feature-Smart Gaussians for Scalable Compression of Ultra-High-Resolution Images
von: Li, Linfei, et al.
Veröffentlicht: (2025)
von: Li, Linfei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Semi-supervised Semantic Segmentation with Multi-Constraint Consistency Learning
von: Yin, Jianjian, et al.
Veröffentlicht: (2025) -
Relating CNN-Transformer Fusion Network for Change Detection
von: Gao, Yuhao, et al.
Veröffentlicht: (2024) -
SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer
von: Zhu, Rui, et al.
Veröffentlicht: (2024) -
Efficient Token Compression for Vision Transformer with Spatial Information Preserved
von: Mao, Junzhu, et al.
Veröffentlicht: (2025) -
Self-supervised Photographic Image Layout Representation Learning
von: Zhao, Zhaoran, et al.
Veröffentlicht: (2024)