Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual Correspondence
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Sunghwan, Cho, Seokju, Kim, Seungryong, Lin, Stephen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation
by: Cho, Seokju, et al.
Published: (2023)
by: Cho, Seokju, et al.
Published: (2023)
Towards Open-Vocabulary Semantic Segmentation Without Semantic Labels
by: Shin, Heeseong, et al.
Published: (2024)
by: Shin, Heeseong, et al.
Published: (2024)
Local All-Pair Correspondence for Point Tracking
by: Cho, Seokju, et al.
Published: (2024)
by: Cho, Seokju, et al.
Published: (2024)
Unifying Correspondence, Pose and NeRF for Pose-Free Novel View Synthesis from Stereo Pairs
by: Hong, Sunghwan, et al.
Published: (2023)
by: Hong, Sunghwan, et al.
Published: (2023)
Cross-View Completion Models are Zero-shot Correspondence Estimators
by: An, Honggyu, et al.
Published: (2024)
by: An, Honggyu, et al.
Published: (2024)
Seurat: From Moving Points to Depth
by: Cho, Seokju, et al.
Published: (2025)
by: Cho, Seokju, et al.
Published: (2025)
Exploring Temporally-Aware Features for Point Tracking
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
Emergent Outlier View Rejection in Visual Geometry Grounded Transformers
by: Han, Jisang, et al.
Published: (2025)
by: Han, Jisang, et al.
Published: (2025)
Multi-Granularity Video Object Segmentation
by: Lim, Sangbeom, et al.
Published: (2024)
by: Lim, Sangbeom, et al.
Published: (2024)
Emergent Temporal Correspondences from Video Diffusion Transformers
by: Nam, Jisu, et al.
Published: (2025)
by: Nam, Jisu, et al.
Published: (2025)
Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers
by: Kim, Chaehyun, et al.
Published: (2025)
by: Kim, Chaehyun, et al.
Published: (2025)
Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models
by: Gröpl, Marcel, et al.
Published: (2026)
by: Gröpl, Marcel, et al.
Published: (2026)
DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation
by: Hong, Susung, et al.
Published: (2023)
by: Hong, Susung, et al.
Published: (2023)
Learning SO(3)-Invariant Semantic Correspondence via Local Shape Transform
by: Park, Chunghyun, et al.
Published: (2024)
by: Park, Chunghyun, et al.
Published: (2024)
TETO: Tracking Events with Teacher Observation for Motion Estimation and Frame Interpolation
by: Yang, Jini, et al.
Published: (2026)
by: Yang, Jini, et al.
Published: (2026)
Match me if you can: Semi-Supervised Semantic Correspondence Learning with Unpaired Images
by: Kim, Jiwon, et al.
Published: (2023)
by: Kim, Jiwon, et al.
Published: (2023)
Similarity-Aware Selective State-Space Modeling for Semantic Correspondence
by: Kim, Seungwook, et al.
Published: (2025)
by: Kim, Seungwook, et al.
Published: (2025)
Unified Diffusion Transformer for High-fidelity Text-Aware Image Restoration
by: Kim, Jin Hyeon, et al.
Published: (2025)
by: Kim, Jin Hyeon, et al.
Published: (2025)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
MV-TAP: Tracking Any Point in Multi-View Videos
by: Koo, Jahyeok, et al.
Published: (2025)
by: Koo, Jahyeok, et al.
Published: (2025)
AnthroTAP: Learning Point Tracking with Real-World Motion
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
by: Kim, Inès Hyeonsu, et al.
Published: (2025)
PF3plat: Pose-Free Feed-Forward 3D Gaussian Splatting
by: Hong, Sunghwan, et al.
Published: (2024)
by: Hong, Sunghwan, et al.
Published: (2024)
SHViT: Single-Head Vision Transformer with Memory Efficient Macro Design
by: Yun, Seokju, et al.
Published: (2024)
by: Yun, Seokju, et al.
Published: (2024)
Visual Representation Alignment for Multimodal Large Language Models
by: Yoon, Heeji, et al.
Published: (2025)
by: Yoon, Heeji, et al.
Published: (2025)
Distillation of Diffusion Features for Semantic Correspondence
by: Fundel, Frank, et al.
Published: (2024)
by: Fundel, Frank, et al.
Published: (2024)
CORAL: Correspondence Alignment for Improved Virtual Try-On
by: Kim, Jiyoung, et al.
Published: (2026)
by: Kim, Jiyoung, et al.
Published: (2026)
Pose-dIVE: Pose-Diversified Augmentation with Diffusion Model for Person Re-Identification
by: Kim, Inès Hyeonsu, et al.
Published: (2024)
by: Kim, Inès Hyeonsu, et al.
Published: (2024)
Semantic Correspondence: Unified Benchmarking and a Strong Baseline
by: Zhang, Kaiyan, et al.
Published: (2025)
by: Zhang, Kaiyan, et al.
Published: (2025)
D$^2$USt3R: Enhancing 3D Reconstruction for Dynamic Scenes
by: Han, Jisang, et al.
Published: (2025)
by: Han, Jisang, et al.
Published: (2025)
Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations
by: Gan, Chaofan, et al.
Published: (2025)
by: Gan, Chaofan, et al.
Published: (2025)
Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation
by: Choi, Jiho, et al.
Published: (2025)
by: Choi, Jiho, et al.
Published: (2025)
CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
by: Kwon, Minkyung, et al.
Published: (2025)
by: Kwon, Minkyung, et al.
Published: (2025)
CorGi: Contribution-Guided Block-Wise Interval Caching for Training-Free Acceleration of Diffusion Transformers
by: Son, Yonglak, et al.
Published: (2025)
by: Son, Yonglak, et al.
Published: (2025)
Gromov Wasserstein Optimal Transport for Semantic Correspondences
by: Snelgar, Francis, et al.
Published: (2026)
by: Snelgar, Francis, et al.
Published: (2026)
RecycleLoRA: Rank-Revealing QR-Based Dual-LoRA Subspace Adaptation for Domain Generalized Semantic Segmentation
by: Cho, Chanseul, et al.
Published: (2026)
by: Cho, Chanseul, et al.
Published: (2026)
DualBEV: Unifying Dual View Transformation with Probabilistic Correspondences
by: Li, Peidong, et al.
Published: (2024)
by: Li, Peidong, et al.
Published: (2024)
3D Scene Prompting for Scene-Consistent Camera-Controllable Video Generation
by: Lee, JoungBin, et al.
Published: (2025)
by: Lee, JoungBin, et al.
Published: (2025)
TAG: Tangential Amplifying Guidance for Hallucination-Resistant Sampling
by: Cho, Hyunmin, et al.
Published: (2025)
by: Cho, Hyunmin, et al.
Published: (2025)
Textual Query-Driven Mask Transformer for Domain Generalized Segmentation
by: Pak, Byeonghyun, et al.
Published: (2024)
by: Pak, Byeonghyun, et al.
Published: (2024)
A Transformer-Based Adaptive Semantic Aggregation Method for UAV Visual Geo-Localization
by: Li, Shishen, et al.
Published: (2024)
by: Li, Shishen, et al.
Published: (2024)
Similar Items
-
CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation
by: Cho, Seokju, et al.
Published: (2023) -
Towards Open-Vocabulary Semantic Segmentation Without Semantic Labels
by: Shin, Heeseong, et al.
Published: (2024) -
Local All-Pair Correspondence for Point Tracking
by: Cho, Seokju, et al.
Published: (2024) -
Unifying Correspondence, Pose and NeRF for Pose-Free Novel View Synthesis from Stereo Pairs
by: Hong, Sunghwan, et al.
Published: (2023) -
Cross-View Completion Models are Zero-shot Correspondence Estimators
by: An, Honggyu, et al.
Published: (2024)