Similarity-Aware Selective State-Space Modeling for Semantic Correspondence
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Seungwook, Cho, Minsu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning SO(3)-Invariant Semantic Correspondence via Local Shape Transform
by: Park, Chunghyun, et al.
Published: (2024)
by: Park, Chunghyun, et al.
Published: (2024)
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
by: Kim, Seungwook, et al.
Published: (2026)
by: Kim, Seungwook, et al.
Published: (2026)
CorrespondentDream: Enhancing 3D Fidelity of Text-to-3D using Cross-View Correspondences
by: Kim, Seungwook, et al.
Published: (2024)
by: Kim, Seungwook, et al.
Published: (2024)
Harnessing the Power of Training-Free Techniques in Text-to-2D Generation for Text-to-3D Generation via Score Distillation Sampling
by: Lee, Junhong, et al.
Published: (2025)
by: Lee, Junhong, et al.
Published: (2025)
FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
by: Kim, Seungwook, et al.
Published: (2025)
by: Kim, Seungwook, et al.
Published: (2025)
RapidMV: Leveraging Spatio-Angular Representations for Efficient and Consistent Text-to-Multi-View Synthesis
by: Kim, Seungwook, et al.
Published: (2025)
by: Kim, Seungwook, et al.
Published: (2025)
Multi-view Image Prompted Multi-view Diffusion for Improved 3D Generation
by: Kim, Seungwook, et al.
Published: (2024)
by: Kim, Seungwook, et al.
Published: (2024)
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
by: Kim, Manjin, et al.
Published: (2026)
by: Kim, Manjin, et al.
Published: (2026)
Exploring High-Order Self-Similarity for Video Understanding
by: Kim, Manjin, et al.
Published: (2026)
by: Kim, Manjin, et al.
Published: (2026)
Locality-Aware Zero-Shot Human-Object Interaction Detection
by: Kim, Sanghyun, et al.
Published: (2025)
by: Kim, Sanghyun, et al.
Published: (2025)
In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation
by: Kang, Dahyun, et al.
Published: (2024)
by: Kang, Dahyun, et al.
Published: (2024)
Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction Detection
by: Kim, Dongkeun, et al.
Published: (2025)
by: Kim, Dongkeun, et al.
Published: (2025)
3D Geometric Shape Assembly via Efficient Point Cloud Matching
by: Lee, Nahyuk, et al.
Published: (2024)
by: Lee, Nahyuk, et al.
Published: (2024)
Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual Correspondence
by: Hong, Sunghwan, et al.
Published: (2024)
by: Hong, Sunghwan, et al.
Published: (2024)
Treating Motion as Option with Output Selection for Unsupervised Video Object Segmentation
by: Cho, Suhwan, et al.
Published: (2023)
by: Cho, Suhwan, et al.
Published: (2023)
Edge-Aware Image Manipulation via Diffusion Models with a Novel Structure-Preservation Loss
by: Gong, Minsu, et al.
Published: (2026)
by: Gong, Minsu, et al.
Published: (2026)
Burst Image Super-Resolution with Base Frame Selection
by: Kim, Sanghyun, et al.
Published: (2024)
by: Kim, Sanghyun, et al.
Published: (2024)
MARCO: Navigating the Unseen Space of Semantic Correspondence
by: Cuttano, Claudia, et al.
Published: (2026)
by: Cuttano, Claudia, et al.
Published: (2026)
Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping
by: Lee, Junmyeong, et al.
Published: (2026)
by: Lee, Junmyeong, et al.
Published: (2026)
Towards More Practical Group Activity Detection: A New Benchmark and Model
by: Kim, Dongkeun, et al.
Published: (2023)
by: Kim, Dongkeun, et al.
Published: (2023)
Selecting and Pruning: A Differentiable Causal Sequentialized State-Space Model for Two-View Correspondence Learning
by: Fang, Xiang, et al.
Published: (2025)
by: Fang, Xiang, et al.
Published: (2025)
Leveraging 3D Geometric Priors in 2D Rotation Symmetry Detection
by: Seo, Ahyun, et al.
Published: (2025)
by: Seo, Ahyun, et al.
Published: (2025)
Global-Aware Monocular Semantic Scene Completion with State Space Models
by: Li, Shijie, et al.
Published: (2025)
by: Li, Shijie, et al.
Published: (2025)
VideoMamba: Spatio-Temporal Selective State Space Model
by: Park, Jinyoung, et al.
Published: (2024)
by: Park, Jinyoung, et al.
Published: (2024)
Diffusion Hyperfeatures: Searching Through Time and Space for Semantic Correspondence
by: Luo, Grace, et al.
Published: (2023)
by: Luo, Grace, et al.
Published: (2023)
Hierarchically Structured Neural Bones for Reconstructing Animatable Objects from Casual Videos
by: Jeon, Subin, et al.
Published: (2024)
by: Jeon, Subin, et al.
Published: (2024)
Video Summarization with Large Language Models
by: Lee, Min Jung, et al.
Published: (2025)
by: Lee, Min Jung, et al.
Published: (2025)
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers
by: Bae, Jongseong, et al.
Published: (2024)
by: Bae, Jongseong, et al.
Published: (2024)
Online Temporal Action Localization with Memory-Augmented Transformer
by: Song, Youngkil, et al.
Published: (2024)
by: Song, Youngkil, et al.
Published: (2024)
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
by: Zhang, Junyi, et al.
Published: (2023)
by: Zhang, Junyi, et al.
Published: (2023)
RSGMamba: Reliability-Aware Self-Gated State Space Model for Multimodal Semantic Segmentation
by: Xu, Guoan, et al.
Published: (2026)
by: Xu, Guoan, et al.
Published: (2026)
ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation
by: Liu, Yong, et al.
Published: (2025)
by: Liu, Yong, et al.
Published: (2025)
Learning Correlation Structures for Vision Transformers
by: Kim, Manjin, et al.
Published: (2024)
by: Kim, Manjin, et al.
Published: (2024)
Axis-level Symmetry Detection with Group-Equivariant Representation
by: Yu, Wongyun, et al.
Published: (2025)
by: Yu, Wongyun, et al.
Published: (2025)
Contrastive Mean-Shift Learning for Generalized Category Discovery
by: Choi, Sua, et al.
Published: (2024)
by: Choi, Sua, et al.
Published: (2024)
Affostruction: 3D Affordance Grounding with Generative Reconstruction
by: Park, Chunghyun, et al.
Published: (2026)
by: Park, Chunghyun, et al.
Published: (2026)
Rethinking Token Reduction for Diffusion Models via Output-Similarity-Awareness
by: Lee, Hangyeol, et al.
Published: (2026)
by: Lee, Hangyeol, et al.
Published: (2026)
ActFusion: a Unified Diffusion Model for Action Segmentation and Anticipation
by: Gong, Dayoung, et al.
Published: (2024)
by: Gong, Dayoung, et al.
Published: (2024)
Local All-Pair Correspondence for Point Tracking
by: Cho, Seokju, et al.
Published: (2024)
by: Cho, Seokju, et al.
Published: (2024)
Similar Items
-
Learning SO(3)-Invariant Semantic Correspondence via Local Shape Transform
by: Park, Chunghyun, et al.
Published: (2024) -
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
by: Kim, Seungwook, et al.
Published: (2026) -
CorrespondentDream: Enhancing 3D Fidelity of Text-to-3D using Cross-View Correspondences
by: Kim, Seungwook, et al.
Published: (2024) -
Harnessing the Power of Training-Free Techniques in Text-to-2D Generation for Text-to-3D Generation via Score Distillation Sampling
by: Lee, Junhong, et al.
Published: (2025) -
FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
by: Kim, Seungwook, et al.
Published: (2025)