Unsupervised Homography Estimation on Multimodal Image Pair via Alternating Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Sanghyeob, Lew, Jaihyun, Jang, Hyemi, Yoon, Sungroh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
by: Baek, Kanghyun, et al.
Published: (2026)
by: Baek, Kanghyun, et al.
Published: (2026)
Towards Generalized Multimodal Homography Estimation
by: You, Jinkun, et al.
Published: (2026)
by: You, Jinkun, et al.
Published: (2026)
Disentangled Motion Modeling for Video Frame Interpolation
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
HeSS: Head Sensitivity Score for Sparsity Redistribution in VGGT
by: Kim, Yongsung, et al.
Published: (2026)
by: Kim, Yongsung, et al.
Published: (2026)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)
by: Jung, Mingi, et al.
Published: (2025)
Unsupervised Multi-View Visual Anomaly Detection via Progressive Homography-Guided Alignment
by: Chen, Xintao, et al.
Published: (2025)
by: Chen, Xintao, et al.
Published: (2025)
Deep Homography Estimation for Visual Place Recognition
by: Lu, Feng, et al.
Published: (2024)
by: Lu, Feng, et al.
Published: (2024)
Negative-Guided Subject Fidelity Optimization for Zero-Shot Subject-Driven Generation
by: Shin, Chaehun, et al.
Published: (2025)
by: Shin, Chaehun, et al.
Published: (2025)
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
by: Park, Sangha, et al.
Published: (2025)
by: Park, Sangha, et al.
Published: (2025)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination
by: Park, Sangha, et al.
Published: (2025)
by: Park, Sangha, et al.
Published: (2025)
SSHNet: Unsupervised Cross-modal Homography Estimation via Problem Reformulation and Split Optimization
by: Yu, Junchen, et al.
Published: (2024)
by: Yu, Junchen, et al.
Published: (2024)
CKNN: Cleansed k-Nearest Neighbor for Unsupervised Video Anomaly Detection
by: Yi, Jihun, et al.
Published: (2024)
by: Yi, Jihun, et al.
Published: (2024)
SF(DA)$^2$: Source-free Domain Adaptation Through the Lens of Data Augmentation
by: Hwang, Uiwon, et al.
Published: (2024)
by: Hwang, Uiwon, et al.
Published: (2024)
Unified Multimodal Understanding via Byte-Pair Visual Encoding
by: Zhang, Wanpeng, et al.
Published: (2025)
by: Zhang, Wanpeng, et al.
Published: (2025)
Application of 2D Homography for High Resolution Traffic Data Collection using CCTV Cameras
by: Zhang, Linlin, et al.
Published: (2024)
by: Zhang, Linlin, et al.
Published: (2024)
Bidirectional Multimodal Prompt Learning with Scale-Aware Training for Few-Shot Multi-Class Anomaly Detection
by: Lee, Yujin, et al.
Published: (2024)
by: Lee, Yujin, et al.
Published: (2024)
I2AM: Interpreting Image-to-Image Latent Diffusion Models via Bi-Attribution Maps
by: Park, Junseo, et al.
Published: (2024)
by: Park, Junseo, et al.
Published: (2024)
Unsupervised Region-Based Image Editing of Denoising Diffusion Models
by: Li, Zixiang, et al.
Published: (2024)
by: Li, Zixiang, et al.
Published: (2024)
DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual Inspection
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
Source-Free and Image-Only Unsupervised Domain Adaptation for Category Level Object Pose Estimation
by: Kaushik, Prakhar, et al.
Published: (2024)
by: Kaushik, Prakhar, et al.
Published: (2024)
Harmonized Tabular-Image Fusion via Gradient-Aligned Alternating Learning
by: Huang, Longfei, et al.
Published: (2026)
by: Huang, Longfei, et al.
Published: (2026)
Alternative Telescopic Displacement: An Efficient Multimodal Alignment Method
by: Qin, Jiahao, et al.
Published: (2023)
by: Qin, Jiahao, et al.
Published: (2023)
LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text Pairing
by: Girella, Federico, et al.
Published: (2025)
by: Girella, Federico, et al.
Published: (2025)
SCPNet: Unsupervised Cross-modal Homography Estimation via Intra-modal Self-supervised Learning
by: Zhang, Runmin, et al.
Published: (2024)
by: Zhang, Runmin, et al.
Published: (2024)
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
by: Mi, Yapeng, et al.
Published: (2025)
by: Mi, Yapeng, et al.
Published: (2025)
Normality Addition via Normality Detection in Industrial Image Anomaly Detection Models
by: Yi, Jihun, et al.
Published: (2024)
by: Yi, Jihun, et al.
Published: (2024)
Unsupervised Synthetic Image Attribution: Alignment and Disentanglement
by: Liu, Zongfang, et al.
Published: (2026)
by: Liu, Zongfang, et al.
Published: (2026)
When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning
by: Wu, Zhengxian, et al.
Published: (2026)
by: Wu, Zhengxian, et al.
Published: (2026)
TextMatch: Enhancing Image-Text Consistency Through Multimodal Optimization
by: Luo, Yucong, et al.
Published: (2024)
by: Luo, Yucong, et al.
Published: (2024)
Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization
by: Chen, Huiyi, et al.
Published: (2025)
by: Chen, Huiyi, et al.
Published: (2025)
Paired Image Generation with Diffusion-Guided Diffusion Models
by: Zhang, Haoxuan, et al.
Published: (2025)
by: Zhang, Haoxuan, et al.
Published: (2025)
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
by: Yu, Shoubin, et al.
Published: (2024)
by: Yu, Shoubin, et al.
Published: (2024)
Med-PerSAM: One-Shot Visual Prompt Tuning for Personalized Segment Anything Model in Medical Domain
by: Yoon, Hangyul, et al.
Published: (2024)
by: Yoon, Hangyul, et al.
Published: (2024)
Towards Label-Free Brain Tumor Segmentation: Unsupervised Learning with Multimodal MRI
by: Comas-Quiles, Gerard, et al.
Published: (2025)
by: Comas-Quiles, Gerard, et al.
Published: (2025)
Unsupervised Federated Domain Adaptation for Segmentation of MRI Images
by: Nananukul, Navapat, et al.
Published: (2024)
by: Nananukul, Navapat, et al.
Published: (2024)
Advancing Depth Anything Model for Unsupervised Monocular Depth Estimation in Endoscopy
by: Li, Bojian, et al.
Published: (2024)
by: Li, Bojian, et al.
Published: (2024)
UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose Estimation
by: Chen, Haopeng, et al.
Published: (2026)
by: Chen, Haopeng, et al.
Published: (2026)
Beyond Wide-Angle Images: Structure-to-Detail Video Portrait Correction via Unsupervised Spatiotemporal Adaptation
by: Nie, Wenbo, et al.
Published: (2025)
by: Nie, Wenbo, et al.
Published: (2025)
Token-Efficient Multimodal Reasoning via Image Prompt Packaging
by: Choi, Joong Ho, et al.
Published: (2026)
by: Choi, Joong Ho, et al.
Published: (2026)
Similar Items
-
Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
by: Baek, Kanghyun, et al.
Published: (2026) -
Towards Generalized Multimodal Homography Estimation
by: You, Jinkun, et al.
Published: (2026) -
Disentangled Motion Modeling for Video Frame Interpolation
by: Lew, Jaihyun, et al.
Published: (2024) -
HeSS: Head Sensitivity Score for Sparsity Redistribution in VGGT
by: Kim, Yongsung, et al.
Published: (2026) -
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)