MuM: Multi-View Masked Image Modeling for 3D Vision
Fuente:
arXiv
Saved in:
| Main Authors: | Nordström, David, Edstedt, Johan, Kahl, Fredrik, Bökman, Georg |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Octic Vision Transformers: Quicker ViTs Through Equivariance
by: Nordström, David, et al.
Published: (2025)
by: Nordström, David, et al.
Published: (2025)
Who Handles Orientation? Investigating Invariance in Feature Matching
by: Nordström, David, et al.
Published: (2026)
by: Nordström, David, et al.
Published: (2026)
Flopping for FLOPs: Leveraging equivariance for computational efficiency
by: Bökman, Georg, et al.
Published: (2025)
by: Bökman, Georg, et al.
Published: (2025)
Steerers: A framework for rotation equivariant keypoint descriptors
by: Bökman, Georg, et al.
Published: (2023)
by: Bökman, Georg, et al.
Published: (2023)
Affine steerers for structured keypoint description
by: Bökman, Georg, et al.
Published: (2024)
by: Bökman, Georg, et al.
Published: (2024)
3D-Consistent Multi-View Editing by Correspondence Guidance
by: Bengtson, Josef, et al.
Published: (2025)
by: Bengtson, Josef, et al.
Published: (2025)
LoMa: Local Feature Matching Revisited
by: Nordström, David, et al.
Published: (2026)
by: Nordström, David, et al.
Published: (2026)
DeDoDe v2: Analyzing and Improving the DeDoDe Keypoint Detector
by: Edstedt, Johan, et al.
Published: (2024)
by: Edstedt, Johan, et al.
Published: (2024)
A Framework for Reducing the Complexity of Geometric Vision Problems and its Application to Two-View Triangulation with Approximation Bounds
by: Rydell, Felix, et al.
Published: (2025)
by: Rydell, Felix, et al.
Published: (2025)
DaD: Distilled Reinforcement Learning for Diverse Keypoint Detection
by: Edstedt, Johan, et al.
Published: (2025)
by: Edstedt, Johan, et al.
Published: (2025)
RoMa v2: Harder Better Faster Denser Feature Matching
by: Edstedt, Johan, et al.
Published: (2025)
by: Edstedt, Johan, et al.
Published: (2025)
RefTr: Recurrent Refinement of Confluent Trajectories for 3D Vascular Tree Centerlines
by: Naeem, Roman, et al.
Published: (2025)
by: Naeem, Roman, et al.
Published: (2025)
RoMa: Robust Dense Feature Matching
by: Edstedt, Johan, et al.
Published: (2023)
by: Edstedt, Johan, et al.
Published: (2023)
ARTA: Adaptive Mixed-Resolution Token Allocation for Efficient Dense Feature Extraction
by: Hagerman, David, et al.
Published: (2026)
by: Hagerman, David, et al.
Published: (2026)
Random Token Fusion for Multi-View Medical Diagnosis
by: Guo, Jingyu, et al.
Published: (2024)
by: Guo, Jingyu, et al.
Published: (2024)
MuTri: Multi-view Tri-alignment for OCT to OCTA 3D Image Translation
by: Chen, Zhuangzhuang, et al.
Published: (2025)
by: Chen, Zhuangzhuang, et al.
Published: (2025)
Deep Models for Multi-View 3D Object Recognition: A Review
by: Alzahrani, Mona, et al.
Published: (2024)
by: Alzahrani, Mona, et al.
Published: (2024)
MuMA-ToM: Multi-modal Multi-Agent Theory of Mind
by: Shi, Haojun, et al.
Published: (2024)
by: Shi, Haojun, et al.
Published: (2024)
MIMIC: Masked Image Modeling with Image Correspondences
by: Marathe, Kalyani, et al.
Published: (2023)
by: Marathe, Kalyani, et al.
Published: (2023)
Masked Image Modeling: A Survey
by: Hondru, Vlad, et al.
Published: (2024)
by: Hondru, Vlad, et al.
Published: (2024)
Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models
by: Lian, Chenyu, et al.
Published: (2025)
by: Lian, Chenyu, et al.
Published: (2025)
Multi-View 3D Reconstruction using Knowledge Distillation
by: Dutt, Aditya, et al.
Published: (2024)
by: Dutt, Aditya, et al.
Published: (2024)
HU-based Foreground Masking for 3D Medical Masked Image Modeling
by: Lee, Jin, et al.
Published: (2025)
by: Lee, Jin, et al.
Published: (2025)
Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
by: Lee, Byung-Kwan, et al.
Published: (2025)
by: Lee, Byung-Kwan, et al.
Published: (2025)
Geometric Consistency Refinement for Single Image Novel View Synthesis via Test-Time Adaptation of Diffusion Models
by: Bengtson, Josef, et al.
Published: (2025)
by: Bengtson, Josef, et al.
Published: (2025)
MINR: Implicit Neural Representations with Masked Image Modelling
by: Lee, Sua, et al.
Published: (2025)
by: Lee, Sua, et al.
Published: (2025)
PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation
by: Wahed, Muntasir, et al.
Published: (2024)
by: Wahed, Muntasir, et al.
Published: (2024)
Less Biased Noise Scale Estimation for Threshold-Robust RANSAC
by: Edstedt, Johan
Published: (2025)
by: Edstedt, Johan
Published: (2025)
Label-free Anomaly Detection in Aerial Agricultural Images with Masked Image Modeling
by: Shikhar, Sambal, et al.
Published: (2024)
by: Shikhar, Sambal, et al.
Published: (2024)
DreamComposer: Controllable 3D Object Generation via Multi-View Conditions
by: Yang, Yunhan, et al.
Published: (2023)
by: Yang, Yunhan, et al.
Published: (2023)
Lightweight Cloud Masking Models for On-Board Inference in Hyperspectral Imaging
by: Ali, Mazen, et al.
Published: (2025)
by: Ali, Mazen, et al.
Published: (2025)
Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation
by: Yariv, Guy, et al.
Published: (2025)
by: Yariv, Guy, et al.
Published: (2025)
Blending 3D Geometry and Machine Learning for Multi-View Stereopsis
by: Vats, Vibhas, et al.
Published: (2025)
by: Vats, Vibhas, et al.
Published: (2025)
Multi-View and Multi-Scale Alignment for Contrastive Language-Image Pre-training in Mammography
by: Du, Yuexi, et al.
Published: (2024)
by: Du, Yuexi, et al.
Published: (2024)
A Markovian View of Iterative-Feedback Loops in Image Generative Models: Neural Resonance and Model Collapse
by: Vats, Vibhas Kumar, et al.
Published: (2026)
by: Vats, Vibhas Kumar, et al.
Published: (2026)
PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated Reporting
by: Maqbool, Danyal, et al.
Published: (2025)
by: Maqbool, Danyal, et al.
Published: (2025)
HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation
by: Kumbong, Hermann, et al.
Published: (2025)
by: Kumbong, Hermann, et al.
Published: (2025)
MAPSeg: Unified Unsupervised Domain Adaptation for Heterogeneous Medical Image Segmentation Based on 3D Masked Autoencoding and Pseudo-Labeling
by: Zhang, Xuzhe, et al.
Published: (2023)
by: Zhang, Xuzhe, et al.
Published: (2023)
Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models
by: Wu, Ziyi, et al.
Published: (2024)
by: Wu, Ziyi, et al.
Published: (2024)
Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models
by: Burgess, James, et al.
Published: (2023)
by: Burgess, James, et al.
Published: (2023)
Similar Items
-
Octic Vision Transformers: Quicker ViTs Through Equivariance
by: Nordström, David, et al.
Published: (2025) -
Who Handles Orientation? Investigating Invariance in Feature Matching
by: Nordström, David, et al.
Published: (2026) -
Flopping for FLOPs: Leveraging equivariance for computational efficiency
by: Bökman, Georg, et al.
Published: (2025) -
Steerers: A framework for rotation equivariant keypoint descriptors
by: Bökman, Georg, et al.
Published: (2023) -
Affine steerers for structured keypoint description
by: Bökman, Georg, et al.
Published: (2024)