MIMIC: Masked Image Modeling with Image Correspondences
Fuente:
arXiv
Saved in:
| Main Authors: | Marathe, Kalyani, Bigverdi, Mahtab, Khan, Nishat, Kundu, Tuhin, Howe, Patrick, S, Sharan Ranjit, Bhattad, Anand, Kembhavi, Aniruddha, Shapiro, Linda G., Krishna, Ranjay |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
by: Zhang, Tianyi, et al.
Published: (2026)
by: Zhang, Tianyi, et al.
Published: (2026)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
by: Bigverdi, Mahtab, et al.
Published: (2024)
by: Bigverdi, Mahtab, et al.
Published: (2024)
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
by: Fan, Xiang, et al.
Published: (2024)
by: Fan, Xiang, et al.
Published: (2024)
MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine
by: Bigverdi, Mahtab, et al.
Published: (2025)
by: Bigverdi, Mahtab, et al.
Published: (2025)
Iterated Learning Improves Compositionality in Large Vision-Language Models
by: Zheng, Chenhao, et al.
Published: (2024)
by: Zheng, Chenhao, et al.
Published: (2024)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
by: Gao, Ziqi, et al.
Published: (2024)
by: Gao, Ziqi, et al.
Published: (2024)
Data Alignment for Zero-Shot Concept Generation in Dermatology AI
by: Gadgil, Soham, et al.
Published: (2024)
by: Gadgil, Soham, et al.
Published: (2024)
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
by: Wallingford, Matthew, et al.
Published: (2024)
by: Wallingford, Matthew, et al.
Published: (2024)
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
by: Li, Linjie, et al.
Published: (2025)
by: Li, Linjie, et al.
Published: (2025)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
by: Yang, Yue, et al.
Published: (2025)
by: Yang, Yue, et al.
Published: (2025)
One Diffusion to Generate Them All
by: Le, Duong H., et al.
Published: (2024)
by: Le, Duong H., et al.
Published: (2024)
Quilt-1M: One Million Image-Text Pairs for Histopathology
by: Ikezogwo, Wisdom Oluchi, et al.
Published: (2023)
by: Ikezogwo, Wisdom Oluchi, et al.
Published: (2023)
CodeNav: Beyond tool-use to using real-world codebases with LLM agents
by: Gupta, Tanmay, et al.
Published: (2024)
by: Gupta, Tanmay, et al.
Published: (2024)
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
by: Seyfioglu, Mehmet Saygin, et al.
Published: (2023)
by: Seyfioglu, Mehmet Saygin, et al.
Published: (2023)
ScribbleLight: Single Image Indoor Relighting with Scribbles
by: Choi, Jun Myeong, et al.
Published: (2024)
by: Choi, Jun Myeong, et al.
Published: (2024)
A General Fixed-Point Theorem for Correspondences
by: Vohra, Ranjit
Published: (2025)
by: Vohra, Ranjit
Published: (2025)
MIMIC: Mask Image Pre-training with Mix Contrastive Fine-tuning for Facial Expression Recognition
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
by: Eftekhar, Ainaz, et al.
Published: (2023)
by: Eftekhar, Ainaz, et al.
Published: (2023)
Task Me Anything
by: Zhang, Jieyu, et al.
Published: (2024)
by: Zhang, Jieyu, et al.
Published: (2024)
Generalizable Sparse-View 3D Reconstruction from Unconstrained Images
by: Gupta, Vinayak, et al.
Published: (2026)
by: Gupta, Vinayak, et al.
Published: (2026)
Gene-Level Representation Learning via Interventional Style Transfer in Optical Pooled Screening
by: Bigverdi, Mahtab, et al.
Published: (2024)
by: Bigverdi, Mahtab, et al.
Published: (2024)
Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation
by: Duggal, Shivam, et al.
Published: (2025)
by: Duggal, Shivam, et al.
Published: (2025)
Agonistic Image Generation: Unsettling the Hegemony of Intention
by: Shaw, Andrew, et al.
Published: (2025)
by: Shaw, Andrew, et al.
Published: (2025)
Laser-cluster interaction in an external magnetic field: the effect of laser polarization
by: Swain, Kalyani, et al.
Published: (2024)
by: Swain, Kalyani, et al.
Published: (2024)
ZeroComp: Zero-shot Object Compositing from Image Intrinsics via Diffusion
by: Zhang, Zitian, et al.
Published: (2024)
by: Zhang, Zitian, et al.
Published: (2024)
MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives
by: Ikezogwo, Wisdom O., et al.
Published: (2025)
by: Ikezogwo, Wisdom O., et al.
Published: (2025)
A Comparative Study of U-Net Architectures for Change Detection in Satellite Images
by: Amin, Yaxita, et al.
Published: (2025)
by: Amin, Yaxita, et al.
Published: (2025)
Keypoint Aware Masked Image Modelling
by: Krishna, Madhava, et al.
Published: (2024)
by: Krishna, Madhava, et al.
Published: (2024)
Neuro-Parametric Spectral Classification of Black Hole and Neutron Star X-ray Binary Systems
by: Garg, Akash, et al.
Published: (2026)
by: Garg, Akash, et al.
Published: (2026)
Visual Jenga: Discovering Object Dependencies via Counterfactual Inpainting
by: Bhattad, Anand, et al.
Published: (2025)
by: Bhattad, Anand, et al.
Published: (2025)
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
by: Maharana, Adyasha, et al.
Published: (2023)
by: Maharana, Adyasha, et al.
Published: (2023)
La producción cultural de artistas y escritores "afrocubanos" en el período revolucionario
by: Linda S. Howe
Published: (2001)
by: Linda S. Howe
Published: (2001)
Semantics-Aware Attention Guidance for Diagnosing Whole Slide Images
by: Liu, Kechun, et al.
Published: (2024)
by: Liu, Kechun, et al.
Published: (2024)
Semantic and Expressive Variation in Image Captions Across Languages
by: Ye, Andre, et al.
Published: (2023)
by: Ye, Andre, et al.
Published: (2023)
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
by: Ray, Arijit, et al.
Published: (2024)
by: Ray, Arijit, et al.
Published: (2024)
Generative Models: What Do They Know? Do They Know Things? Let's Find Out!
by: Du, Xiaodan, et al.
Published: (2023)
by: Du, Xiaodan, et al.
Published: (2023)
“Vulva,” Not “Private Part”: The Importance of Accurate Genital Terminology
by: Hannah R. Chang, et al.
Published: (2024)
by: Hannah R. Chang, et al.
Published: (2024)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
by: Islam, Tanzima Z., et al.
Published: (2024)
by: Islam, Tanzima Z., et al.
Published: (2024)
Harmonic Mobile Manipulation
by: Yang, Ruihan, et al.
Published: (2023)
by: Yang, Ruihan, et al.
Published: (2023)
Attaining the Sabatier Sweet Spot via d‐Orbital Occupancy Engineering for High‐Performance Zn–Air Batteries
by: Srijib Das, et al.
Published: (2026)
by: Srijib Das, et al.
Published: (2026)
Similar Items
-
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
by: Zhang, Tianyi, et al.
Published: (2026) -
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
by: Bigverdi, Mahtab, et al.
Published: (2024) -
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
by: Fan, Xiang, et al.
Published: (2024) -
MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine
by: Bigverdi, Mahtab, et al.
Published: (2025) -
Iterated Learning Improves Compositionality in Large Vision-Language Models
by: Zheng, Chenhao, et al.
Published: (2024)