MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jain, Animesh, Stergiou, Alexandros |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
von: Stergiou, Alexandros
Veröffentlicht: (2025)
von: Stergiou, Alexandros
Veröffentlicht: (2025)
LAVIB: A Large-scale Video Interpolation Benchmark
von: Stergiou, Alexandros
Veröffentlicht: (2024)
von: Stergiou, Alexandros
Veröffentlicht: (2024)
About Time: Advances, Challenges, and Outlooks of Action Understanding
von: Stergiou, Alexandros, et al.
Veröffentlicht: (2024)
von: Stergiou, Alexandros, et al.
Veröffentlicht: (2024)
MIMIC: Multimodal Islamophobic Meme Identification and Classification
von: Islam, S M Jishanul, et al.
Veröffentlicht: (2024)
von: Islam, S M Jishanul, et al.
Veröffentlicht: (2024)
Every Shot Counts: Using Exemplars for Repetition Counting in Videos
von: Sinha, Saptarshi, et al.
Veröffentlicht: (2024)
von: Sinha, Saptarshi, et al.
Veröffentlicht: (2024)
EDITOR: Effective and Interpretable Prompt Inversion for Text-to-Image Diffusion Models
von: Li, Mingzhe, et al.
Veröffentlicht: (2025)
von: Li, Mingzhe, et al.
Veröffentlicht: (2025)
Benchmarking CXR Foundation Models With Publicly Available MIMIC-CXR and NIH-CXR14 Datasets
von: Shin, Jiho, et al.
Veröffentlicht: (2025)
von: Shin, Jiho, et al.
Veröffentlicht: (2025)
MIMIC: Masked Image Modeling with Image Correspondences
von: Marathe, Kalyani, et al.
Veröffentlicht: (2023)
von: Marathe, Kalyani, et al.
Veröffentlicht: (2023)
Interpretable 3D Neural Object Volumes for Robust Conceptual Reasoning
von: Pham, Nhi, et al.
Veröffentlicht: (2025)
von: Pham, Nhi, et al.
Veröffentlicht: (2025)
Synthetic Prior for Few-Shot Drivable Head Avatar Inversion
von: Zielonka, Wojciech, et al.
Veröffentlicht: (2025)
von: Zielonka, Wojciech, et al.
Veröffentlicht: (2025)
MIMIC: Mask Image Pre-training with Mix Contrastive Fine-tuning for Facial Expression Recognition
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion
von: Zhang, Libo, et al.
Veröffentlicht: (2025)
von: Zhang, Libo, et al.
Veröffentlicht: (2025)
Leveraging MIMIC Datasets for Better Digital Health: A Review on Open Problems, Progress Highlights, and Future Promises
von: Khaled, Afifa, et al.
Veröffentlicht: (2025)
von: Khaled, Afifa, et al.
Veröffentlicht: (2025)
Interpretable Bilingual Multimodal Large Language Model for Diverse Biomedical Tasks
von: Wang, Lehan, et al.
Veröffentlicht: (2024)
von: Wang, Lehan, et al.
Veröffentlicht: (2024)
SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation
von: Roberts, Jonathan, et al.
Veröffentlicht: (2024)
von: Roberts, Jonathan, et al.
Veröffentlicht: (2024)
Echo-α: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation
von: Zhang, Jing, et al.
Veröffentlicht: (2026)
von: Zhang, Jing, et al.
Veröffentlicht: (2026)
VISTA: A Visual and Textual Attention Dataset for Interpreting Multimodal Models
von: Harshit, et al.
Veröffentlicht: (2024)
von: Harshit, et al.
Veröffentlicht: (2024)
Multimodal Object Detection Under Sparse Forest-Canopy Occlusion
von: Jain, Nitik, et al.
Veröffentlicht: (2026)
von: Jain, Nitik, et al.
Veröffentlicht: (2026)
Tuned Contrastive Learning
von: Animesh, Chaitanya, et al.
Veröffentlicht: (2023)
von: Animesh, Chaitanya, et al.
Veröffentlicht: (2023)
MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering
von: Peng, Jingwei, et al.
Veröffentlicht: (2025)
von: Peng, Jingwei, et al.
Veröffentlicht: (2025)
Conceptual Learning via Embedding Approximations for Reinforcing Interpretability and Transparency
von: Dikter, Maor, et al.
Veröffentlicht: (2024)
von: Dikter, Maor, et al.
Veröffentlicht: (2024)
Oscillation Inversion: Understand the structure of Large Flow Model through the Lens of Inversion Method
von: Zheng, Yan, et al.
Veröffentlicht: (2024)
von: Zheng, Yan, et al.
Veröffentlicht: (2024)
Negative-prompt Inversion: Fast Image Inversion for Editing with Text-guided Diffusion Models
von: Miyake, Daiki, et al.
Veröffentlicht: (2023)
von: Miyake, Daiki, et al.
Veröffentlicht: (2023)
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models
von: Zhang, Guosheng, et al.
Veröffentlicht: (2025)
von: Zhang, Guosheng, et al.
Veröffentlicht: (2025)
Compositional Inversion for Stable Diffusion Models
von: Zhang, Xulu, et al.
Veröffentlicht: (2023)
von: Zhang, Xulu, et al.
Veröffentlicht: (2023)
Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions
von: Ntinou, Ioanna, et al.
Veröffentlicht: (2025)
von: Ntinou, Ioanna, et al.
Veröffentlicht: (2025)
Optimizing Multimodal Language Models through Attention-based Interpretability
von: Sergeev, Alexander, et al.
Veröffentlicht: (2025)
von: Sergeev, Alexander, et al.
Veröffentlicht: (2025)
FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human Recognition
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
Conceptual Codebook Learning for Vision-Language Models
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
Bridging the Semantic Chasm: Synergistic Conceptual Anchoring for Generalized Few-Shot and Zero-Shot OOD Perception
von: Christoforos, Alexandros, et al.
Veröffentlicht: (2026)
von: Christoforos, Alexandros, et al.
Veröffentlicht: (2026)
Model-Agnostic Human Preference Inversion in Diffusion Models
von: Kim, Jeeyung, et al.
Veröffentlicht: (2024)
von: Kim, Jeeyung, et al.
Veröffentlicht: (2024)
Interpretable Measures of Conceptual Similarity by Complexity-Constrained Descriptive Auto-Encoding
von: Achille, Alessandro, et al.
Veröffentlicht: (2024)
von: Achille, Alessandro, et al.
Veröffentlicht: (2024)
Visual Graph Arena: Evaluating Visual Conceptualization of Vision and Multimodal Large Language Models
von: Babaiee, Zahra, et al.
Veröffentlicht: (2025)
von: Babaiee, Zahra, et al.
Veröffentlicht: (2025)
InJecteD: Analyzing Trajectories and Drift Dynamics in Denoising Diffusion Probabilistic Models for 2D Point Cloud Generation
von: Jain, Sanyam, et al.
Veröffentlicht: (2025)
von: Jain, Sanyam, et al.
Veröffentlicht: (2025)
Conditional Evidence Reconstruction and Decomposition for Interpretable Multimodal Diagnosis
von: Wan, Shaowen, et al.
Veröffentlicht: (2026)
von: Wan, Shaowen, et al.
Veröffentlicht: (2026)
On the Vulnerability of Skip Connections to Model Inversion Attacks
von: Koh, Jun Hao, et al.
Veröffentlicht: (2024)
von: Koh, Jun Hao, et al.
Veröffentlicht: (2024)
Probing Conceptual Understanding of Large Visual-Language Models
von: Schiappa, Madeline, et al.
Veröffentlicht: (2023)
von: Schiappa, Madeline, et al.
Veröffentlicht: (2023)
Exploring the Limits of Zero Shot Vision Language Models for Hate Meme Detection: The Vulnerabilities and their Interpretations
von: Rizwan, Naquee, et al.
Veröffentlicht: (2024)
von: Rizwan, Naquee, et al.
Veröffentlicht: (2024)
Vision-Enhanced Large Language Models for High-Resolution Image Synthesis and Multimodal Data Interpretation
von: KV, Karthikeya
Veröffentlicht: (2025)
von: KV, Karthikeya
Veröffentlicht: (2025)
Universal Fingerprint Generation: Controllable Diffusion Model with Multimodal Conditions
von: Grosz, Steven A., et al.
Veröffentlicht: (2024)
von: Grosz, Steven A., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
von: Stergiou, Alexandros
Veröffentlicht: (2025) -
LAVIB: A Large-scale Video Interpolation Benchmark
von: Stergiou, Alexandros
Veröffentlicht: (2024) -
About Time: Advances, Challenges, and Outlooks of Action Understanding
von: Stergiou, Alexandros, et al.
Veröffentlicht: (2024) -
MIMIC: Multimodal Islamophobic Meme Identification and Classification
von: Islam, S M Jishanul, et al.
Veröffentlicht: (2024) -
Every Shot Counts: Using Exemplars for Repetition Counting in Videos
von: Sinha, Saptarshi, et al.
Veröffentlicht: (2024)