What You See is (Usually) What You Get: Multimodal Prototype Networks that Abstain from Expensive Modalities
Fuente:
arXiv
Saved in:
| Main Authors: | Bahng, Muchang, Berens, Charlie, Donnelly, Jon, Chen, Eric, Chen, Chaofan, Rudin, Cynthia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interpretable Image Classification with Adaptive Prototype-based Vision Transformers
by: Ma, Chiyu, et al.
Published: (2024)
by: Ma, Chiyu, et al.
Published: (2024)
Rashomon Sets for Prototypical-Part Networks: Editing Interpretable Models in Real-Time
by: Donnelly, Jon, et al.
Published: (2025)
by: Donnelly, Jon, et al.
Published: (2025)
Is What You Ask For What You Get? Investigating Concept Associations in Text-to-Image Models
by: Magid, Salma Abdel, et al.
Published: (2024)
by: Magid, Salma Abdel, et al.
Published: (2024)
What You See is What You Ask: Evaluating Audio Descriptions
by: Kala, Divy, et al.
Published: (2025)
by: Kala, Divy, et al.
Published: (2025)
What You See is What You Classify: Black Box Attributions
by: Stalder, Steven, et al.
Published: (2022)
by: Stalder, Steven, et al.
Published: (2022)
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
by: Li, Senmao, et al.
Published: (2024)
by: Li, Senmao, et al.
Published: (2024)
What You Have is What You Track: Adaptive and Robust Multimodal Tracking
by: Tan, Yuedong, et al.
Published: (2025)
by: Tan, Yuedong, et al.
Published: (2025)
Deformable ProtoPNet: An Interpretable Image Classifier Using Deformable Prototypes
by: Donnelly, Jon, et al.
Published: (2021)
by: Donnelly, Jon, et al.
Published: (2021)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
by: Song, Junha, et al.
Published: (2026)
by: Song, Junha, et al.
Published: (2026)
What Do You See in Common? Learning Hierarchical Prototypes over Tree-of-Life to Discover Evolutionary Traits
by: Manogaran, Harish Babu, et al.
Published: (2024)
by: Manogaran, Harish Babu, et al.
Published: (2024)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
Chameleon: Images Are What You Need For Multimodal Learning Robust To Missing Modalities
by: Liaqat, Muhammad Irzam, et al.
Published: (2024)
by: Liaqat, Muhammad Irzam, et al.
Published: (2024)
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
by: De Nadai, Marco, et al.
Published: (2025)
by: De Nadai, Marco, et al.
Published: (2025)
What Do You See? Enhancing Zero-Shot Image Classification with Multimodal Large Language Models
by: Abdelhamed, Abdelrahman, et al.
Published: (2024)
by: Abdelhamed, Abdelrahman, et al.
Published: (2024)
SfM on-the-fly: Get better 3D from What You Capture
by: Zhan, Zongqian, et al.
Published: (2024)
by: Zhan, Zongqian, et al.
Published: (2024)
Can You Trust What You See? Alpha Channel No-Box Attacks on Video Object Detection
by: Yi, Ariana, et al.
Published: (2025)
by: Yi, Ariana, et al.
Published: (2025)
What Do You See in Vehicle? Comprehensive Vision Solution for In-Vehicle Gaze Estimation
by: Cheng, Yihua, et al.
Published: (2024)
by: Cheng, Yihua, et al.
Published: (2024)
What are You Looking at? Modality Contribution in Multimodal Medical Deep Learning
by: Gapp, Christian, et al.
Published: (2025)
by: Gapp, Christian, et al.
Published: (2025)
SeTformer is What You Need for Vision and Language
by: Shamsolmoali, Pourya, et al.
Published: (2024)
by: Shamsolmoali, Pourya, et al.
Published: (2024)
What You See Is What Matters: A Novel Visual and Physics-Based Metric for Evaluating Video Generation Quality
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
What You See is What You GAN: Rendering Every Pixel for High-Fidelity Geometry in 3D GANs
by: Trevithick, Alex, et al.
Published: (2024)
by: Trevithick, Alex, et al.
Published: (2024)
Tell What You Hear From What You See -- Video to Audio Generation Through Text
by: Liu, Xiulong, et al.
Published: (2024)
by: Liu, Xiulong, et al.
Published: (2024)
See What You Seek: Semantic Contextual Integration for Cloth-Changing Person Re-Identification
by: Han, Xiyu, et al.
Published: (2024)
by: Han, Xiyu, et al.
Published: (2024)
Ground What You See: Hallucination-Resistant MLLMs via Caption Feedback, Diversity-Aware Sampling, and Conflict Regularization
by: Pan, Miao, et al.
Published: (2026)
by: Pan, Miao, et al.
Published: (2026)
Smart Feature is What You Need
by: Hu, Zhaoxin, et al.
Published: (2024)
by: Hu, Zhaoxin, et al.
Published: (2024)
Seeing What You Say: Expressive Image Generation from Speech
by: Lee, Jiyoung, et al.
Published: (2025)
by: Lee, Jiyoung, et al.
Published: (2025)
What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image Segmentation
by: Lin, Jianghang, et al.
Published: (2025)
by: Lin, Jianghang, et al.
Published: (2025)
FPN-IAIA-BL: A Multi-Scale Interpretable Deep Learning Model for Classification of Mass Margins in Digital Mammography
by: Yang, Julia, et al.
Published: (2024)
by: Yang, Julia, et al.
Published: (2024)
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
by: Choi, Yura, et al.
Published: (2026)
by: Choi, Yura, et al.
Published: (2026)
Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
by: Bora, Maheswar, et al.
Published: (2025)
by: Bora, Maheswar, et al.
Published: (2025)
NeIn: Telling What You Don't Want
by: Bui, Nhat-Tan, et al.
Published: (2024)
by: Bui, Nhat-Tan, et al.
Published: (2024)
Tell Me What You See: Text-Guided Real-World Image Denoising
by: Yosef, Erez, et al.
Published: (2023)
by: Yosef, Erez, et al.
Published: (2023)
Eye-See-You: Reverse Pass-Through VR and Head Avatars
by: Dash, Ankan, et al.
Published: (2025)
by: Dash, Ankan, et al.
Published: (2025)
Multi-View Representation is What You Need for Point-Cloud Pre-Training
by: Yan, Siming, et al.
Published: (2023)
by: Yan, Siming, et al.
Published: (2023)
Get In Video: Add Anything You Want to the Video
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
NijiGAN: Transform What You See into Anime with Contrastive Semi-Supervised Learning and Neural Ordinary Differential Equations
by: Santoso, Kevin Putra, et al.
Published: (2024)
by: Santoso, Kevin Putra, et al.
Published: (2024)
How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction
by: Jun, Sejoon, et al.
Published: (2026)
by: Jun, Sejoon, et al.
Published: (2026)
You Only Speak Once to See
by: Yang, Wenhao, et al.
Published: (2024)
by: Yang, Wenhao, et al.
Published: (2024)
Multi-Modal Prototypes for Open-World Semantic Segmentation
by: Yang, Yuhuan, et al.
Published: (2023)
by: Yang, Yuhuan, et al.
Published: (2023)
Point What You Mean: Visually Grounded Instruction Policy
by: Yu, Hang, et al.
Published: (2025)
by: Yu, Hang, et al.
Published: (2025)
Similar Items
-
Interpretable Image Classification with Adaptive Prototype-based Vision Transformers
by: Ma, Chiyu, et al.
Published: (2024) -
Rashomon Sets for Prototypical-Part Networks: Editing Interpretable Models in Real-Time
by: Donnelly, Jon, et al.
Published: (2025) -
Is What You Ask For What You Get? Investigating Concept Associations in Text-to-Image Models
by: Magid, Salma Abdel, et al.
Published: (2024) -
What You See is What You Ask: Evaluating Audio Descriptions
by: Kala, Divy, et al.
Published: (2025) -
What You See is What You Classify: Black Box Attributions
by: Stalder, Steven, et al.
Published: (2022)