Multimodal Learning Without Labeled Multimodal Data: Guarantees and Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Paul Pu, Ling, Chun Kai, Cheng, Yun, Obolenskiy, Alex, Liu, Yudong, Pandey, Rohan, Wilf, Alex, Morency, Louis-Philippe, Salakhutdinov, Ruslan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
by: Yu, Haofei, et al.
Published: (2023)
by: Yu, Haofei, et al.
Published: (2023)
HEMM: Holistic Evaluation of Multimodal Foundation Models
by: Liang, Paul Pu, et al.
Published: (2024)
by: Liang, Paul Pu, et al.
Published: (2024)
MultiIoT: Benchmarking Machine Learning for the Internet of Things
by: Mo, Shentong, et al.
Published: (2023)
by: Mo, Shentong, et al.
Published: (2023)
IoT-LM: Large Multisensory Language Models for the Internet of Things
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Social Genome: Grounded Social Reasoning Abilities of Multimodal Models
by: Mathur, Leena, et al.
Published: (2025)
by: Mathur, Leena, et al.
Published: (2025)
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
by: Ma, Martin Q., et al.
Published: (2026)
by: Ma, Martin Q., et al.
Published: (2026)
Act2See: Emergent Active Visual Perception for Video Reasoning
by: Ma, Martin Q., et al.
Published: (2026)
by: Ma, Martin Q., et al.
Published: (2026)
Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions
by: Mathur, Leena, et al.
Published: (2024)
by: Mathur, Leena, et al.
Published: (2024)
Social Caption: Evaluating Social Understanding in Multimodal Models
by: Thumu, Bhaavanaa, et al.
Published: (2026)
by: Thumu, Bhaavanaa, et al.
Published: (2026)
OpenFace 3.0: A Lightweight Multitask System for Comprehensive Facial Behavior Analysis
by: Hu, Jiewen, et al.
Published: (2025)
by: Hu, Jiewen, et al.
Published: (2025)
From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking
by: Kim, Gyeongwon James, et al.
Published: (2025)
by: Kim, Gyeongwon James, et al.
Published: (2025)
Aligning Dialogue Agents with Global Feedback via Large Language Model Multimodal Reward Decomposition
by: Lee, Dong Won, et al.
Published: (2025)
by: Lee, Dong Won, et al.
Published: (2025)
Propose, Solve, Verify: Self-Play Through Formal Verification
by: Wilf, Alex, et al.
Published: (2025)
by: Wilf, Alex, et al.
Published: (2025)
Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback
by: Lee, Dong Won, et al.
Published: (2024)
by: Lee, Dong Won, et al.
Published: (2024)
Dissecting Adversarial Robustness of Multimodal LM Agents
by: Wu, Chen Henry, et al.
Published: (2024)
by: Wu, Chen Henry, et al.
Published: (2024)
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
by: Kapoor, Raghav, et al.
Published: (2024)
by: Kapoor, Raghav, et al.
Published: (2024)
MultiMed: Massively Multimodal and Multitask Medical Understanding
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Exploring Curriculum Learning for Vision-Language Tasks: A Study on Small-Scale Multimodal Training
by: Saha, Rohan, et al.
Published: (2024)
by: Saha, Rohan, et al.
Published: (2024)
Isolated Causal Effects of Natural Language
by: Lin, Victoria, et al.
Published: (2024)
by: Lin, Victoria, et al.
Published: (2024)
Optimizing Language Models for Human Preferences is a Causal Inference Problem
by: Lin, Victoria, et al.
Published: (2024)
by: Lin, Victoria, et al.
Published: (2024)
Omitted Variable Bias in Language Models Under Distribution Shift
by: Lin, Victoria, et al.
Published: (2026)
by: Lin, Victoria, et al.
Published: (2026)
Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
by: Zhang, Ce, et al.
Published: (2025)
by: Zhang, Ce, et al.
Published: (2025)
Automatic Question-Answer Generation for Long-Tail Knowledge
by: Kumar, Rohan, et al.
Published: (2024)
by: Kumar, Rohan, et al.
Published: (2024)
Ranked from Within: Ranking Large Multimodal Models Without Labels
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
Multi-Agent Computer Use
by: Koh, Jing Yu, et al.
Published: (2026)
by: Koh, Jing Yu, et al.
Published: (2026)
Acquiring Linguistic Knowledge from Multimodal Input
by: Amariucai, Theodor, et al.
Published: (2024)
by: Amariucai, Theodor, et al.
Published: (2024)
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
by: Wu, Te-Lin, et al.
Published: (2021)
by: Wu, Te-Lin, et al.
Published: (2021)
Effective Data Augmentation With Diffusion Models
by: Trabucco, Brandon, et al.
Published: (2023)
by: Trabucco, Brandon, et al.
Published: (2023)
Synthetic Patients: Simulating Difficult Conversations with Multimodal Generative AI for Medical Education
by: Chu, Simon N., et al.
Published: (2024)
by: Chu, Simon N., et al.
Published: (2024)
Understanding Visual Concepts Across Models
by: Trabucco, Brandon, et al.
Published: (2024)
by: Trabucco, Brandon, et al.
Published: (2024)
Interactive Sketchpad: A Multimodal Tutoring System for Collaborative, Visual Problem-Solving
by: Chen, Steven-Shine, et al.
Published: (2025)
by: Chen, Steven-Shine, et al.
Published: (2025)
Plan-Seq-Learn: Language Model Guided RL for Solving Long Horizon Robotics Tasks
by: Dalal, Murtaza, et al.
Published: (2024)
by: Dalal, Murtaza, et al.
Published: (2024)
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
by: Jang, Lawrence Keunho, et al.
Published: (2026)
by: Jang, Lawrence Keunho, et al.
Published: (2026)
Balancing Multimodal Training Through Game-Theoretic Regularization
by: Kontras, Konstantinos, et al.
Published: (2024)
by: Kontras, Konstantinos, et al.
Published: (2024)
CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks
by: Suharitdamrong, Wish, et al.
Published: (2026)
by: Suharitdamrong, Wish, et al.
Published: (2026)
POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration
by: Qu, Yuxiao, et al.
Published: (2026)
by: Qu, Yuxiao, et al.
Published: (2026)
Tree Search for Language Model Agents
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
Lessons Learned on the Path to Guaranteeing the Error Bound in Lossy Quantizers
by: Fallin, Alex, et al.
Published: (2024)
by: Fallin, Alex, et al.
Published: (2024)
A Stochastic Approach to the Definition of the Path Integral Measure
by: Obolenskiy, Timur
Published: (2025)
by: Obolenskiy, Timur
Published: (2025)
Similar Items
-
MMoE: Enhancing Multimodal Models with Mixtures of Multimodal Interaction Experts
by: Yu, Haofei, et al.
Published: (2023) -
HEMM: Holistic Evaluation of Multimodal Foundation Models
by: Liang, Paul Pu, et al.
Published: (2024) -
MultiIoT: Benchmarking Machine Learning for the Internet of Things
by: Mo, Shentong, et al.
Published: (2023) -
IoT-LM: Large Multisensory Language Models for the Internet of Things
by: Mo, Shentong, et al.
Published: (2024) -
Social Genome: Grounded Social Reasoning Abilities of Multimodal Models
by: Mathur, Leena, et al.
Published: (2025)