Saved in:
| Main Authors: | Yazdian, Payam Jome, Stanley, Zoe, Lim, Angelica |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.29219 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Salsa as a Nonverbal Embodied Language -- The CoMPAS3D Dataset and Benchmarks
by: Burkanova, Bermet, et al.
Published: (2025)
by: Burkanova, Bermet, et al.
Published: (2025)
MotionScript: Natural Language Descriptions for Expressive 3D Human Motions
by: Yazdian, Payam Jome, et al.
Published: (2023)
by: Yazdian, Payam Jome, et al.
Published: (2023)
Thinker: A vision-language foundation model for embodied intelligence
by: Pan, Baiyu, et al.
Published: (2026)
by: Pan, Baiyu, et al.
Published: (2026)
On the robustness of multimodal language model towards distractions
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
MedViLaM: A multimodal large language model with advanced generalizability and explainability for medical data understanding and generation
by: Xu, Lijian, et al.
Published: (2024)
by: Xu, Lijian, et al.
Published: (2024)
MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
by: Chen, Yanyuan, et al.
Published: (2025)
by: Chen, Yanyuan, et al.
Published: (2025)
In-context learning enables multimodal large language models to classify cancer pathology images
by: Ferber, Dyke, et al.
Published: (2024)
by: Ferber, Dyke, et al.
Published: (2024)
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025)
by: Lv, Haoxin, et al.
Published: (2025)
Chain-of-Caption: Training-free improvement of multimodal large language model on referring expression comprehension
by: Pang, Yik Lung, et al.
Published: (2026)
by: Pang, Yik Lung, et al.
Published: (2026)
Evaluating point-light biological motion in multimodal large language models
by: Kadambi, Akila, et al.
Published: (2025)
by: Kadambi, Akila, et al.
Published: (2025)
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
Elucidating the design space of language models for image generation
by: Liu, Xuantong, et al.
Published: (2024)
by: Liu, Xuantong, et al.
Published: (2024)
GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding
by: Wu, Yiqi, et al.
Published: (2024)
by: Wu, Yiqi, et al.
Published: (2024)
Your Robot Will Feel You Now: Empathy in Robots and Embodied Agents
by: Lim, Angelica, et al.
Published: (2026)
by: Lim, Angelica, et al.
Published: (2026)
Attacks on multimodal models
by: Iablochnikov, Viacheslav, et al.
Published: (2024)
by: Iablochnikov, Viacheslav, et al.
Published: (2024)
Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models
by: Padlewski, Piotr, et al.
Published: (2024)
by: Padlewski, Piotr, et al.
Published: (2024)
What do vision-language models see in the context? Investigating multimodal in-context learning
by: Santos, Gabriel O. dos, et al.
Published: (2025)
by: Santos, Gabriel O. dos, et al.
Published: (2025)
A multimodal gesture recognition dataset for desktop human-computer interaction
by: Wang, Qi, et al.
Published: (2024)
by: Wang, Qi, et al.
Published: (2024)
MAIRA-1: A specialised large multimodal model for radiology report generation
by: Hyland, Stephanie L., et al.
Published: (2023)
by: Hyland, Stephanie L., et al.
Published: (2023)
Explaining latent representations of generative models with large multimodal models
by: Zhu, Mengdan, et al.
Published: (2024)
by: Zhu, Mengdan, et al.
Published: (2024)
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
by: Zhang, Ruixuan, et al.
Published: (2025)
by: Zhang, Ruixuan, et al.
Published: (2025)
Scaling medical imaging report generation with multimodal reinforcement learning
by: Liu, Qianchu, et al.
Published: (2026)
by: Liu, Qianchu, et al.
Published: (2026)
Densification and forecasting of Sentinel-2 time series from multimodal SAR and Optical satellite data using deep generative models
by: Defonte, Véronique, et al.
Published: (2026)
by: Defonte, Véronique, et al.
Published: (2026)
SalsaNext: Fast, Uncertainty-aware Semantic Segmentation of LiDAR Point Clouds for Autonomous Driving
by: Cortinhal, Tiago, et al.
Published: (2020)
by: Cortinhal, Tiago, et al.
Published: (2020)
Astrophotography turbulence mitigation via generative models
by: Kim, Joonyeoup, et al.
Published: (2025)
by: Kim, Joonyeoup, et al.
Published: (2025)
Pulp Motion: Framing-aware multimodal camera and human motion generation
by: Courant, Robin, et al.
Published: (2025)
by: Courant, Robin, et al.
Published: (2025)
MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation
by: Yang, Junjie, et al.
Published: (2025)
by: Yang, Junjie, et al.
Published: (2025)
MM2Latent: Text-to-facial image generation and editing in GANs with multimodal assistance
by: Meng, Debin, et al.
Published: (2024)
by: Meng, Debin, et al.
Published: (2024)
Vision-language models for decoding provider attention during neonatal resuscitation
by: Parodi, Felipe, et al.
Published: (2024)
by: Parodi, Felipe, et al.
Published: (2024)
DeTrack: A Benchmark and Altitude-Aware Dual World Model for Drone-embodied Tracking
by: Hu, Guyue, et al.
Published: (2026)
by: Hu, Guyue, et al.
Published: (2026)
AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation
by: Zhou, Milton, et al.
Published: (2026)
by: Zhou, Milton, et al.
Published: (2026)
ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents
by: Nie, Chang, et al.
Published: (2025)
by: Nie, Chang, et al.
Published: (2025)
OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
by: Zhao, Tiancheng, et al.
Published: (2022)
by: Zhao, Tiancheng, et al.
Published: (2022)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
by: Xing, Yang, et al.
Published: (2026)
by: Xing, Yang, et al.
Published: (2026)
Explaining multimodal LLMs via intra-modal token interactions
by: Liang, Jiawei, et al.
Published: (2025)
by: Liang, Jiawei, et al.
Published: (2025)
ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval
by: Fanelli, Nicola, et al.
Published: (2025)
by: Fanelli, Nicola, et al.
Published: (2025)
Multi-SIGATnet: A multimodal schizophrenia MRI classification algorithm using sparse interaction mechanisms and graph attention networks
by: Jiao, Yuhong, et al.
Published: (2024)
by: Jiao, Yuhong, et al.
Published: (2024)
GenRL: Multimodal-foundation world models for generalization in embodied agents
by: Mazzaglia, Pietro, et al.
Published: (2024)
by: Mazzaglia, Pietro, et al.
Published: (2024)
Emotional Theory of Mind: Bridging Fast Visual Processing with Slow Linguistic Reasoning
by: Etesam, Yasaman, et al.
Published: (2023)
by: Etesam, Yasaman, et al.
Published: (2023)
Contextual Emotion Recognition using Large Vision Language Models
by: Etesam, Yasaman, et al.
Published: (2024)
by: Etesam, Yasaman, et al.
Published: (2024)
Similar Items
-
Salsa as a Nonverbal Embodied Language -- The CoMPAS3D Dataset and Benchmarks
by: Burkanova, Bermet, et al.
Published: (2025) -
MotionScript: Natural Language Descriptions for Expressive 3D Human Motions
by: Yazdian, Payam Jome, et al.
Published: (2023) -
Thinker: A vision-language foundation model for embodied intelligence
by: Pan, Baiyu, et al.
Published: (2026) -
On the robustness of multimodal language model towards distractions
by: Liu, Ming, et al.
Published: (2025) -
MedViLaM: A multimodal large language model with advanced generalizability and explainability for medical data understanding and generation
by: Xu, Lijian, et al.
Published: (2024)