Saved in:
| Main Authors: | Verghese, Mrinal, Chen, Brian, Eghbalzadeh, Hamid, Nagarajan, Tushar, Desai, Ruta |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2408.03160 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Instructions to Assistance: a Dataset Aligning Instruction Manuals with Assembly Videos for Evaluating Multimodal LLMs
by: Toschi, Federico, et al.
Published: (2026)
by: Toschi, Federico, et al.
Published: (2026)
VITED: Video Temporal Evidence Distillation
by: Lu, Yujie, et al.
Published: (2025)
by: Lu, Yujie, et al.
Published: (2025)
Learning Latent Action World Models In The Wild
by: Garrido, Quentin, et al.
Published: (2026)
by: Garrido, Quentin, et al.
Published: (2026)
Comparative Evaluation of Hard and Soft Clustering for Precise Brain Tumor Segmentation in MR Imaging
by: Bora, Dibya Jyoti, et al.
Published: (2025)
by: Bora, Dibya Jyoti, et al.
Published: (2025)
Simulating Clinical AI Assistance using Multimodal LLMs: A Case Study in Diabetic Retinopathy
by: Barakat, Nadim, et al.
Published: (2025)
by: Barakat, Nadim, et al.
Published: (2025)
Vision-Integrated LLMs for Autonomous Driving Assistance : Human Performance Comparison and Trust Evaluation
by: Kim, Namhee, et al.
Published: (2025)
by: Kim, Namhee, et al.
Published: (2025)
Beyond Diagnosis: Evaluating Multimodal LLMs for Pathology Localization in Chest Radiographs
by: Gosai, Advait, et al.
Published: (2025)
by: Gosai, Advait, et al.
Published: (2025)
Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
by: Kurz, Paul Jonas, et al.
Published: (2026)
by: Kurz, Paul Jonas, et al.
Published: (2026)
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
by: Peng, Tianhao, et al.
Published: (2025)
by: Peng, Tianhao, et al.
Published: (2025)
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
by: Hong, Jack, et al.
Published: (2025)
by: Hong, Jack, et al.
Published: (2025)
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
by: Patil, Vaidehi, et al.
Published: (2025)
by: Patil, Vaidehi, et al.
Published: (2025)
Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
by: Guo, Xingang, et al.
Published: (2025)
by: Guo, Xingang, et al.
Published: (2025)
Multimodal Crowd Counting with Pix2Pix GANs
by: Khan, Muhammad Asif, et al.
Published: (2024)
by: Khan, Muhammad Asif, et al.
Published: (2024)
Beyond CLIP: Knowledge-Enhanced Multimodal Transformers for Cross-Modal Alignment in Diabetic Retinopathy Diagnosis
by: Samanta, Argha Kamal, et al.
Published: (2025)
by: Samanta, Argha Kamal, et al.
Published: (2025)
SO-Bench: A Structural Output Evaluation of Multimodal LLMs
by: Feng, Di, et al.
Published: (2025)
by: Feng, Di, et al.
Published: (2025)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
by: Fu, Chaoyou, et al.
Published: (2024)
by: Fu, Chaoyou, et al.
Published: (2024)
Personalized Video Summarization by Multimodal Video Understanding
by: Chen, Brian, et al.
Published: (2024)
by: Chen, Brian, et al.
Published: (2024)
Are Multimodal LLMs Ready for Clinical Dermatology? A Real-World Evaluation in Dermatology
by: Jiang, Roy, et al.
Published: (2026)
by: Jiang, Roy, et al.
Published: (2026)
MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
by: Wei, Yake, et al.
Published: (2024)
by: Wei, Yake, et al.
Published: (2024)
Real Time Offside Detection using a Single Camera in Soccer
by: Desai, Shounak
Published: (2025)
by: Desai, Shounak
Published: (2025)
BayesSDF: Surface-Based Laplacian Uncertainty Estimation for 3D Geometry with Neural Signed Distance Fields
by: Desai, Rushil
Published: (2025)
by: Desai, Rushil
Published: (2025)
A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures
by: Terver, Basile, et al.
Published: (2026)
by: Terver, Basile, et al.
Published: (2026)
Seeing the Big Picture: Evaluating Multimodal LLMs' Ability to Interpret and Grade Handwritten Student Work
by: Henkel, Owen, et al.
Published: (2025)
by: Henkel, Owen, et al.
Published: (2025)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
by: Xie, Yuxuan, et al.
Published: (2024)
by: Xie, Yuxuan, et al.
Published: (2024)
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
by: Song, Tingyu, et al.
Published: (2025)
by: Song, Tingyu, et al.
Published: (2025)
Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
by: Sarto, Sara, et al.
Published: (2025)
by: Sarto, Sara, et al.
Published: (2025)
MIBench: Evaluating LMMs on Multimodal Interaction
by: Miao, Yu, et al.
Published: (2026)
by: Miao, Yu, et al.
Published: (2026)
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion
by: Chen, Zhuokun, et al.
Published: (2024)
by: Chen, Zhuokun, et al.
Published: (2024)
Soiling detection for Advanced Driver Assistance Systems
by: Beránek, Filip, et al.
Published: (2025)
by: Beránek, Filip, et al.
Published: (2025)
Text-Guided Layer Fusion Mitigates Hallucination in Multimodal LLMs
by: Lin, Chenchen, et al.
Published: (2026)
by: Lin, Chenchen, et al.
Published: (2026)
Hierarchical and Multimodal Data for Daily Activity Understanding
by: Kaviani, Ghazal, et al.
Published: (2025)
by: Kaviani, Ghazal, et al.
Published: (2025)
FLOW: Fusing and Shuffling Global and Local Views for Cross-User Human Activity Recognition with IMUs
by: Qiu, Qi, et al.
Published: (2024)
by: Qiu, Qi, et al.
Published: (2024)
Leveraging ChatGPT's Multimodal Vision Capabilities to Rank Satellite Images by Poverty Level: Advancing Tools for Social Science Research
by: Sarmadi, Hamid, et al.
Published: (2025)
by: Sarmadi, Hamid, et al.
Published: (2025)
Rethinking Visual Information Processing in Multimodal LLMs
by: Kim, Dongwan, et al.
Published: (2025)
by: Kim, Dongwan, et al.
Published: (2025)
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
by: Pan, Yaning, et al.
Published: (2025)
by: Pan, Yaning, et al.
Published: (2025)
Tri-Reader: An Open-Access, Multi-Stage AI Pipeline for First-Pass Lung Nodule Annotation in Screening CT
by: Tushar, Fakrul Islam, et al.
Published: (2026)
by: Tushar, Fakrul Islam, et al.
Published: (2026)
Incorporating Eye-Tracking Signals Into Multimodal Deep Visual Models For Predicting User Aesthetic Experience In Residential Interiors
by: Chien, Chen-Ying, et al.
Published: (2026)
by: Chien, Chen-Ying, et al.
Published: (2026)
Demographic User Modeling for Social Robotics with Multimodal Pre-trained Models
by: Rahimi, Hamed, et al.
Published: (2025)
by: Rahimi, Hamed, et al.
Published: (2025)
Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
by: Fan, Zhaoyu, et al.
Published: (2025)
by: Fan, Zhaoyu, et al.
Published: (2025)
Can AI Assistance Aid in the Grading of Handwritten Answer Sheets?
by: Sil, Pritam, et al.
Published: (2024)
by: Sil, Pritam, et al.
Published: (2024)
Similar Items
-
From Instructions to Assistance: a Dataset Aligning Instruction Manuals with Assembly Videos for Evaluating Multimodal LLMs
by: Toschi, Federico, et al.
Published: (2026) -
VITED: Video Temporal Evidence Distillation
by: Lu, Yujie, et al.
Published: (2025) -
Learning Latent Action World Models In The Wild
by: Garrido, Quentin, et al.
Published: (2026) -
Comparative Evaluation of Hard and Soft Clustering for Precise Brain Tumor Segmentation in MR Imaging
by: Bora, Dibya Jyoti, et al.
Published: (2025) -
Simulating Clinical AI Assistance using Multimodal LLMs: A Case Study in Diabetic Retinopathy
by: Barakat, Nadim, et al.
Published: (2025)