Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Xingang, Tyagi, Utkarsh, Gosai, Advait, Vergara, Paula, Park, Jayeon, Montoya, Ernesto Gabriel Hernández, Zhang, Chen Bo Calvin, Hu, Bin, He, Yunzhong, Liu, Bing, Srinivasa, Rakshith Sharma |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
by: Gosai, Advait, et al.
Published: (2025)
by: Gosai, Advait, et al.
Published: (2025)
Beyond Diagnosis: Evaluating Multimodal LLMs for Pathology Localization in Chest Radiographs
by: Gosai, Advait, et al.
Published: (2025)
by: Gosai, Advait, et al.
Published: (2025)
TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models
by: Srinivasa, Rakshith S, et al.
Published: (2025)
by: Srinivasa, Rakshith S, et al.
Published: (2025)
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
by: Tyagi, Utkarsh, et al.
Published: (2026)
by: Tyagi, Utkarsh, et al.
Published: (2026)
Automated ensemble method for pediatric brain tumor segmentation
by: Javaji, Shashidhar Reddy, et al.
Published: (2023)
by: Javaji, Shashidhar Reddy, et al.
Published: (2023)
Seeing with You: Perception-Reasoning Coevolution for Multimodal Reasoning
by: Miao, Ziqi, et al.
Published: (2026)
by: Miao, Ziqi, et al.
Published: (2026)
Scaling Laws for Neural Material Models
by: Trikha, Akshay, et al.
Published: (2025)
by: Trikha, Akshay, et al.
Published: (2025)
To See or To Read: User Behavior Reasoning in Multimodal LLMs
by: Dong, Tianning, et al.
Published: (2025)
by: Dong, Tianning, et al.
Published: (2025)
ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory
by: Xiao, Yunzhong, et al.
Published: (2025)
by: Xiao, Yunzhong, et al.
Published: (2025)
See, Think, Learn: A Self-Taught Multimodal Reasoner
by: Sharma, Sourabh, et al.
Published: (2025)
by: Sharma, Sourabh, et al.
Published: (2025)
PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning
by: Akyürek, Afra Feyza, et al.
Published: (2025)
by: Akyürek, Afra Feyza, et al.
Published: (2025)
From Hallucination to Articulation: Language Model-Driven Losses for Ultra Low-Bitrate Neural Speech Coding
by: Yi, Jayeon, et al.
Published: (2026)
by: Yi, Jayeon, et al.
Published: (2026)
SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?
by: Sehwag, Udari Madhushani, et al.
Published: (2026)
by: Sehwag, Udari Madhushani, et al.
Published: (2026)
Multimodal LLMs See Sentiment
by: da Silva, Neemias B., et al.
Published: (2025)
by: da Silva, Neemias B., et al.
Published: (2025)
Seeing and Reasoning with Confidence: Supercharging Multimodal LLMs with an Uncertainty-Aware Agentic Framework
by: Zhi, Zhuo, et al.
Published: (2025)
by: Zhi, Zhuo, et al.
Published: (2025)
Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs
by: Kanade, Aditya, et al.
Published: (2025)
by: Kanade, Aditya, et al.
Published: (2025)
SeeingEye: Agentic Information Flow Unlocks Multimodal Reasoning In Text-only LLMs
by: Zhang, Weijia, et al.
Published: (2025)
by: Zhang, Weijia, et al.
Published: (2025)
Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing
by: Huang, Kai, et al.
Published: (2025)
by: Huang, Kai, et al.
Published: (2025)
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
by: Lu, Jinda, et al.
Published: (2026)
by: Lu, Jinda, et al.
Published: (2026)
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
by: Kim, Wonjoong, et al.
Published: (2025)
by: Kim, Wonjoong, et al.
Published: (2025)
Enhanced Survival Prediction in Head and Neck Cancer Using Convolutional Block Attention and Multimodal Data Fusion
by: Farooq, Aiman, et al.
Published: (2024)
by: Farooq, Aiman, et al.
Published: (2024)
Proof-of-Perception: Certified Tool-Using Multimodal Reasoning with Compositional Conformal Guarantees
by: Fayyazi, Arya, et al.
Published: (2026)
by: Fayyazi, Arya, et al.
Published: (2026)
DDD: A Perceptually Superior Low-Response-Time DNN-based Declipper
by: Yi, Jayeon, et al.
Published: (2024)
by: Yi, Jayeon, et al.
Published: (2024)
Multi-Mode Inverters: A Unified Control Design for Grid-Forming, Grid-Following, and Beyond
by: Askarian, Alireza, et al.
Published: (2024)
by: Askarian, Alireza, et al.
Published: (2024)
See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
by: Zhang, Shuoshuo, et al.
Published: (2025)
by: Zhang, Shuoshuo, et al.
Published: (2025)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
by: Gunjal, Anisha, et al.
Published: (2025)
by: Gunjal, Anisha, et al.
Published: (2025)
ProxyDet: Synthesizing Proxy Novel Classes via Classwise Mixup for Open-Vocabulary Object Detection
by: Jeong, Joonhyun, et al.
Published: (2023)
by: Jeong, Joonhyun, et al.
Published: (2023)
RestoreGrad: Signal Restoration Using Conditional Denoising Diffusion Models with Jointly Learned Prior
by: Lee, Ching-Hua, et al.
Published: (2025)
by: Lee, Ching-Hua, et al.
Published: (2025)
Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos
by: Shen, Yixuan, et al.
Published: (2026)
by: Shen, Yixuan, et al.
Published: (2026)
Data Lineage and Impact Analysis: Tools and Techniques for Data Governance
by: Srinivasa Rao Karanam
Published: (2022)
by: Srinivasa Rao Karanam
Published: (2022)
MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
by: Selvakumar, Ramaneswaran, et al.
Published: (2025)
by: Selvakumar, Ramaneswaran, et al.
Published: (2025)
SAModified: A Foundation Model-Based Zero-Shot Approach for Refining Noisy Land-Use Land-Cover Maps
by: Pekhale, Sparsh, et al.
Published: (2024)
by: Pekhale, Sparsh, et al.
Published: (2024)
Wings of Change: Motivations, Challenges, and Environmental Concerns Among Birdwatchers
by: Kamal Raj Gosai, et al.
Published: (2026)
by: Kamal Raj Gosai, et al.
Published: (2026)
Enabling Modularity for Spin Qubits via Driven Quantum Dot-Mediated Entanglement
by: Srinivasa, V.
Published: (2026)
by: Srinivasa, V.
Published: (2026)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
by: Juvekar, Kush, et al.
Published: (2025)
by: Juvekar, Kush, et al.
Published: (2025)
COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
by: Guo, Xingang, et al.
Published: (2024)
by: Guo, Xingang, et al.
Published: (2024)
Agentic Rubrics as Contextual Verifiers for SWE Agents
by: Raghavendra, Mohit, et al.
Published: (2026)
by: Raghavendra, Mohit, et al.
Published: (2026)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Matching correlated VAR time series
by: Araya, Ernesto, et al.
Published: (2025)
by: Araya, Ernesto, et al.
Published: (2025)
FAST-Prefill: FPGA Accelerated Sparse Attention for Long Context LLM Prefill
by: Jayanth, Rakshith, et al.
Published: (2026)
by: Jayanth, Rakshith, et al.
Published: (2026)
Similar Items
-
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
by: Gosai, Advait, et al.
Published: (2025) -
Beyond Diagnosis: Evaluating Multimodal LLMs for Pathology Localization in Chest Radiographs
by: Gosai, Advait, et al.
Published: (2025) -
TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models
by: Srinivasa, Rakshith S, et al.
Published: (2025) -
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
by: Tyagi, Utkarsh, et al.
Published: (2026) -
Automated ensemble method for pediatric brain tumor segmentation
by: Javaji, Shashidhar Reddy, et al.
Published: (2023)