Evaluating Contextual Intelligence in Recyclability: A Comprehensive Study of Image-Based Reasoning Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Park, Eliot, Kumar, Abhi, Rajpurkar, Pranav |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ColonCrafter: A Depth Estimation Model for Colonoscopy Videos Using Diffusion Priors
von: Hardy, Romain, et al.
Veröffentlicht: (2025)
von: Hardy, Romain, et al.
Veröffentlicht: (2025)
The Progression of Transformers from Language to Vision to MOT: A Literature Review on Multi-Object Tracking with Transformers
von: Kamboj, Abhi
Veröffentlicht: (2024)
von: Kamboj, Abhi
Veröffentlicht: (2024)
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
von: Wang, Xucheng, et al.
Veröffentlicht: (2026)
von: Wang, Xucheng, et al.
Veröffentlicht: (2026)
Multimodal Foundation Models Exploit Text to Make Medical Image Predictions
von: Buckley, Thomas, et al.
Veröffentlicht: (2023)
von: Buckley, Thomas, et al.
Veröffentlicht: (2023)
3DReasonKnee: Advancing Grounded Reasoning in Medical Vision Language Models
von: Sambara, Sraavya, et al.
Veröffentlicht: (2025)
von: Sambara, Sraavya, et al.
Veröffentlicht: (2025)
Uncovering Knowledge Gaps in Radiology Report Generation Models through Knowledge Graphs
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2024)
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
von: Baharoon, Mohammed, et al.
Veröffentlicht: (2026)
von: Baharoon, Mohammed, et al.
Veröffentlicht: (2026)
a2z-1 for Multi-Disease Detection in Abdomen-Pelvis CT: External Validation and Performance Analysis Across 21 Conditions
von: Rajpurkar, Pranav, et al.
Veröffentlicht: (2024)
von: Rajpurkar, Pranav, et al.
Veröffentlicht: (2024)
Enhancing Image Retrieval : A Comprehensive Study on Photo Search using the CLIP Mode
von: Lahajal, Naresh Kumar, et al.
Veröffentlicht: (2024)
von: Lahajal, Naresh Kumar, et al.
Veröffentlicht: (2024)
A Brief Survey on Leveraging Large Scale Vision Models for Enhanced Robot Grasping
von: Kamboj, Abhi, et al.
Veröffentlicht: (2024)
von: Kamboj, Abhi, et al.
Veröffentlicht: (2024)
SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
von: Wu, Haoning, et al.
Veröffentlicht: (2025)
von: Wu, Haoning, et al.
Veröffentlicht: (2025)
Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation
von: Lei, Yinjie, et al.
Veröffentlicht: (2023)
von: Lei, Yinjie, et al.
Veröffentlicht: (2023)
VIRO: Robust and Efficient Neuro-Symbolic Reasoning with Verification for Referring Expression Comprehension
von: Park, Hyejin, et al.
Veröffentlicht: (2026)
von: Park, Hyejin, et al.
Veröffentlicht: (2026)
ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2024)
Exploring Intrinsic Properties of Medical Images for Self-Supervised Binary Semantic Segmentation
von: Singh, Pranav, et al.
Veröffentlicht: (2024)
von: Singh, Pranav, et al.
Veröffentlicht: (2024)
ReX-MLE: The Autonomous Agent Benchmark for Medical Imaging Challenges
von: Kenia, Roshan, et al.
Veröffentlicht: (2025)
von: Kenia, Roshan, et al.
Veröffentlicht: (2025)
Evaluating Adversarial Protections for Diffusion Personalization: A Comprehensive Study
von: Ye, Kai, et al.
Veröffentlicht: (2025)
von: Ye, Kai, et al.
Veröffentlicht: (2025)
A Large Model for Non-invasive and Personalized Management of Breast Cancer from Multiparametric MRI
von: Luo, Luyang, et al.
Veröffentlicht: (2024)
von: Luo, Luyang, et al.
Veröffentlicht: (2024)
RL-MoE: An Image-Based Privacy Preserving Approach In Intelligent Transportation System
von: Rezaei, Abdolazim, et al.
Veröffentlicht: (2025)
von: Rezaei, Abdolazim, et al.
Veröffentlicht: (2025)
Research on Intelligent Aided Diagnosis System of Medical Image Based on Computer Deep Learning
von: Yuan, Jiajie, et al.
Veröffentlicht: (2024)
von: Yuan, Jiajie, et al.
Veröffentlicht: (2024)
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI
von: Fan, Fanda, et al.
Veröffentlicht: (2024)
von: Fan, Fanda, et al.
Veröffentlicht: (2024)
VARS: Vision-based Assessment of Risk in Security Systems
von: Gupta, Pranav, et al.
Veröffentlicht: (2024)
von: Gupta, Pranav, et al.
Veröffentlicht: (2024)
Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
von: Guo, Xingang, et al.
Veröffentlicht: (2025)
von: Guo, Xingang, et al.
Veröffentlicht: (2025)
CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
von: Li, Jingyao, et al.
Veröffentlicht: (2025)
von: Li, Jingyao, et al.
Veröffentlicht: (2025)
VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
von: Qi, Yukun, et al.
Veröffentlicht: (2025)
von: Qi, Yukun, et al.
Veröffentlicht: (2025)
CUPID: Contextual Understanding of Prompt-conditioned Image Distributions
von: Zhao, Yayan, et al.
Veröffentlicht: (2024)
von: Zhao, Yayan, et al.
Veröffentlicht: (2024)
ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models
von: Yi, Jingwei, et al.
Veröffentlicht: (2025)
von: Yi, Jingwei, et al.
Veröffentlicht: (2025)
YO-CSA-T: A Real-time Badminton Tracking System Utilizing YOLO Based on Contextual and Spatial Attention
von: Lai, Yuan, et al.
Veröffentlicht: (2025)
von: Lai, Yuan, et al.
Veröffentlicht: (2025)
CoReVAD: A Contextual Reasoning Framework for Training-Free Video Anomaly Detection
von: Lim, Hyeongmuk, et al.
Veröffentlicht: (2026)
von: Lim, Hyeongmuk, et al.
Veröffentlicht: (2026)
Transformers Meet Hyperspectral Imaging: A Comprehensive Study of Models, Challenges and Open Problems
von: Zhang, Guyang, et al.
Veröffentlicht: (2025)
von: Zhang, Guyang, et al.
Veröffentlicht: (2025)
Gen-LangSplat: Generalized Language Gaussian Splatting with Pre-Trained Feature Compression
von: Saxena, Pranav
Veröffentlicht: (2025)
von: Saxena, Pranav
Veröffentlicht: (2025)
SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation
von: Pei, Yuhan, et al.
Veröffentlicht: (2024)
von: Pei, Yuhan, et al.
Veröffentlicht: (2024)
Multimodal Contextualized Support for Enhancing Video Retrieval System
von: Nguyen-Le, Quoc-Bao, et al.
Veröffentlicht: (2024)
von: Nguyen-Le, Quoc-Bao, et al.
Veröffentlicht: (2024)
Image Segmentation with Large Language Models: A Survey with Perspectives for Intelligent Transportation Systems
von: Akter, Sanjeda, et al.
Veröffentlicht: (2025)
von: Akter, Sanjeda, et al.
Veröffentlicht: (2025)
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
von: Park, Jinho, et al.
Veröffentlicht: (2026)
von: Park, Jinho, et al.
Veröffentlicht: (2026)
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification
von: Xu, Zhaopan, et al.
Veröffentlicht: (2025)
von: Xu, Zhaopan, et al.
Veröffentlicht: (2025)
SurGo-R1: Benchmarking and Modeling Contextual Reasoning for Operative Zone in Surgical Video
von: Qin, Guanyi, et al.
Veröffentlicht: (2026)
von: Qin, Guanyi, et al.
Veröffentlicht: (2026)
Artificial Intelligence for Geometry-Based Feature Extraction, Analysis and Synthesis in Artistic Images: A Survey
von: Vijendran, Mridula, et al.
Veröffentlicht: (2024)
von: Vijendran, Mridula, et al.
Veröffentlicht: (2024)
A Survey of IMU Based Cross-Modal Transfer Learning in Human Activity Recognition
von: Kamboj, Abhi, et al.
Veröffentlicht: (2024)
von: Kamboj, Abhi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ColonCrafter: A Depth Estimation Model for Colonoscopy Videos Using Diffusion Priors
von: Hardy, Romain, et al.
Veröffentlicht: (2025) -
The Progression of Transformers from Language to Vision to MOT: A Literature Review on Multi-Object Tracking with Transformers
von: Kamboj, Abhi
Veröffentlicht: (2024) -
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
von: Wang, Xucheng, et al.
Veröffentlicht: (2026) -
Multimodal Foundation Models Exploit Text to Make Medical Image Predictions
von: Buckley, Thomas, et al.
Veröffentlicht: (2023) -
3DReasonKnee: Advancing Grounded Reasoning in Medical Vision Language Models
von: Sambara, Sraavya, et al.
Veröffentlicht: (2025)