Saved in:
| Main Authors: | Goldshmidt, Roni, Scott, Hamish, Niccolini, Lorenzo, Matzner, Hernan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.05767 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BADAS: Context Aware Collision Prediction Using Real-World Dashcam Data
by: Goldshmidt, Roni, et al.
Published: (2025)
by: Goldshmidt, Roni, et al.
Published: (2025)
Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On
by: Goldshmidt, Roni
Published: (2025)
by: Goldshmidt, Roni
Published: (2025)
VARCO-VISION-2.0 Technical Report
by: Cha, Young-rok, et al.
Published: (2025)
by: Cha, Young-rok, et al.
Published: (2025)
Privacy-Aware Camera 2.0 Technical Report
by: Song, Huan, et al.
Published: (2026)
by: Song, Huan, et al.
Published: (2026)
ChartCheck: Explainable Fact-Checking over Real-World Chart Images
by: Akhtar, Mubashara, et al.
Published: (2023)
by: Akhtar, Mubashara, et al.
Published: (2023)
Beyond Real Weights: Hypercomplex Representations for Stable Quantization
by: Ahad, Jawad Ibn, et al.
Published: (2025)
by: Ahad, Jawad Ibn, et al.
Published: (2025)
Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation
by: Pasca, Razvan-George, et al.
Published: (2023)
by: Pasca, Razvan-George, et al.
Published: (2023)
TokenSHAP: Interpreting Large Language Models with Monte Carlo Shapley Value Estimation
by: Goldshmidt, Roni, et al.
Published: (2024)
by: Goldshmidt, Roni, et al.
Published: (2024)
MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval
by: Kriz, Reno, et al.
Published: (2024)
by: Kriz, Reno, et al.
Published: (2024)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
by: Sun, Yubo, et al.
Published: (2025)
by: Sun, Yubo, et al.
Published: (2025)
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
by: Xing, Long, et al.
Published: (2025)
by: Xing, Long, et al.
Published: (2025)
Cerberus: Real-Time Video Anomaly Detection via Cascaded Vision-Language Models
by: Zheng, Yue, et al.
Published: (2025)
by: Zheng, Yue, et al.
Published: (2025)
RiskProp: Collision-Anchored Self-Supervised Risk Propagation for Early Accident Anticipation
by: Zou, Yiyang, et al.
Published: (2026)
by: Zou, Yiyang, et al.
Published: (2026)
Cross-Modal Rationale Transfer for Explainable Humanitarian Classification on Social Media
by: Nguyen, Thi Huyen, et al.
Published: (2026)
by: Nguyen, Thi Huyen, et al.
Published: (2026)
Towards Transparent AI: A Survey on Explainable Large Language Models
by: Palikhe, Avash, et al.
Published: (2025)
by: Palikhe, Avash, et al.
Published: (2025)
Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models
by: Lin, Junyan, et al.
Published: (2026)
by: Lin, Junyan, et al.
Published: (2026)
Unhackable Temporal Rewarding for Scalable Video MLLMs
by: Yu, En, et al.
Published: (2025)
by: Yu, En, et al.
Published: (2025)
RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
by: Chen, Tianxing, et al.
Published: (2025)
by: Chen, Tianxing, et al.
Published: (2025)
Towards Zero-Shot & Explainable Video Description by Reasoning over Graphs of Events in Space and Time
by: Masala, Mihai, et al.
Published: (2025)
by: Masala, Mihai, et al.
Published: (2025)
From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
by: Masala, Mihai, et al.
Published: (2025)
by: Masala, Mihai, et al.
Published: (2025)
Selectively Answering Visual Questions
by: Eisenschlos, Julian Martin, et al.
Published: (2024)
by: Eisenschlos, Julian Martin, et al.
Published: (2024)
Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration
by: Yang, Honglong, et al.
Published: (2025)
by: Yang, Honglong, et al.
Published: (2025)
Real-time Traffic Accident Anticipation with Feature Reuse
by: Song, Inpyo, et al.
Published: (2025)
by: Song, Inpyo, et al.
Published: (2025)
Beyond Occlusion: In Search for Near Real-Time Explainability of CNN-Based Prostate Cancer Classification
by: Krebs, Martin, et al.
Published: (2025)
by: Krebs, Martin, et al.
Published: (2025)
HazardNet: A Small-Scale Vision Language Model for Real-Time Traffic Safety Detection at Edge Devices
by: Tami, Mohammad Abu, et al.
Published: (2025)
by: Tami, Mohammad Abu, et al.
Published: (2025)
Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability
by: Shu, Dong, et al.
Published: (2025)
by: Shu, Dong, et al.
Published: (2025)
Real Deep Research for AI, Robotics and Beyond
by: Zou, Xueyan, et al.
Published: (2025)
by: Zou, Xueyan, et al.
Published: (2025)
CPJ: Explainable Agricultural Pest Diagnosis via Caption-Prompt-Judge with LLM-Judged Refinement
by: Zhang, Wentao, et al.
Published: (2025)
by: Zhang, Wentao, et al.
Published: (2025)
A CNN-Based Malaria Diagnosis from Blood Cell Images with SHAP and LIME Explainability
by: Abir, Md. Ismiel Hossen, et al.
Published: (2025)
by: Abir, Md. Ismiel Hossen, et al.
Published: (2025)
Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI
by: Shakhadri, Syed Abdul Gaffar, et al.
Published: (2025)
by: Shakhadri, Syed Abdul Gaffar, et al.
Published: (2025)
An Explainable Biomedical Foundation Model via Large-Scale Concept-Enhanced Vision-Language Pre-training
by: Nie, Yuxiang, et al.
Published: (2025)
by: Nie, Yuxiang, et al.
Published: (2025)
A Two-Stage Multitask Vision-Language Framework for Explainable Crop Disease Visual Question Answering
by: Hossain, Md. Zahid, et al.
Published: (2026)
by: Hossain, Md. Zahid, et al.
Published: (2026)
AquaFusionNet: Lightweight VisionSensor Fusion Framework for Real-Time Pathogen Detection and Water Quality Anomaly Prediction on Edge Devices
by: Kristanto, Sepyan Purnama, et al.
Published: (2025)
by: Kristanto, Sepyan Purnama, et al.
Published: (2025)
Real-Time Multimodal Cognitive Assistant for Emergency Medical Services
by: Weerasinghe, Keshara, et al.
Published: (2024)
by: Weerasinghe, Keshara, et al.
Published: (2024)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
by: Xu, Ruyi, et al.
Published: (2025)
by: Xu, Ruyi, et al.
Published: (2025)
BabyVision: Visual Reasoning Beyond Language
by: Chen, Liang, et al.
Published: (2026)
by: Chen, Liang, et al.
Published: (2026)
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
Implementation of Real-Time Lane Detection on Autonomous Mobile Robot
by: Mirdanies, Midriem, et al.
Published: (2024)
by: Mirdanies, Midriem, et al.
Published: (2024)
ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation
by: Wu, Mengyang, et al.
Published: (2024)
by: Wu, Mengyang, et al.
Published: (2024)
Quality-Aware Image-Text Alignment for Opinion-Unaware Image Quality Assessment
by: Agnolucci, Lorenzo, et al.
Published: (2024)
by: Agnolucci, Lorenzo, et al.
Published: (2024)
Similar Items
-
BADAS: Context Aware Collision Prediction Using Real-World Dashcam Data
by: Goldshmidt, Roni, et al.
Published: (2025) -
Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On
by: Goldshmidt, Roni
Published: (2025) -
VARCO-VISION-2.0 Technical Report
by: Cha, Young-rok, et al.
Published: (2025) -
Privacy-Aware Camera 2.0 Technical Report
by: Song, Huan, et al.
Published: (2026) -
ChartCheck: Explainable Fact-Checking over Real-World Chart Images
by: Akhtar, Mubashara, et al.
Published: (2023)