A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Kurpath, Mohammed Irfan, Kaithakkodan, Jaseel Muhammad, Zhou, Jinxing, Mullappilly, Sahal Shaji, Almansoori, Mohammad, Ahsan, Noor, Kalmakhanbet, Beknur, Shikhar, Sambal, Lalla, Rishabh, Lahoud, Jean, Awad, Mariette, Khan, Fahad Shahbaz, Khan, Salman, Anwer, Rao Muhammad, Cholakkal, Hisham |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM
by: Shikhar, Sambal, et al.
Published: (2025)
by: Shikhar, Sambal, et al.
Published: (2025)
Semi-supervised Open-World Object Detection
by: Mullappilly, Sahal Shaji, et al.
Published: (2024)
by: Mullappilly, Sahal Shaji, et al.
Published: (2024)
MAviS: A Multimodal Conversational Assistant For Avian Species
by: Kryklyvets, Yevheniia, et al.
Published: (2026)
by: Kryklyvets, Yevheniia, et al.
Published: (2026)
BiMediX: Bilingual Medical Mixture of Experts LLM
by: Pieri, Sara, et al.
Published: (2024)
by: Pieri, Sara, et al.
Published: (2024)
MediX-R1: Open Ended Medical Reinforcement Learning
by: Mullappilly, Sahal Shaji, et al.
Published: (2026)
by: Mullappilly, Sahal Shaji, et al.
Published: (2026)
XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
by: Thawakar, Omkar, et al.
Published: (2023)
by: Thawakar, Omkar, et al.
Published: (2023)
BiMediX2: Bio-Medical EXpert LMM for Diverse Medical Modalities
by: Mullappilly, Sahal Shaji, et al.
Published: (2024)
by: Mullappilly, Sahal Shaji, et al.
Published: (2024)
Tracking Meets Large Multimodal Models for Driving Scenario Understanding
by: Ishaq, Ayesha, et al.
Published: (2025)
by: Ishaq, Ayesha, et al.
Published: (2025)
MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
by: Sheikh, Tooba Tehreem, et al.
Published: (2025)
by: Sheikh, Tooba Tehreem, et al.
Published: (2025)
GLaMM: Pixel Grounding Large Multimodal Model
by: Rasheed, Hanoona, et al.
Published: (2023)
by: Rasheed, Hanoona, et al.
Published: (2023)
Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation
by: Boudjoghra, Mohamed El Amine, et al.
Published: (2024)
by: Boudjoghra, Mohamed El Amine, et al.
Published: (2024)
Open3DTrack: Towards Open-Vocabulary 3D Multi-Object Tracking
by: Ishaq, Ayesha, et al.
Published: (2024)
by: Ishaq, Ayesha, et al.
Published: (2024)
CDChat: A Large Multimodal Model for Remote Sensing Change Description
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
by: Kumar, Komal, et al.
Published: (2025)
by: Kumar, Komal, et al.
Published: (2025)
AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
by: Nawaz, Umair, et al.
Published: (2025)
by: Nawaz, Umair, et al.
Published: (2025)
CLIMB-3D: Continual Learning for Imbalanced 3D Instance Segmentation
by: Thengane, Vishal, et al.
Published: (2025)
by: Thengane, Vishal, et al.
Published: (2025)
AIN: The Arabic INclusive Large Multimodal Model
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
Self-Evolving Multi-Agent Simulations for Realistic Clinical Interactions
by: Almansoori, Mohammad, et al.
Published: (2025)
by: Almansoori, Mohammad, et al.
Published: (2025)
Efficient 3D-Aware Facial Image Editing via Attribute-Specific Prompt Learning
by: Kumar, Amandeep, et al.
Published: (2024)
by: Kumar, Amandeep, et al.
Published: (2024)
DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding
by: Ishaq, Ayesha, et al.
Published: (2025)
by: Ishaq, Ayesha, et al.
Published: (2025)
TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models
by: Luo, Ziyang, et al.
Published: (2025)
by: Luo, Ziyang, et al.
Published: (2025)
How Good are Foundation Models in Step-by-Step Embodied Reasoning?
by: Dissanayake, Dinura, et al.
Published: (2025)
by: Dissanayake, Dinura, et al.
Published: (2025)
PARIS3D: Reasoning-based 3D Part Segmentation Using Large Multimodal Model
by: Kareem, Amrin, et al.
Published: (2024)
by: Kareem, Amrin, et al.
Published: (2024)
Rethinking Transformers Pre-training for Multi-Spectral Satellite Imagery
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
by: Shaker, Abdelrahman, et al.
Published: (2026)
by: Shaker, Abdelrahman, et al.
Published: (2026)
Salient Mask-Guided Vision Transformer for Fine-Grained Classification
by: Demidov, Dmitry, et al.
Published: (2023)
by: Demidov, Dmitry, et al.
Published: (2023)
Multi-modal Generation via Cross-Modal In-Context Learning
by: Kumar, Amandeep, et al.
Published: (2024)
by: Kumar, Amandeep, et al.
Published: (2024)
Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts
by: Ghaboura, Sara, et al.
Published: (2025)
by: Ghaboura, Sara, et al.
Published: (2025)
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
by: Ashraf, Tajamul, et al.
Published: (2025)
by: Ashraf, Tajamul, et al.
Published: (2025)
ELGC-Net: Efficient Local-Global Context Aggregation for Remote Sensing Change Detection
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework
by: Kumar, Komal, et al.
Published: (2026)
by: Kumar, Komal, et al.
Published: (2026)
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
by: Thawakar, Omkar, et al.
Published: (2025)
by: Thawakar, Omkar, et al.
Published: (2025)
Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation
by: Zhou, Jinxing, et al.
Published: (2025)
by: Zhou, Jinxing, et al.
Published: (2025)
Label-free Anomaly Detection in Aerial Agricultural Images with Masked Image Modeling
by: Shikhar, Sambal, et al.
Published: (2024)
by: Shikhar, Sambal, et al.
Published: (2024)
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
by: Thawakar, Omkar, et al.
Published: (2025)
by: Thawakar, Omkar, et al.
Published: (2025)
TPA: Temporal Prompt Alignment for Fetal Congenital Heart Defect Classification
by: Taratynova, Darya, et al.
Published: (2025)
by: Taratynova, Darya, et al.
Published: (2025)
CONDA: Condensed Deep Association Learning for Co-Salient Object Detection
by: Li, Long, et al.
Published: (2024)
by: Li, Long, et al.
Published: (2024)
Audit After Segmentation: Reference-Free Mask Quality Assessment for Language-Referred Audio-Visual Segmentation
by: Zhou, Jinxing, et al.
Published: (2026)
by: Zhou, Jinxing, et al.
Published: (2026)
LLM Post-Training: A Deep Dive into Reasoning Large Language Models
by: Kumar, Komal, et al.
Published: (2025)
by: Kumar, Komal, et al.
Published: (2025)
AgroGPT: Efficient Agricultural Vision-Language Model with Expert Tuning
by: Awais, Muhammad, et al.
Published: (2024)
by: Awais, Muhammad, et al.
Published: (2024)
Similar Items
-
LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM
by: Shikhar, Sambal, et al.
Published: (2025) -
Semi-supervised Open-World Object Detection
by: Mullappilly, Sahal Shaji, et al.
Published: (2024) -
MAviS: A Multimodal Conversational Assistant For Avian Species
by: Kryklyvets, Yevheniia, et al.
Published: (2026) -
BiMediX: Bilingual Medical Mixture of Experts LLM
by: Pieri, Sara, et al.
Published: (2024) -
MediX-R1: Open Ended Medical Reinforcement Learning
by: Mullappilly, Sahal Shaji, et al.
Published: (2026)