LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Thawakar, Omkar, Dissanayake, Dinura, More, Ketan, Thawkar, Ritesh, Heakl, Ahmed, Ahsan, Noor, Li, Yuhao, Zumri, Mohammed, Lahoud, Jean, Anwer, Rao Muhammad, Cholakkal, Hisham, Laptev, Ivan, Shah, Mubarak, Khan, Fahad Shahbaz, Khan, Salman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Good are Foundation Models in Step-by-Step Embodied Reasoning?
von: Dissanayake, Dinura, et al.
Veröffentlicht: (2025)
von: Dissanayake, Dinura, et al.
Veröffentlicht: (2025)
DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding
von: Ishaq, Ayesha, et al.
Veröffentlicht: (2025)
von: Ishaq, Ayesha, et al.
Veröffentlicht: (2025)
Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts
von: Ghaboura, Sara, et al.
Veröffentlicht: (2025)
von: Ghaboura, Sara, et al.
Veröffentlicht: (2025)
Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs
von: Alghallabi, Wafa, et al.
Veröffentlicht: (2025)
von: Alghallabi, Wafa, et al.
Veröffentlicht: (2025)
AIN: The Arabic INclusive Large Multimodal Model
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)
Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2026)
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2026)
Tracking Meets Large Multimodal Models for Driving Scenario Understanding
von: Ishaq, Ayesha, et al.
Veröffentlicht: (2025)
von: Ishaq, Ayesha, et al.
Veröffentlicht: (2025)
ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark
von: Ghaboura, Sara, et al.
Veröffentlicht: (2025)
von: Ghaboura, Sara, et al.
Veröffentlicht: (2025)
DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities
von: Sheikh, Tooba Tehreem, et al.
Veröffentlicht: (2025)
von: Sheikh, Tooba Tehreem, et al.
Veröffentlicht: (2025)
MobiLlama: Towards Accurate and Lightweight Fully Transparent GPT
von: Thawakar, Omkar, et al.
Veröffentlicht: (2024)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2024)
CDChat: A Large Multimodal Model for Remote Sensing Change Description
von: Noman, Mubashir, et al.
Veröffentlicht: (2024)
von: Noman, Mubashir, et al.
Veröffentlicht: (2024)
XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
von: Thawakar, Omkar, et al.
Veröffentlicht: (2023)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2023)
Composed Video Retrieval via Enriched Context and Discriminative Embeddings
von: Thawakar, Omkar, et al.
Veröffentlicht: (2024)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2024)
LLM Post-Training: A Deep Dive into Reasoning Large Language Models
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
von: Kumar, Komal, et al.
Veröffentlicht: (2025)
Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation
von: Boudjoghra, Mohamed El Amine, et al.
Veröffentlicht: (2024)
von: Boudjoghra, Mohamed El Amine, et al.
Veröffentlicht: (2024)
Open3DTrack: Towards Open-Vocabulary 3D Multi-Object Tracking
von: Ishaq, Ayesha, et al.
Veröffentlicht: (2024)
von: Ishaq, Ayesha, et al.
Veröffentlicht: (2024)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
von: Demidov, Dmitry, et al.
Veröffentlicht: (2025)
von: Demidov, Dmitry, et al.
Veröffentlicht: (2025)
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
ELGC-Net: Efficient Local-Global Context Aggregation for Remote Sensing Change Detection
von: Noman, Mubashir, et al.
Veröffentlicht: (2024)
von: Noman, Mubashir, et al.
Veröffentlicht: (2024)
Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework
von: Kumar, Komal, et al.
Veröffentlicht: (2026)
von: Kumar, Komal, et al.
Veröffentlicht: (2026)
AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
von: Nawaz, Umair, et al.
Veröffentlicht: (2025)
von: Nawaz, Umair, et al.
Veröffentlicht: (2025)
BiMediX: Bilingual Medical Mixture of Experts LLM
von: Pieri, Sara, et al.
Veröffentlicht: (2024)
von: Pieri, Sara, et al.
Veröffentlicht: (2024)
LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM
von: Shikhar, Sambal, et al.
Veröffentlicht: (2025)
von: Shikhar, Sambal, et al.
Veröffentlicht: (2025)
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark
von: Ghaboura, Sara, et al.
Veröffentlicht: (2024)
von: Ghaboura, Sara, et al.
Veröffentlicht: (2024)
TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models
von: Luo, Ziyang, et al.
Veröffentlicht: (2025)
von: Luo, Ziyang, et al.
Veröffentlicht: (2025)
CLIMB-3D: Continual Learning for Imbalanced 3D Instance Segmentation
von: Thengane, Vishal, et al.
Veröffentlicht: (2025)
von: Thengane, Vishal, et al.
Veröffentlicht: (2025)
Rethinking Transformers Pre-training for Multi-Spectral Satellite Imagery
von: Noman, Mubashir, et al.
Veröffentlicht: (2024)
von: Noman, Mubashir, et al.
Veröffentlicht: (2024)
Dynamic Pre-training: Towards Efficient and Scalable All-in-One Image Restoration
von: Dudhane, Akshay, et al.
Veröffentlicht: (2024)
von: Dudhane, Akshay, et al.
Veröffentlicht: (2024)
Semi-supervised Open-World Object Detection
von: Mullappilly, Sahal Shaji, et al.
Veröffentlicht: (2024)
von: Mullappilly, Sahal Shaji, et al.
Veröffentlicht: (2024)
CONDA: Condensed Deep Association Learning for Co-Salient Object Detection
von: Li, Long, et al.
Veröffentlicht: (2024)
von: Li, Long, et al.
Veröffentlicht: (2024)
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning
von: Deria, Ankan, et al.
Veröffentlicht: (2026)
von: Deria, Ankan, et al.
Veröffentlicht: (2026)
WorldCache: Content-Aware Caching for Accelerated Video World Models
von: Nawaz, Umair, et al.
Veröffentlicht: (2026)
von: Nawaz, Umair, et al.
Veröffentlicht: (2026)
Video-CoM: Interactive Video Reasoning via Chain of Manipulations
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
von: Shabbir, Akashah, et al.
Veröffentlicht: (2025)
von: Shabbir, Akashah, et al.
Veröffentlicht: (2025)
A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
von: Kurpath, Mohammed Irfan, et al.
Veröffentlicht: (2025)
von: Kurpath, Mohammed Irfan, et al.
Veröffentlicht: (2025)
PARIS3D: Reasoning-based 3D Part Segmentation Using Large Multimodal Model
von: Kareem, Amrin, et al.
Veröffentlicht: (2024)
von: Kareem, Amrin, et al.
Veröffentlicht: (2024)
MediX-R1: Open Ended Medical Reinforcement Learning
von: Mullappilly, Sahal Shaji, et al.
Veröffentlicht: (2026)
von: Mullappilly, Sahal Shaji, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
How Good are Foundation Models in Step-by-Step Embodied Reasoning?
von: Dissanayake, Dinura, et al.
Veröffentlicht: (2025) -
DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding
von: Ishaq, Ayesha, et al.
Veröffentlicht: (2025) -
Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts
von: Ghaboura, Sara, et al.
Veröffentlicht: (2025) -
Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs
von: Alghallabi, Wafa, et al.
Veröffentlicht: (2025) -
AIN: The Arabic INclusive Large Multimodal Model
von: Heakl, Ahmed, et al.
Veröffentlicht: (2025)