Enhancing Agentic Autonomous Scientific Discovery with Vision-Language Model Capabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Gandhi, Kahaan, Bolliet, Boris, Zubeldia, Inigo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Open Source Planning & Control System with Language Agents for Autonomous Scientific Discovery
by: Xu, Licong, et al.
Published: (2025)
by: Xu, Licong, et al.
Published: (2025)
Autonomous Computer Vision Development with Agentic AI
by: Kim, Jin, et al.
Published: (2025)
by: Kim, Jin, et al.
Published: (2025)
Hydra: An Agentic Reasoning Approach for Enhancing Adversarial Robustness and Mitigating Hallucinations in Vision-Language Models
by: Chung-En, et al.
Published: (2025)
by: Chung-En, et al.
Published: (2025)
ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models
by: Yu, Chung-En Johnny, et al.
Published: (2025)
by: Yu, Chung-En Johnny, et al.
Published: (2025)
Metropolis-Hastings Captioning Game: Knowledge Fusion of Vision Language Models via Decentralized Bayesian Inference
by: Matsui, Yuta, et al.
Published: (2025)
by: Matsui, Yuta, et al.
Published: (2025)
An Agentic System for Rare Disease Diagnosis with Traceable Reasoning
by: Zhao, Weike, et al.
Published: (2025)
by: Zhao, Weike, et al.
Published: (2025)
Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers
by: Pang, Wei, et al.
Published: (2025)
by: Pang, Wei, et al.
Published: (2025)
AIDE: Agentically Improve Visual Language Model with Domain Experts
by: Chiu, Ming-Chang, et al.
Published: (2025)
by: Chiu, Ming-Chang, et al.
Published: (2025)
Paper2Video: Automatic Video Generation from Scientific Papers
by: Zhu, Zeyu, et al.
Published: (2025)
by: Zhu, Zeyu, et al.
Published: (2025)
Agentic Knowledgeable Self-awareness
by: Qiao, Shuofei, et al.
Published: (2025)
by: Qiao, Shuofei, et al.
Published: (2025)
Towards Rationality in Language and Multimodal Agents: A Survey
by: Jiang, Bowen, et al.
Published: (2024)
by: Jiang, Bowen, et al.
Published: (2024)
REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents using Information Relevance and Relative Proximity
by: Seo, SeungWon, et al.
Published: (2024)
by: Seo, SeungWon, et al.
Published: (2024)
COMIC: Agentic Sketch Comedy Generation
by: Hong, Susung, et al.
Published: (2026)
by: Hong, Susung, et al.
Published: (2026)
SynWorld: Virtual Scenario Synthesis for Agentic Action Knowledge Refinement
by: Fang, Runnan, et al.
Published: (2025)
by: Fang, Runnan, et al.
Published: (2025)
Concept-RuleNet: Grounded Multi-Agent Neurosymbolic Reasoning in Vision Language Models
by: Sinha, Sanchit, et al.
Published: (2025)
by: Sinha, Sanchit, et al.
Published: (2025)
Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
by: Qin, Yulei, et al.
Published: (2025)
by: Qin, Yulei, et al.
Published: (2025)
Exploring Model Kinship for Merging Large Language Models
by: Hu, Yedi, et al.
Published: (2024)
by: Hu, Yedi, et al.
Published: (2024)
Automated Vehicles Should be Connected with Natural Language
by: Gao, Xiangbo, et al.
Published: (2025)
by: Gao, Xiangbo, et al.
Published: (2025)
Many Heads Are Better Than One: Improved Scientific Idea Generation by A LLM-Based Multi-Agent System
by: Su, Haoyang, et al.
Published: (2024)
by: Su, Haoyang, et al.
Published: (2024)
Chain of Questions: Guiding Multimodal Curiosity in Language Models
by: Iji, Nima, et al.
Published: (2025)
by: Iji, Nima, et al.
Published: (2025)
SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control
by: Lu, Quanfeng, et al.
Published: (2025)
by: Lu, Quanfeng, et al.
Published: (2025)
Curriculum Guided Massive Multi Agent System Solving For Robust Long Horizon Tasks
by: Kar, Indrajit, et al.
Published: (2025)
by: Kar, Indrajit, et al.
Published: (2025)
MARIC: Multi-Agent Reasoning for Image Classification
by: Seo, Wonduk, et al.
Published: (2025)
by: Seo, Wonduk, et al.
Published: (2025)
EH-Benchmark Ophthalmic Hallucination Benchmark and Agent-Driven Top-Down Traceable Reasoning Workflow
by: Pan, Xiaoyu, et al.
Published: (2025)
by: Pan, Xiaoyu, et al.
Published: (2025)
Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
by: Wu, Shengguang, et al.
Published: (2025)
by: Wu, Shengguang, et al.
Published: (2025)
PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to Histopathology
by: Ghezloo, Fatemeh, et al.
Published: (2025)
by: Ghezloo, Fatemeh, et al.
Published: (2025)
PreMind: Multi-Agent Video Understanding for Advanced Indexing of Presentation-style Videos
by: Wei, Kangda, et al.
Published: (2025)
by: Wei, Kangda, et al.
Published: (2025)
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
by: Song, Mingyang, et al.
Published: (2026)
by: Song, Mingyang, et al.
Published: (2026)
TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft
by: Long, Qian, et al.
Published: (2024)
by: Long, Qian, et al.
Published: (2024)
See it. Say it. Sorted: Agentic System for Compositional Diagram Generation
by: Zhang, Hantao, et al.
Published: (2025)
by: Zhang, Hantao, et al.
Published: (2025)
Agentic-J: An AI Agent for Biological Microscopy Image Analysis
by: Johanns, Lukas, et al.
Published: (2026)
by: Johanns, Lukas, et al.
Published: (2026)
PhotoFlow: Agentic 3D Virtual Photography Missions
by: Guo, Jiarui, et al.
Published: (2026)
by: Guo, Jiarui, et al.
Published: (2026)
ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search
by: Kim, Myungchul, et al.
Published: (2026)
by: Kim, Myungchul, et al.
Published: (2026)
MAP: Evaluation and Multi-Agent Enhancement of Large Language Models for Inpatient Pathways
by: Chen, Zhen, et al.
Published: (2025)
by: Chen, Zhen, et al.
Published: (2025)
CoLMDriver: LLM-based Negotiation Benefits Cooperative Autonomous Driving
by: Liu, Changxing, et al.
Published: (2025)
by: Liu, Changxing, et al.
Published: (2025)
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
by: Sun, Zeyi, et al.
Published: (2025)
by: Sun, Zeyi, et al.
Published: (2025)
RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
by: Chen, Tianxing, et al.
Published: (2025)
by: Chen, Tianxing, et al.
Published: (2025)
Agent Planning with World Knowledge Model
by: Qiao, Shuofei, et al.
Published: (2024)
by: Qiao, Shuofei, et al.
Published: (2024)
Judge Model for Large-scale Multimodality Benchmarks
by: Shih, Min-Han, et al.
Published: (2026)
by: Shih, Min-Han, et al.
Published: (2026)
Visual Reasoning Agent: Robust Vision Systems in Remote Sensing via Inference-Time Scaling
by: Yu, Chung-En Johnny, et al.
Published: (2025)
by: Yu, Chung-En Johnny, et al.
Published: (2025)
Similar Items
-
Open Source Planning & Control System with Language Agents for Autonomous Scientific Discovery
by: Xu, Licong, et al.
Published: (2025) -
Autonomous Computer Vision Development with Agentic AI
by: Kim, Jin, et al.
Published: (2025) -
Hydra: An Agentic Reasoning Approach for Enhancing Adversarial Robustness and Mitigating Hallucinations in Vision-Language Models
by: Chung-En, et al.
Published: (2025) -
ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models
by: Yu, Chung-En Johnny, et al.
Published: (2025) -
Metropolis-Hastings Captioning Game: Knowledge Fusion of Vision Language Models via Decentralized Bayesian Inference
by: Matsui, Yuta, et al.
Published: (2025)