Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bajpai, Ashutosh, Majumder, Tamal, Nambi, Akshay, Chakraborty, Tanmoy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2026)
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2026)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
von: Li, Danyang, et al.
Veröffentlicht: (2025)
von: Li, Danyang, et al.
Veröffentlicht: (2025)
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2025)
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2025)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
von: Yang, Shan
Veröffentlicht: (2026)
von: Yang, Shan
Veröffentlicht: (2026)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
von: Lim, Shoon Kit, et al.
Veröffentlicht: (2025)
von: Lim, Shoon Kit, et al.
Veröffentlicht: (2025)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
Learning the meanings of function words from grounded language using a visual question answering model
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
GroundCap: A Visually Grounded Image Captioning Dataset
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
von: Dai, Song, et al.
Veröffentlicht: (2025)
von: Dai, Song, et al.
Veröffentlicht: (2025)
Universal Adversarial Attack on Aligned Multimodal LLMs
von: Rahmatullaev, Temurbek, et al.
Veröffentlicht: (2025)
von: Rahmatullaev, Temurbek, et al.
Veröffentlicht: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
Memory-Efficient Differentially Private Training with Gradient Random Projection
von: Mulrooney, Alex, et al.
Veröffentlicht: (2025)
von: Mulrooney, Alex, et al.
Veröffentlicht: (2025)
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
von: Zhang, Sinin, et al.
Veröffentlicht: (2026)
von: Zhang, Sinin, et al.
Veröffentlicht: (2026)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
ReaderLM-v2: Small Language Model for HTML to Markdown and JSON
von: Wang, Feng, et al.
Veröffentlicht: (2025)
von: Wang, Feng, et al.
Veröffentlicht: (2025)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
von: Anh, Duy Le Dinh, et al.
Veröffentlicht: (2024)
von: Anh, Duy Le Dinh, et al.
Veröffentlicht: (2024)
NAAQA: A Neural Architecture for Acoustic Question Answering
von: Abdelnour, Jerome, et al.
Veröffentlicht: (2021)
von: Abdelnour, Jerome, et al.
Veröffentlicht: (2021)
MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing
von: Skripkin, Matvey, et al.
Veröffentlicht: (2025)
von: Skripkin, Matvey, et al.
Veröffentlicht: (2025)
Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
von: Asanuma, Haruka, et al.
Veröffentlicht: (2025)
von: Asanuma, Haruka, et al.
Veröffentlicht: (2025)
Context-Dependent Affordance Computation in Vision-Language Models
von: Farzulla, Murad
Veröffentlicht: (2026)
von: Farzulla, Murad
Veröffentlicht: (2026)
Unpacking Hateful Memes: Presupposed Context and False Claims
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
von: Simões, Lucca Emmanuel Pineli, et al.
Veröffentlicht: (2024)
von: Simões, Lucca Emmanuel Pineli, et al.
Veröffentlicht: (2024)
Survey Transfer Learning: Recycling Data with Silicon Responses
von: Amini, Ali
Veröffentlicht: (2025)
von: Amini, Ali
Veröffentlicht: (2025)
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
von: Portelance, Eva, et al.
Veröffentlicht: (2024)
von: Portelance, Eva, et al.
Veröffentlicht: (2024)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
von: Huo, Dongjie, et al.
Veröffentlicht: (2026)
von: Huo, Dongjie, et al.
Veröffentlicht: (2026)
Leum-VL Technical Report
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
von: Liu, Bingnan, et al.
Veröffentlicht: (2026)
von: Liu, Bingnan, et al.
Veröffentlicht: (2026)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
von: Yanambakkam, Hemanth Teja, et al.
Veröffentlicht: (2025)
von: Yanambakkam, Hemanth Teja, et al.
Veröffentlicht: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
von: Viveiros, André G., et al.
Veröffentlicht: (2025)
von: Viveiros, André G., et al.
Veröffentlicht: (2025)
Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
von: Agia, Christopher, et al.
Veröffentlicht: (2024)
von: Agia, Christopher, et al.
Veröffentlicht: (2024)
EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM Agents
von: Chen, Junting, et al.
Veröffentlicht: (2024)
von: Chen, Junting, et al.
Veröffentlicht: (2024)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
von: Shah, Nisarg A., et al.
Veröffentlicht: (2025)
von: Shah, Nisarg A., et al.
Veröffentlicht: (2025)
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
von: Bucher, Martin JJ., et al.
Veröffentlicht: (2025)
von: Bucher, Martin JJ., et al.
Veröffentlicht: (2025)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
von: Balasubramanian, Sriram, et al.
Veröffentlicht: (2025)
von: Balasubramanian, Sriram, et al.
Veröffentlicht: (2025)
MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation
von: Bian, Zhipeng, et al.
Veröffentlicht: (2025)
von: Bian, Zhipeng, et al.
Veröffentlicht: (2025)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
von: Le, Van-Truong
Veröffentlicht: (2026)
von: Le, Van-Truong
Veröffentlicht: (2026)
Ähnliche Einträge
-
SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2026) -
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
von: Li, Danyang, et al.
Veröffentlicht: (2025) -
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2025) -
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
von: Kim, Soyeon, et al.
Veröffentlicht: (2026) -
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
von: Yang, Shan
Veröffentlicht: (2026)