Context-Dependent Affordance Computation in Vision-Language Models
Fuente:
arXiv
Guardado en:
| Autor principal: | Farzulla, Murad |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
por: Viveiros, André G., et al.
Publicado: (2025)
por: Viveiros, André G., et al.
Publicado: (2025)
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025)
por: Cai, Weibin, et al.
Publicado: (2025)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
por: Huo, Dongjie, et al.
Publicado: (2026)
por: Huo, Dongjie, et al.
Publicado: (2026)
A Survey on Vision-Language-Action Models for Embodied AI
por: Ma, Yueen, et al.
Publicado: (2024)
por: Ma, Yueen, et al.
Publicado: (2024)
RACAS: Controlling Diverse Robots With a Single Agentic System
por: Ashley, Dylan R., et al.
Publicado: (2026)
por: Ashley, Dylan R., et al.
Publicado: (2026)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
por: Qi, Dekang, et al.
Publicado: (2026)
por: Qi, Dekang, et al.
Publicado: (2026)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
por: Fu, Tianyu, et al.
Publicado: (2024)
por: Fu, Tianyu, et al.
Publicado: (2024)
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
por: Dumpala, Sri Harsha, et al.
Publicado: (2024)
por: Dumpala, Sri Harsha, et al.
Publicado: (2024)
A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection
por: Malikussaid, et al.
Publicado: (2026)
por: Malikussaid, et al.
Publicado: (2026)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
por: Yim, Wen-wai, et al.
Publicado: (2025)
por: Yim, Wen-wai, et al.
Publicado: (2025)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
por: Sáez, Arnau Igualde, et al.
Publicado: (2025)
por: Sáez, Arnau Igualde, et al.
Publicado: (2025)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
por: Kim, Soyeon, et al.
Publicado: (2026)
por: Kim, Soyeon, et al.
Publicado: (2026)
Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling
por: Ouyang, Rongxin, et al.
Publicado: (2024)
por: Ouyang, Rongxin, et al.
Publicado: (2024)
Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
Autonomous Navigation and Collision Avoidance for Mobile Robots: Classification and Review
por: de Carvalho, Marcus Vinicius Leal, et al.
Publicado: (2024)
por: de Carvalho, Marcus Vinicius Leal, et al.
Publicado: (2024)
The MSR-Video to Text Dataset with Clean Annotations
por: Chen, Haoran, et al.
Publicado: (2021)
por: Chen, Haoran, et al.
Publicado: (2021)
Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
por: Gkountouras, John, et al.
Publicado: (2025)
por: Gkountouras, John, et al.
Publicado: (2025)
ReaderLM-v2: Small Language Model for HTML to Markdown and JSON
por: Wang, Feng, et al.
Publicado: (2025)
por: Wang, Feng, et al.
Publicado: (2025)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
por: Zhang, Junwen, et al.
Publicado: (2025)
por: Zhang, Junwen, et al.
Publicado: (2025)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
To Whom are You Talking? A Deep Learning Model to Endow Social Robots with Addressee Estimation Skills
por: Mazzola, Carlo, et al.
Publicado: (2023)
por: Mazzola, Carlo, et al.
Publicado: (2023)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
por: Yanambakkam, Hemanth Teja, et al.
Publicado: (2025)
por: Yanambakkam, Hemanth Teja, et al.
Publicado: (2025)
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
por: Chahine, Makram, et al.
Publicado: (2024)
por: Chahine, Makram, et al.
Publicado: (2024)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
por: Chen, Kewei, et al.
Publicado: (2026)
por: Chen, Kewei, et al.
Publicado: (2026)
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
por: Zhang, Haichao, et al.
Publicado: (2026)
por: Zhang, Haichao, et al.
Publicado: (2026)
PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications
por: Patel, Hitesh Laxmichand, et al.
Publicado: (2025)
por: Patel, Hitesh Laxmichand, et al.
Publicado: (2025)
SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
por: Gautam, Sushant, et al.
Publicado: (2025)
por: Gautam, Sushant, et al.
Publicado: (2025)
Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
por: Sharma, Aditya, et al.
Publicado: (2025)
por: Sharma, Aditya, et al.
Publicado: (2025)
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
por: Parcalabescu, Letitia, et al.
Publicado: (2023)
por: Parcalabescu, Letitia, et al.
Publicado: (2023)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
por: Agarwal, Amit, et al.
Publicado: (2025)
por: Agarwal, Amit, et al.
Publicado: (2025)
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
por: Cui, Hejie, et al.
Publicado: (2024)
por: Cui, Hejie, et al.
Publicado: (2024)
TensLoRA: Tensor Alternatives for Low-Rank Adaptation
por: Marmoret, Axel, et al.
Publicado: (2025)
por: Marmoret, Axel, et al.
Publicado: (2025)
Measuring What Matters: Scenario-Driven Evaluation for Trajectory Predictors in Autonomous Driving
por: Da, Longchao, et al.
Publicado: (2025)
por: Da, Longchao, et al.
Publicado: (2025)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
Generating Natural-Language Surgical Feedback: From Structured Representation to Domain-Grounded Evaluation
por: Nasriddinov, Firdavs, et al.
Publicado: (2025)
por: Nasriddinov, Firdavs, et al.
Publicado: (2025)
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
por: Papyan, Narek, et al.
Publicado: (2024)
por: Papyan, Narek, et al.
Publicado: (2024)
MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing
por: Skripkin, Matvey, et al.
Publicado: (2025)
por: Skripkin, Matvey, et al.
Publicado: (2025)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
por: Chen, Kewei, et al.
Publicado: (2025)
por: Chen, Kewei, et al.
Publicado: (2025)
Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
por: Menon, Anjali R., et al.
Publicado: (2025)
por: Menon, Anjali R., et al.
Publicado: (2025)
Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
por: Agia, Christopher, et al.
Publicado: (2024)
por: Agia, Christopher, et al.
Publicado: (2024)
Ejemplares similares
-
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
por: Viveiros, André G., et al.
Publicado: (2025) -
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025) -
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
por: Huo, Dongjie, et al.
Publicado: (2026) -
A Survey on Vision-Language-Action Models for Embodied AI
por: Ma, Yueen, et al.
Publicado: (2024) -
RACAS: Controlling Diverse Robots With a Single Agentic System
por: Ashley, Dylan R., et al.
Publicado: (2026)