Gespeichert in:
| Hauptverfasser: | Wu, Jing, Barretto, Daphne, Chen, Yiye, Gydé, Nicholas, Jian, Yanan, He, Yuhang, Vineet, Vibhav |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2601.20650 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
von: Azad, Shehreen, et al.
Veröffentlicht: (2025)
von: Azad, Shehreen, et al.
Veröffentlicht: (2025)
Physics Knowledge in Frontier Models: A Diagnostic Study of Failure Modes
von: Bagdonaviciute, Ieva, et al.
Veröffentlicht: (2025)
von: Bagdonaviciute, Ieva, et al.
Veröffentlicht: (2025)
StreamReady: Learning What to Answer and When in Long Streaming Videos
von: Azad, Shehreen, et al.
Veröffentlicht: (2026)
von: Azad, Shehreen, et al.
Veröffentlicht: (2026)
GASP: Gaussian Avatars with Synthetic Priors
von: Saunders, Jack, et al.
Veröffentlicht: (2024)
von: Saunders, Jack, et al.
Veröffentlicht: (2024)
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
von: Modi, Rajat, et al.
Veröffentlicht: (2024)
von: Modi, Rajat, et al.
Veröffentlicht: (2024)
Navigating Hallucinations for Reasoning of Unintentional Activities
von: Grover, Shresth, et al.
Veröffentlicht: (2024)
von: Grover, Shresth, et al.
Veröffentlicht: (2024)
Grounding Task Assistance with Multimodal Cues from a Single Demonstration
von: Sarch, Gabriel, et al.
Veröffentlicht: (2025)
von: Sarch, Gabriel, et al.
Veröffentlicht: (2025)
Schema-Guided Scene-Graph Reasoning based on Multi-Agent Large Language Model System
von: Chen, Yiye, et al.
Veröffentlicht: (2025)
von: Chen, Yiye, et al.
Veröffentlicht: (2025)
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
OmViD: Omni-supervised active learning for video action detection
von: Rana, Aayush, et al.
Veröffentlicht: (2025)
von: Rana, Aayush, et al.
Veröffentlicht: (2025)
PEEKABOO: Interactive Video Generation via Masked-Diffusion
von: Jain, Yash, et al.
Veröffentlicht: (2023)
von: Jain, Yash, et al.
Veröffentlicht: (2023)
CarePilot: A Multi-Agent Framework for Long-Horizon Computer Task Automation in Healthcare
von: Ghosh, Akash, et al.
Veröffentlicht: (2026)
von: Ghosh, Akash, et al.
Veröffentlicht: (2026)
Defeasible Visual Entailment: Benchmark, Evaluator, and Reward-Driven Optimization
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
von: Grover, Shresth, et al.
Veröffentlicht: (2025)
von: Grover, Shresth, et al.
Veröffentlicht: (2025)
Understanding Depth and Height Perception in Large Visual-Language Models
von: Azad, Shehreen, et al.
Veröffentlicht: (2024)
von: Azad, Shehreen, et al.
Veröffentlicht: (2024)
MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation
von: Joshi, Siddharth, et al.
Veröffentlicht: (2025)
von: Joshi, Siddharth, et al.
Veröffentlicht: (2025)
BOSS: Benchmark for Observation Space Shift in Long-Horizon Task
von: Yang, Yue, et al.
Veröffentlicht: (2025)
von: Yang, Yue, et al.
Veröffentlicht: (2025)
Fara-7B: An Efficient Agentic Model for Computer Use
von: Awadallah, Ahmed, et al.
Veröffentlicht: (2025)
von: Awadallah, Ahmed, et al.
Veröffentlicht: (2025)
VISTA: Enhancing Visual Conditioning via Track-Following Preference Optimization in Vision-Language-Action Models
von: Chen, Yiye, et al.
Veröffentlicht: (2026)
von: Chen, Yiye, et al.
Veröffentlicht: (2026)
DreamDistribution: Learning Prompt Distribution for Diverse In-distribution Generation
von: Zhao, Brian Nlong, et al.
Veröffentlicht: (2023)
von: Zhao, Brian Nlong, et al.
Veröffentlicht: (2023)
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
von: Ravi, Sahithya, et al.
Veröffentlicht: (2025)
von: Ravi, Sahithya, et al.
Veröffentlicht: (2025)
ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2026)
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2026)
A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
von: Zhang, Yue, et al.
Veröffentlicht: (2026)
OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
von: Hu, Xueyu, et al.
Veröffentlicht: (2025)
von: Hu, Xueyu, et al.
Veröffentlicht: (2025)
Robustness Analysis on Foundational Segmentation Models
von: Schiappa, Madeline Chantry, et al.
Veröffentlicht: (2023)
von: Schiappa, Madeline Chantry, et al.
Veröffentlicht: (2023)
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
von: Wang, Jiayu, et al.
Veröffentlicht: (2024)
von: Wang, Jiayu, et al.
Veröffentlicht: (2024)
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
von: Sun, Zeyi, et al.
Veröffentlicht: (2025)
von: Sun, Zeyi, et al.
Veröffentlicht: (2025)
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
von: Zhang, Shiduo, et al.
Veröffentlicht: (2024)
von: Zhang, Shiduo, et al.
Veröffentlicht: (2024)
EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks
von: Liu, Lulin, et al.
Veröffentlicht: (2026)
von: Liu, Lulin, et al.
Veröffentlicht: (2026)
Generalizable Dense Reward for Long-Horizon Robotic Tasks
von: Yong, Silong, et al.
Veröffentlicht: (2026)
von: Yong, Silong, et al.
Veröffentlicht: (2026)
Chameleon: Episodic Memory for Long-Horizon Robotic Manipulation
von: Guo, Xinying, et al.
Veröffentlicht: (2026)
von: Guo, Xinying, et al.
Veröffentlicht: (2026)
Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method
von: Song, Xinshuai, et al.
Veröffentlicht: (2024)
von: Song, Xinshuai, et al.
Veröffentlicht: (2024)
ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks
von: Wang, Kaijun, et al.
Veröffentlicht: (2025)
von: Wang, Kaijun, et al.
Veröffentlicht: (2025)
Future Predictive Success-or-Failure Classification for Long-Horizon Robotic Tasks
von: Sogi, Naoya, et al.
Veröffentlicht: (2024)
von: Sogi, Naoya, et al.
Veröffentlicht: (2024)
Curriculum Guided Massive Multi Agent System Solving For Robust Long Horizon Tasks
von: Kar, Indrajit, et al.
Veröffentlicht: (2025)
von: Kar, Indrajit, et al.
Veröffentlicht: (2025)
VSearcher: Long-Horizon Multimodal Search Agent via Reinforcement Learning
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2026)
Task Consistent Prototype Learning for Incremental Few-shot Semantic Segmentation
von: Xu, Wenbo, et al.
Veröffentlicht: (2024)
von: Xu, Wenbo, et al.
Veröffentlicht: (2024)
AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI Agents
von: Shi, Yibo, et al.
Veröffentlicht: (2026)
von: Shi, Yibo, et al.
Veröffentlicht: (2026)
CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning
von: Sun, Zeyi, et al.
Veröffentlicht: (2025)
von: Sun, Zeyi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
von: Azad, Shehreen, et al.
Veröffentlicht: (2025) -
Physics Knowledge in Frontier Models: A Diagnostic Study of Failure Modes
von: Bagdonaviciute, Ieva, et al.
Veröffentlicht: (2025) -
StreamReady: Learning What to Answer and When in Long Streaming Videos
von: Azad, Shehreen, et al.
Veröffentlicht: (2026) -
GASP: Gaussian Avatars with Synthetic Priors
von: Saunders, Jack, et al.
Veröffentlicht: (2024) -
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
von: Modi, Rajat, et al.
Veröffentlicht: (2024)