Grounding Computer Use Agents on Human Demonstrations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feizi, Aarash, Nayak, Shravan, Jian, Xiangru, Lin, Kevin Qinghong, Li, Kaixin, Awal, Rabiul, Lù, Xing Han, Obando-Ceron, Johan, Rodriguez, Juan A., Chapados, Nicolas, Vazquez, David, Romero-Soriano, Adriana, Rabbany, Reihaneh, Taslakian, Perouz, Pal, Christopher, Gella, Spandana, Rajeswar, Sai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
von: Jian, Xiangru, et al.
Veröffentlicht: (2026)
von: Jian, Xiangru, et al.
Veröffentlicht: (2026)
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
von: Feizi, Aarash, et al.
Veröffentlicht: (2025)
von: Feizi, Aarash, et al.
Veröffentlicht: (2025)
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
von: Nayak, Shravan, et al.
Veröffentlicht: (2025)
von: Nayak, Shravan, et al.
Veröffentlicht: (2025)
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
von: Awal, Rabiul, et al.
Veröffentlicht: (2025)
von: Awal, Rabiul, et al.
Veröffentlicht: (2025)
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
von: Rodriguez, Juan A., et al.
Veröffentlicht: (2025)
von: Rodriguez, Juan A., et al.
Veröffentlicht: (2025)
GPS-SSL: Guided Positive Sampling to Inject Prior Into Self-Supervised Learning
von: Feizi, Aarash, et al.
Veröffentlicht: (2024)
von: Feizi, Aarash, et al.
Veröffentlicht: (2024)
Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
von: Wang, Suyuchen, et al.
Veröffentlicht: (2025)
von: Wang, Suyuchen, et al.
Veröffentlicht: (2025)
StarFlow: Generating Structured Workflow Outputs From Sketch Images
von: Bechard, Patrice, et al.
Veröffentlicht: (2025)
von: Bechard, Patrice, et al.
Veröffentlicht: (2025)
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
von: Gurung, Alexander, et al.
Veröffentlicht: (2026)
von: Gurung, Alexander, et al.
Veröffentlicht: (2026)
Mem-$π$: Adaptive Memory through Learning When and What to Generate
von: Wang, Xiaoqiang, et al.
Veröffentlicht: (2026)
von: Wang, Xiaoqiang, et al.
Veröffentlicht: (2026)
RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content
von: Monteiro, Joao, et al.
Veröffentlicht: (2024)
von: Monteiro, Joao, et al.
Veröffentlicht: (2024)
AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
von: Masry, Ahmed, et al.
Veröffentlicht: (2025)
von: Masry, Ahmed, et al.
Veröffentlicht: (2025)
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
von: Masry, Ahmed, et al.
Veröffentlicht: (2025)
von: Masry, Ahmed, et al.
Veröffentlicht: (2025)
VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing
von: Rodriguez, Juan, et al.
Veröffentlicht: (2026)
von: Rodriguez, Juan, et al.
Veröffentlicht: (2026)
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
von: Masry, Ahmed, et al.
Veröffentlicht: (2025)
von: Masry, Ahmed, et al.
Veröffentlicht: (2025)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
von: Monteiro, João, et al.
Veröffentlicht: (2024)
von: Monteiro, João, et al.
Veröffentlicht: (2024)
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
von: Rodriguez, Juan, et al.
Veröffentlicht: (2024)
von: Rodriguez, Juan, et al.
Veröffentlicht: (2024)
InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation
von: Sahu, Gaurav, et al.
Veröffentlicht: (2024)
von: Sahu, Gaurav, et al.
Veröffentlicht: (2024)
VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
von: Zhang, Tianyu, et al.
Veröffentlicht: (2024)
von: Zhang, Tianyu, et al.
Veröffentlicht: (2024)
Weak Supervision for Real World Graphs
von: Nair, Pratheeksha, et al.
Veröffentlicht: (2025)
von: Nair, Pratheeksha, et al.
Veröffentlicht: (2025)
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA
von: Maheshwary, Rishabh, et al.
Veröffentlicht: (2025)
von: Maheshwary, Rishabh, et al.
Veröffentlicht: (2025)
Scope: Selective Cross-modal Orchestration of Visual Perception Experts
von: Zhang, Tianyu, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyu, et al.
Veröffentlicht: (2025)
Kurtosis-Guided Denoising Score Matching for Tabular Anomaly Detection
von: Livernoche, Victor, et al.
Veröffentlicht: (2026)
von: Livernoche, Victor, et al.
Veröffentlicht: (2026)
Higher Order Transformers: Enhancing Stock Movement Prediction On Multimodal Time-Series Data
von: Omranpour, Soroush, et al.
Veröffentlicht: (2024)
von: Omranpour, Soroush, et al.
Veröffentlicht: (2024)
Unified Game Moderation: Soft-Prompting and LLM-Assisted Label Transfer for Resource-Efficient Toxicity Detection
von: Yang, Zachary, et al.
Veröffentlicht: (2025)
von: Yang, Zachary, et al.
Veröffentlicht: (2025)
Higher-Order Transformers With Kronecker-Structured Attention
von: Omranpour, Soroush, et al.
Veröffentlicht: (2024)
von: Omranpour, Soroush, et al.
Veröffentlicht: (2024)
Benchmarking Vision Language Models for Cultural Understanding
von: Nayak, Shravan, et al.
Veröffentlicht: (2024)
von: Nayak, Shravan, et al.
Veröffentlicht: (2024)
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding
von: Zhang, Le, et al.
Veröffentlicht: (2023)
von: Zhang, Le, et al.
Veröffentlicht: (2023)
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
von: Awal, Rabiul, et al.
Veröffentlicht: (2023)
von: Awal, Rabiul, et al.
Veröffentlicht: (2023)
The Promise of RL for Autoregressive Image Editing
von: Ahmadi, Saba, et al.
Veröffentlicht: (2025)
von: Ahmadi, Saba, et al.
Veröffentlicht: (2025)
FairLoRA: Unpacking Bias Mitigation in Vision Models with Fairness-Driven Low-Rank Adaptation
von: Sukumaran, Rohan, et al.
Veröffentlicht: (2024)
von: Sukumaran, Rohan, et al.
Veröffentlicht: (2024)
BiXSE: Improving Dense Retrieval via Probabilistic Graded Relevance Distillation
von: Tsirigotis, Christos, et al.
Veröffentlicht: (2025)
von: Tsirigotis, Christos, et al.
Veröffentlicht: (2025)
Towards Detecting Contextual Real-Time Toxicity for In-Game Chat
von: Yang, Zachary, et al.
Veröffentlicht: (2023)
von: Yang, Zachary, et al.
Veröffentlicht: (2023)
Grammar Search for Multi-Agent Systems
von: Singh, Mayank, et al.
Veröffentlicht: (2025)
von: Singh, Mayank, et al.
Veröffentlicht: (2025)
VisMin: Visual Minimal-Change Understanding
von: Awal, Rabiul, et al.
Veröffentlicht: (2024)
von: Awal, Rabiul, et al.
Veröffentlicht: (2024)
Combining Confidence Elicitation and Sample-based Methods for Uncertainty Quantification in Misinformation Mitigation
von: Rivera, Mauricio, et al.
Veröffentlicht: (2024)
von: Rivera, Mauricio, et al.
Veröffentlicht: (2024)
Comparing GPT-4 and Open-Source Language Models in Misinformation Mitigation
von: Vergho, Tyler, et al.
Veröffentlicht: (2024)
von: Vergho, Tyler, et al.
Veröffentlicht: (2024)
In value-based deep reinforcement learning, a pruned network is a good network
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
Ask before you Build: Rethinking AI-for-Good in Human Trafficking Interventions
von: Nair, Pratheeksha, et al.
Veröffentlicht: (2025)
von: Nair, Pratheeksha, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
von: Jian, Xiangru, et al.
Veröffentlicht: (2026) -
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
von: Feizi, Aarash, et al.
Veröffentlicht: (2025) -
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
von: Nayak, Shravan, et al.
Veröffentlicht: (2025) -
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
von: Awal, Rabiul, et al.
Veröffentlicht: (2025) -
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
von: Rodriguez, Juan A., et al.
Veröffentlicht: (2025)