OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Xuetian, Chen, Yinghao, Yuan, Xinfeng, Peng, Zhuo, Chen, Lu, Li, Yuekeng, Zhang, Zhoujia, Huang, Yingqian, Huang, Leyan, Liang, Jiaqing, Xie, Tianbao, Wu, Zhiyong, Sun, Qiushi, Qi, Biqing, Zhou, Bowen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
von: Lu, Yi, et al.
Veröffentlicht: (2025)
von: Lu, Yi, et al.
Veröffentlicht: (2025)
How Far Can In-Context Alignment Go? Exploring the State of In-Context Alignment
von: Huang, Heyan, et al.
Veröffentlicht: (2024)
von: Huang, Heyan, et al.
Veröffentlicht: (2024)
Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?
von: Li, Bangzheng, et al.
Veröffentlicht: (2023)
von: Li, Bangzheng, et al.
Veröffentlicht: (2023)
LLM Compression: How Far Can We Go in Balancing Size and Performance?
von: Sk, Sahil, et al.
Veröffentlicht: (2025)
von: Sk, Sahil, et al.
Veröffentlicht: (2025)
Financial Named Entity Recognition: How Far Can LLM Go?
von: Lu, Yi-Te, et al.
Veröffentlicht: (2025)
von: Lu, Yi-Te, et al.
Veröffentlicht: (2025)
Compute Allocation in Evolutionary Search: From Depth-Breadth to Multi-Armed Bandits
von: Xing, Sixue, et al.
Veröffentlicht: (2026)
von: Xing, Sixue, et al.
Veröffentlicht: (2026)
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment
von: Wang, Sizhe, et al.
Veröffentlicht: (2024)
von: Wang, Sizhe, et al.
Veröffentlicht: (2024)
Dual Engines of Thoughts: A Depth-Breadth Integration Framework for Open-Ended Analysis
von: Yu, Fei-Hsuan, et al.
Veröffentlicht: (2025)
von: Yu, Fei-Hsuan, et al.
Veröffentlicht: (2025)
In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2025)
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2025)
Runtime Failure Hunting for Physics Engine Based Software Systems: How Far Can We Go?
von: Li, Shuqing, et al.
Veröffentlicht: (2025)
von: Li, Shuqing, et al.
Veröffentlicht: (2025)
The Narrow Depth and Breadth of Corporate Responsible AI Research
von: Ahmed, Nur, et al.
Veröffentlicht: (2024)
von: Ahmed, Nur, et al.
Veröffentlicht: (2024)
D$^2$HScore: Reasoning-Aware Hallucination Detection via Semantic Breadth and Depth Analysis in LLMs
von: Ding, Yue, et al.
Veröffentlicht: (2025)
von: Ding, Yue, et al.
Veröffentlicht: (2025)
How Far Can Unsupervised RLVR Scale LLM Training?
von: He, Bingxiang, et al.
Veröffentlicht: (2026)
von: He, Bingxiang, et al.
Veröffentlicht: (2026)
How Far Can We Compress Instant-NGP-Based NeRF?
von: Chen, Yihang, et al.
Veröffentlicht: (2024)
von: Chen, Yihang, et al.
Veröffentlicht: (2024)
Depth-Copy-Paste: Multimodal and Depth-Aware Compositing for Robust Face Detection
von: Guo, Qiushi
Veröffentlicht: (2025)
von: Guo, Qiushi
Veröffentlicht: (2025)
AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go?
von: Karmakar, Ranit, et al.
Veröffentlicht: (2026)
von: Karmakar, Ranit, et al.
Veröffentlicht: (2026)
How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering?
von: Lando, Giuseppe, et al.
Veröffentlicht: (2025)
von: Lando, Giuseppe, et al.
Veröffentlicht: (2025)
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
von: Lyu, Zhiheng, et al.
Veröffentlicht: (2025)
von: Lyu, Zhiheng, et al.
Veröffentlicht: (2025)
Towards Building Specialized Generalist AI with System 1 and System 2 Fusion
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2024)
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2024)
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
von: Wu, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wu, Zhiyong, et al.
Veröffentlicht: (2024)
Growing a Neural Network in Breadth, Depth, and Time
von: Butkus, Eivinas, et al.
Veröffentlicht: (2026)
von: Butkus, Eivinas, et al.
Veröffentlicht: (2026)
RETUYT-INCO at BEA 2025 Shared Task: How Far Can Lightweight Models Go in AI-powered Tutor Evaluation?
von: Góngora, Santiago, et al.
Veröffentlicht: (2025)
von: Góngora, Santiago, et al.
Veröffentlicht: (2025)
Implying Volatility: How Fast Can We Go?
von: Floc'h, Fabien Le, et al.
Veröffentlicht: (2026)
von: Floc'h, Fabien Le, et al.
Veröffentlicht: (2026)
How Far Can 100 Samples Go? Unlocking Overall Zero-Shot Multilingual Translation via Tiny Multi-Parallel Data
von: Wu, Di, et al.
Veröffentlicht: (2024)
von: Wu, Di, et al.
Veröffentlicht: (2024)
PaD: Program-aided Distillation Can Teach Small Models Reasoning Better than Chain-of-thought Fine-tuning
von: Zhu, Xuekai, et al.
Veröffentlicht: (2023)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2023)
SAM-PD: How Far Can SAM Take Us in Tracking and Segmenting Anything in Videos by Prompt Denoising
von: Zhou, Tao, et al.
Veröffentlicht: (2024)
von: Zhou, Tao, et al.
Veröffentlicht: (2024)
KGHaluBench: A Knowledge Graph-Based Hallucination Benchmark for Evaluating the Breadth and Depth of LLM Knowledge
von: Robertson, Alex, et al.
Veröffentlicht: (2026)
von: Robertson, Alex, et al.
Veröffentlicht: (2026)
How Far are LLMs from Being Our Digital Twins? A Benchmark for Persona-Based Behavior Chain Simulation
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
How You Can Protect Public Access Computers "and" Their Users
von: Huang, Phil
Veröffentlicht: (2007)
von: Huang, Phil
Veröffentlicht: (2007)
ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?
von: Han, Haonan, et al.
Veröffentlicht: (2026)
von: Han, Haonan, et al.
Veröffentlicht: (2026)
Delving into the Reversal Curse: How Far Can Large Language Models Generalize?
von: Lin, Zhengkai, et al.
Veröffentlicht: (2024)
von: Lin, Zhengkai, et al.
Veröffentlicht: (2024)
Large Language Models for Predictive Analysis: How Far Are They?
von: Chen, Qin, et al.
Veröffentlicht: (2025)
von: Chen, Qin, et al.
Veröffentlicht: (2025)
Can Pre-trained Language Models Understand Chinese Humor?
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
How Far Can VLMs Go for Visual Bug Detection? Studying 19,738 Keyframes from 41 Hours of Gameplay Videos
von: Lu, Wentao, et al.
Veröffentlicht: (2026)
von: Lu, Wentao, et al.
Veröffentlicht: (2026)
IDEA-Bench: How Far are Generative Models from Professional Designing?
von: Liang, Chen, et al.
Veröffentlicht: (2024)
von: Liang, Chen, et al.
Veröffentlicht: (2024)
How Far Can We Extract Diverse Perspectives from Large Language Models?
von: Hayati, Shirley Anugrah, et al.
Veröffentlicht: (2023)
von: Hayati, Shirley Anugrah, et al.
Veröffentlicht: (2023)
Vulnerability Detection with Code Language Models: How Far Are We?
von: Ding, Yangruibo, et al.
Veröffentlicht: (2024)
von: Ding, Yangruibo, et al.
Veröffentlicht: (2024)
As Far as Eye See: Vergence-Pupil Coupling in Near-Far Depth Switching
von: Maquiling, Virmarie, et al.
Veröffentlicht: (2026)
von: Maquiling, Virmarie, et al.
Veröffentlicht: (2026)
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
von: Liu, Runze, et al.
Veröffentlicht: (2025)
von: Liu, Runze, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
von: Lu, Yi, et al.
Veröffentlicht: (2025) -
How Far Can In-Context Alignment Go? Exploring the State of In-Context Alignment
von: Huang, Heyan, et al.
Veröffentlicht: (2024) -
Deceptive Semantic Shortcuts on Reasoning Chains: How Far Can Models Go without Hallucination?
von: Li, Bangzheng, et al.
Veröffentlicht: (2023) -
LLM Compression: How Far Can We Go in Balancing Size and Performance?
von: Sk, Sahil, et al.
Veröffentlicht: (2025) -
Financial Named Entity Recognition: How Far Can LLM Go?
von: Lu, Yi-Te, et al.
Veröffentlicht: (2025)