CUA-Skill: Develop Skills for Computer Using Agent
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Tianyi, Li, Yinheng, Solodko, Michael, Wang, Sen, Jiang, Nan, Cui, Tingyuan, Hao, Junheng, Ko, Jongwoo, Abdali, Sara, Xu, Leon, Zheng, Suzhen, Fan, Hao, Cameron, Pashmina, Wagle, Justin, Koishida, Kazuhito |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AppSelectBench: Application-Level Tool Selection Benchmark
by: Chen, Tianyi, et al.
Published: (2025)
by: Chen, Tianyi, et al.
Published: (2025)
Data Generation Using Large Language Models for Text Classification: An Empirical Case Study
by: Li, Yinheng, et al.
Published: (2024)
by: Li, Yinheng, et al.
Published: (2024)
Instruction Agent: Enhancing Agent with Expert Demonstration
by: Li, Yinheng, et al.
Published: (2025)
by: Li, Yinheng, et al.
Published: (2025)
Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
by: Amizadeh, Saeed, et al.
Published: (2025)
by: Amizadeh, Saeed, et al.
Published: (2025)
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
by: Ko, Jongwoo, et al.
Published: (2026)
by: Ko, Jongwoo, et al.
Published: (2026)
Self-reflecting Large Language Models: A Hegelian Dialectical Approach
by: Abdali, Sara, et al.
Published: (2025)
by: Abdali, Sara, et al.
Published: (2025)
Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
by: Bonatti, Rogerio, et al.
Published: (2024)
by: Bonatti, Rogerio, et al.
Published: (2024)
ScreenSearch: Uncertainty-Aware OS Exploration
by: Solodko, Michael, et al.
Published: (2026)
by: Solodko, Michael, et al.
Published: (2026)
WinClick: GUI Grounding with Multimodal Large Language Models
by: Hui, Zheng, et al.
Published: (2025)
by: Hui, Zheng, et al.
Published: (2025)
VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
by: Jang, Lawrence, et al.
Published: (2024)
by: Jang, Lawrence, et al.
Published: (2024)
IntentCUA: Learning Intent-level Representations for Skill Abstraction and Multi-Agent Planning in Computer-Use Agents
by: Lee, Seoyoung, et al.
Published: (2026)
by: Lee, Seoyoung, et al.
Published: (2026)
WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference
by: Chen, Sihan, et al.
Published: (2025)
by: Chen, Sihan, et al.
Published: (2025)
StableQAT: Stable Quantization-Aware Training at Ultra-Low Bitwidths
by: Chen, Tianyi, et al.
Published: (2026)
by: Chen, Tianyi, et al.
Published: (2026)
PRO-CUA: Process-Reward Optimization for Computer Use Agents
by: He, Yifei, et al.
Published: (2026)
by: He, Yifei, et al.
Published: (2026)
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
by: Wang, Bowen, et al.
Published: (2026)
by: Wang, Bowen, et al.
Published: (2026)
Zero-Shot Text-to-Speech from Continuous Text Streams
by: Dang, Trung, et al.
Published: (2024)
by: Dang, Trung, et al.
Published: (2024)
OpenCUA: Open Foundations for Computer-Use Agents
by: Wang, Xinyuan, et al.
Published: (2025)
by: Wang, Xinyuan, et al.
Published: (2025)
Single-channel speech enhancement using learnable loss mixup
by: Chang, Oscar, et al.
Published: (2023)
by: Chang, Oscar, et al.
Published: (2023)
Learned Image Compression with Text Quality Enhancement
by: Lai, Chih-Yu, et al.
Published: (2024)
by: Lai, Chih-Yu, et al.
Published: (2024)
LiteCUA: Computer as MCP Server for Computer-Use Agent on AIOS
by: Mei, Kai, et al.
Published: (2025)
by: Mei, Kai, et al.
Published: (2025)
Weakly-supervised Audio Separation via Bi-modal Semantic Similarity
by: Mahmud, Tanvir, et al.
Published: (2024)
by: Mahmud, Tanvir, et al.
Published: (2024)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
by: Dang, Trung, et al.
Published: (2024)
by: Dang, Trung, et al.
Published: (2024)
EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
by: Ju, Ruofei, et al.
Published: (2026)
by: Ju, Ruofei, et al.
Published: (2026)
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
by: Jian, Xiangru, et al.
Published: (2026)
by: Jian, Xiangru, et al.
Published: (2026)
UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
by: Yang, Yuhao, et al.
Published: (2025)
by: Yang, Yuhao, et al.
Published: (2025)
Organizing, Orchestrating, and Benchmarking Agent Skills at Ecosystem Scale
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures
by: Tabassum, Afrina, et al.
Published: (2024)
by: Tabassum, Afrina, et al.
Published: (2024)
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
by: Bai, Yatong, et al.
Published: (2023)
by: Bai, Yatong, et al.
Published: (2023)
ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data
by: Liu, Zhaoyang, et al.
Published: (2025)
by: Liu, Zhaoyang, et al.
Published: (2025)
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents
by: Hu, Xuhao, et al.
Published: (2026)
by: Hu, Xuhao, et al.
Published: (2026)
A11y-CUA Dataset: Characterizing the Accessibility Gap in Computer Use Agents
by: Mohanbabu, Ananya Gubbi, et al.
Published: (2026)
by: Mohanbabu, Ananya Gubbi, et al.
Published: (2026)
AnySkill: Learning Open-Vocabulary Physical Skill for Interactive Agents
by: Cui, Jieming, et al.
Published: (2024)
by: Cui, Jieming, et al.
Published: (2024)
SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents
by: Zhou, Yifan, et al.
Published: (2026)
by: Zhou, Yifan, et al.
Published: (2026)
SkillScope: Toward Fine-Grained Least-Privilege Enforcement for Agent Skills
by: Wu, Jiangrong, et al.
Published: (2026)
by: Wu, Jiangrong, et al.
Published: (2026)
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
DistiLLM: Towards Streamlined Distillation for Large Language Models
by: Ko, Jongwoo, et al.
Published: (2024)
by: Ko, Jongwoo, et al.
Published: (2024)
SkillGraph: Graph Foundation Priors for LLM Agent Tool Sequence Recommendation
by: Liu, Hao, et al.
Published: (2026)
by: Liu, Hao, et al.
Published: (2026)
EvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experience
by: Xue, Taofeng, et al.
Published: (2026)
by: Xue, Taofeng, et al.
Published: (2026)
How to Augment Language Skills
by: Pym, Anthony, et al.
Published: (2024)
by: Pym, Anthony, et al.
Published: (2024)
MoLoRA: Composable Specialization via Per-Token Adapter Routing
by: Shah, Shrey, et al.
Published: (2026)
by: Shah, Shrey, et al.
Published: (2026)
Similar Items
-
AppSelectBench: Application-Level Tool Selection Benchmark
by: Chen, Tianyi, et al.
Published: (2025) -
Data Generation Using Large Language Models for Text Classification: An Empirical Case Study
by: Li, Yinheng, et al.
Published: (2024) -
Instruction Agent: Enhancing Agent with Expert Demonstration
by: Li, Yinheng, et al.
Published: (2025) -
Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
by: Amizadeh, Saeed, et al.
Published: (2025) -
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
by: Ko, Jongwoo, et al.
Published: (2026)