Covering Human Action Space for Computer Use: Data Synthesis and Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Miaosen, Zhao, Xiaohan, Tan, Zhihong, Huoshen, Zhou, Fan, Yijia, Yang, Yifan, Qiu, Kai, Liu, Bei, Wagle, Justin, Yin, Chenzhong, Cheng, Mingxi, Li, Ji, Dai, Qi, Luo, Chong, Yang, Xu, Geng, Xin, Guo, Baining |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
by: Zhang, Miaosen, et al.
Published: (2025)
by: Zhang, Miaosen, et al.
Published: (2025)
MageBench: Bridging Large Multimodal Models to Agents
by: Zhang, Miaosen, et al.
Published: (2024)
by: Zhang, Miaosen, et al.
Published: (2024)
Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms
by: Zhang, Miaosen, et al.
Published: (2024)
by: Zhang, Miaosen, et al.
Published: (2024)
Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM Training
by: Zhang, Miaosen, et al.
Published: (2026)
by: Zhang, Miaosen, et al.
Published: (2026)
RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents
by: Zhu, Jialiang, et al.
Published: (2026)
by: Zhu, Jialiang, et al.
Published: (2026)
InfoAgent: Advancing Autonomous Information-Seeking Agents
by: Zhang, Gongrui, et al.
Published: (2025)
by: Zhang, Gongrui, et al.
Published: (2025)
Language-Guided Face Animation by Recurrent StyleGAN-based Generator
by: Hang, Tiankai, et al.
Published: (2022)
by: Hang, Tiankai, et al.
Published: (2022)
MoLoRA: Composable Specialization via Per-Token Adapter Routing
by: Shah, Shrey, et al.
Published: (2026)
by: Shah, Shrey, et al.
Published: (2026)
ScreenSearch: Uncertainty-Aware OS Exploration
by: Solodko, Michael, et al.
Published: (2026)
by: Solodko, Michael, et al.
Published: (2026)
A Review of Brain-Computer Interface Technologies: Signal Acquisition Methods and Interaction Paradigms
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis
by: Cheng, Anzhe, et al.
Published: (2025)
by: Cheng, Anzhe, et al.
Published: (2025)
AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
by: Zhou, Ziwei, et al.
Published: (2026)
by: Zhou, Ziwei, et al.
Published: (2026)
Dielectric Properties of thin Film Al/Sb2Pb1Se7/Al Devices
by: Shaila Wagle
Published: (2000)
by: Shaila Wagle
Published: (2000)
An Advantage-based Optimization Method for Reinforcement Learning in Large Action Space
by: Lin, Hai, et al.
Published: (2024)
by: Lin, Hai, et al.
Published: (2024)
ForestSim: A Synthetic Benchmark for Intelligent Vehicle Perception in Unstructured Forest Environments
by: Wagle, Pragat, et al.
Published: (2026)
by: Wagle, Pragat, et al.
Published: (2026)
Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting
by: Luo, Miaosen, et al.
Published: (2025)
by: Luo, Miaosen, et al.
Published: (2025)
Exploration of Circadian Clock‐Related Genes in the Pathogenesis of Psoriatic Arthritis to Identify Potential Therapeutic Targets From Multi‐Omics Insight: A Mendelian Randomization Study
by: Aimei Liu, et al.
Published: (2025)
by: Aimei Liu, et al.
Published: (2025)
Association Between Heroin Use and Depression: NHANES 2005–2018
by: Bei Li, et al.
Published: (2026)
by: Bei Li, et al.
Published: (2026)
CogPlanner: Unveiling the Potential of Agentic Multimodal Retrieval Augmented Generation with Planning
by: Yu, Xiaohan, et al.
Published: (2025)
by: Yu, Xiaohan, et al.
Published: (2025)
HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models
by: Zhou, Ziqin, et al.
Published: (2025)
by: Zhou, Ziqin, et al.
Published: (2025)
Improved Noise Schedule for Diffusion Training
by: Hang, Tiankai, et al.
Published: (2024)
by: Hang, Tiankai, et al.
Published: (2024)
Fluoroalkoxylating Reagents in Organic Synthesis: Recent Advances
by: Mingxi Chen, et al.
Published: (2024)
by: Mingxi Chen, et al.
Published: (2024)
Instruction Agent: Enhancing Agent with Expert Demonstration
by: Li, Yinheng, et al.
Published: (2025)
by: Li, Yinheng, et al.
Published: (2025)
Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems
by: Jha, Rishi, et al.
Published: (2025)
by: Jha, Rishi, et al.
Published: (2025)
EMoE: Eigenbasis-Guided Routing for Mixture-of-Experts
by: Cheng, Anzhe, et al.
Published: (2026)
by: Cheng, Anzhe, et al.
Published: (2026)
Når musikken gir mening.
by: Wagle Christensen, Torild
Published: (2017)
by: Wagle Christensen, Torild
Published: (2017)
A Guide to Research Collections in Microform in the University of Toronto Library. Reference Series No. 19.
by: Wagle, Iqbal, Comp.
Published: (1974)
by: Wagle, Iqbal, Comp.
Published: (1974)
LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
by: Peng, Yingzhe, et al.
Published: (2025)
by: Peng, Yingzhe, et al.
Published: (2025)
Multimodal machine learning with large language embedding model for polymer property prediction
by: Zhang, Tianren, et al.
Published: (2025)
by: Zhang, Tianren, et al.
Published: (2025)
Benchmarking LLM Tool-Use in the Wild
by: Yu, Peijie, et al.
Published: (2026)
by: Yu, Peijie, et al.
Published: (2026)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
by: Yang, Yifan, et al.
Published: (2026)
by: Yang, Yifan, et al.
Published: (2026)
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
by: Zhao, Baining, et al.
Published: (2025)
by: Zhao, Baining, et al.
Published: (2025)
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning
by: Xu, Ziqiang, et al.
Published: (2025)
by: Xu, Ziqiang, et al.
Published: (2025)
REDUCIO! Generating 1K Video within 16 Seconds using Extremely Compressed Motion Latents
by: Tian, Rui, et al.
Published: (2024)
by: Tian, Rui, et al.
Published: (2024)
HLB: Benchmarking LLMs' Humanlikeness in Language Use
by: Duan, Xufeng, et al.
Published: (2024)
by: Duan, Xufeng, et al.
Published: (2024)
Adaptive Segment-level Reward: Bridging the Gap Between Action and Reward Space in Alignment
by: Li, Yanshi, et al.
Published: (2024)
by: Li, Yanshi, et al.
Published: (2024)
Adaptive Bounded Exploration and Intermediate Actions for Data Debiasing
by: Yang, Yifan, et al.
Published: (2025)
by: Yang, Yifan, et al.
Published: (2025)
PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action
by: Shao, Yijia, et al.
Published: (2024)
by: Shao, Yijia, et al.
Published: (2024)
Similar Items
-
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
by: Zhang, Miaosen, et al.
Published: (2025) -
MageBench: Bridging Large Multimodal Models to Agents
by: Zhang, Miaosen, et al.
Published: (2024) -
Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms
by: Zhang, Miaosen, et al.
Published: (2024) -
Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM Training
by: Zhang, Miaosen, et al.
Published: (2026) -
RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents
by: Zhu, Jialiang, et al.
Published: (2026)