Benchmarking Mobile Device Control Agents across Diverse Configurations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Juyong, Min, Taywon, An, Minyong, Hahm, Dongyoon, Lee, Haeone, Kim, Changyeon, Lee, Kimin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding Impact of Human Feedback via Influence Functions
von: Min, Taywon, et al.
Veröffentlicht: (2025)
von: Min, Taywon, et al.
Veröffentlicht: (2025)
State Your Intention to Steer Your Attention: An AI Assistant for Intentional Digital Living
von: Choi, Juheon, et al.
Veröffentlicht: (2025)
von: Choi, Juheon, et al.
Veröffentlicht: (2025)
Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2025)
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2025)
DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
von: Kim, Changyeon, et al.
Veröffentlicht: (2025)
von: Kim, Changyeon, et al.
Veröffentlicht: (2025)
MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control
von: Lee, Juyong, et al.
Veröffentlicht: (2024)
von: Lee, Juyong, et al.
Veröffentlicht: (2024)
From Accuracy to Readiness: Metrics and Benchmarks for Human-AI Decision-Making
von: Lee, Min Hun
Veröffentlicht: (2026)
von: Lee, Min Hun
Veröffentlicht: (2026)
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2026)
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2026)
Enhancing LLM Agent Safety via Causal Influence Prompting
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2025)
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2025)
AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents
von: Kim, Jeonghyeon, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghyeon, et al.
Veröffentlicht: (2026)
Towards Uncertainty Aware Task Delegation and Human-AI Collaborative Decision-Making
von: Lee, Min Hun, et al.
Veröffentlicht: (2025)
von: Lee, Min Hun, et al.
Veröffentlicht: (2025)
RuleEdit: Failure-Guided Human-AI Model Editing with Prospective Impact Preview
von: Lee, Min Hun, et al.
Veröffentlicht: (2026)
von: Lee, Min Hun, et al.
Veröffentlicht: (2026)
Investigating an Intelligent System to Monitor \& Explain Abnormal Activity Patterns of Older Adults
von: Lee, Min Hun, et al.
Veröffentlicht: (2025)
von: Lee, Min Hun, et al.
Veröffentlicht: (2025)
Mobile Fitting Room: On-device Virtual Try-on via Diffusion Models
von: Blalock, Justin, et al.
Veröffentlicht: (2024)
von: Blalock, Justin, et al.
Veröffentlicht: (2024)
Catch Me if You Search: When Contextual Web Search Results Affect the Detection of Hallucinations
von: Nahar, Mahjabin, et al.
Veröffentlicht: (2025)
von: Nahar, Mahjabin, et al.
Veröffentlicht: (2025)
Towards Interactive Reinforcement Learning with Intrinsic Feedback
von: Poole, Benjamin, et al.
Veröffentlicht: (2021)
von: Poole, Benjamin, et al.
Veröffentlicht: (2021)
Improving Health Professionals' Onboarding with AI and XAI for Trustworthy Human-AI Collaborative Decision Making
von: Lee, Min Hun, et al.
Veröffentlicht: (2024)
von: Lee, Min Hun, et al.
Veröffentlicht: (2024)
Interactive Example-based Explanations to Improve Health Professionals' Onboarding with AI for Human-AI Collaborative Decision Making
von: Lee, Min Hun, et al.
Veröffentlicht: (2024)
von: Lee, Min Hun, et al.
Veröffentlicht: (2024)
HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
DigiData: Training and Evaluating General-Purpose Mobile Control Agents
von: Sun, Yuxuan, et al.
Veröffentlicht: (2025)
von: Sun, Yuxuan, et al.
Veröffentlicht: (2025)
Improving Dialogue Agents by Decomposing One Global Explicit Annotation with Local Implicit Multimodal Feedback
von: Lee, Dong Won, et al.
Veröffentlicht: (2024)
von: Lee, Dong Won, et al.
Veröffentlicht: (2024)
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
von: Zhang, Lechen, et al.
Veröffentlicht: (2025)
von: Zhang, Lechen, et al.
Veröffentlicht: (2025)
AI Agents for Inventory Control: Human-LLM-OR Complementarity
von: Baek, Jackie, et al.
Veröffentlicht: (2026)
von: Baek, Jackie, et al.
Veröffentlicht: (2026)
MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents
von: Wang, Luyuan, et al.
Veröffentlicht: (2024)
von: Wang, Luyuan, et al.
Veröffentlicht: (2024)
Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human Feedback
von: Yuan, Yifu, et al.
Veröffentlicht: (2024)
von: Yuan, Yifu, et al.
Veröffentlicht: (2024)
Predicting Trust In Autonomous Vehicles: Modeling Young Adult Psychosocial Traits, Risk-Benefit Attitudes, And Driving Factors With Machine Learning
von: Kaufman, Robert, et al.
Veröffentlicht: (2024)
von: Kaufman, Robert, et al.
Veröffentlicht: (2024)
TalkToAgent: A Human-centric Explanation of Reinforcement Learning Agents with Large Language Models
von: Kim, Haechang, et al.
Veröffentlicht: (2025)
von: Kim, Haechang, et al.
Veröffentlicht: (2025)
NeuroNet: A Novel Hybrid Self-Supervised Learning Framework for Sleep Stage Classification Using Single-Channel EEG
von: Lee, Cheol-Hui, et al.
Veröffentlicht: (2024)
von: Lee, Cheol-Hui, et al.
Veröffentlicht: (2024)
U-Define: Designing User Workflows for Hard and Soft Constraints in LLM-Based Planning
von: Lee, Christine P, et al.
Veröffentlicht: (2026)
von: Lee, Christine P, et al.
Veröffentlicht: (2026)
When Generative Artificial Intelligence meets Extended Reality: A Systematic Review
von: Ning, Xinyu, et al.
Veröffentlicht: (2025)
von: Ning, Xinyu, et al.
Veröffentlicht: (2025)
Enabling On-Device LLMs Personalization with Smartphone Sensing
von: Zhang, Shiquan, et al.
Veröffentlicht: (2024)
von: Zhang, Shiquan, et al.
Veröffentlicht: (2024)
Optimizing Data Delivery: Insights from User Preferences on Visuals, Tables, and Text
von: Luera, Reuben, et al.
Veröffentlicht: (2024)
von: Luera, Reuben, et al.
Veröffentlicht: (2024)
Rethinking Generalized BCIs: Benchmarking 340,000+ Unique Algorithmic Configurations for EEG Mental Command Decoding
von: Barbaste, Paul, et al.
Veröffentlicht: (2025)
von: Barbaste, Paul, et al.
Veröffentlicht: (2025)
Advancing Brainwave Modeling with a Codebook-Based Foundation Model
von: Barmpas, Konstantinos, et al.
Veröffentlicht: (2025)
von: Barmpas, Konstantinos, et al.
Veröffentlicht: (2025)
Dataset Refinement for Improving the Generalization Ability of the EEG Decoding Model
von: Kim, Sung-Jin, et al.
Veröffentlicht: (2024)
von: Kim, Sung-Jin, et al.
Veröffentlicht: (2024)
Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback
von: Yang, Yongjin, et al.
Veröffentlicht: (2025)
von: Yang, Yongjin, et al.
Veröffentlicht: (2025)
Toward Human-AI Complementarity Across Diverse Tasks
von: Xu, Yuzheng, et al.
Veröffentlicht: (2026)
von: Xu, Yuzheng, et al.
Veröffentlicht: (2026)
DiscoverLLM: From Executing Intents to Discovering Them
von: Kim, Tae Soo, et al.
Veröffentlicht: (2026)
von: Kim, Tae Soo, et al.
Veröffentlicht: (2026)
A Resilient Solution for Sewer Overflow Monitoring across Cloud and Edge
von: Singh, Vipin, et al.
Veröffentlicht: (2026)
von: Singh, Vipin, et al.
Veröffentlicht: (2026)
Benchmark It Yourself (BIY): Preparing a Dataset and Benchmarking AI Models for Scatterplot-Related Tasks
von: Palmeiro, João, et al.
Veröffentlicht: (2025)
von: Palmeiro, João, et al.
Veröffentlicht: (2025)
Quality over Quantity: Demonstration Curation via Influence Functions for Data-Centric Robot Learning
von: Lee, Haeone, et al.
Veröffentlicht: (2026)
von: Lee, Haeone, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Understanding Impact of Human Feedback via Influence Functions
von: Min, Taywon, et al.
Veröffentlicht: (2025) -
State Your Intention to Steer Your Attention: An AI Assistant for Intentional Digital Living
von: Choi, Juheon, et al.
Veröffentlicht: (2025) -
Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2025) -
DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
von: Kim, Changyeon, et al.
Veröffentlicht: (2025) -
MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control
von: Lee, Juyong, et al.
Veröffentlicht: (2024)