NeuralOS: Towards Simulating Operating Systems via Neural Generative Models
Fuente:
arXiv
Saved in:
| Main Authors: | Rivard, Luke, Sun, Sun, Guo, Hongyu, Chen, Wenhu, Deng, Yuntian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
by: Wu, Zhiyong, et al.
Published: (2024)
by: Wu, Zhiyong, et al.
Published: (2024)
OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
by: Sun, Qiushi, et al.
Published: (2025)
by: Sun, Qiushi, et al.
Published: (2025)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
by: Wang, Siting, et al.
Published: (2025)
by: Wang, Siting, et al.
Published: (2025)
AURORA: Navigating UI Tarpits via Automated Neural Screen Understanding
by: Khan, Safwat Ali, et al.
Published: (2024)
by: Khan, Safwat Ali, et al.
Published: (2024)
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
by: Sun, Qiushi, et al.
Published: (2024)
by: Sun, Qiushi, et al.
Published: (2024)
OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
by: Chen, Xuetian, et al.
Published: (2025)
by: Chen, Xuetian, et al.
Published: (2025)
Computer-Use Agents as Judges for Generative User Interface
by: Lin, Kevin Qinghong, et al.
Published: (2025)
by: Lin, Kevin Qinghong, et al.
Published: (2025)
Visual Neural Decoding via Improved Visual-EEG Semantic Consistency
by: Chen, Hongzhou, et al.
Published: (2024)
by: Chen, Hongzhou, et al.
Published: (2024)
Long-Term Ad Memorability: Understanding & Generating Memorable Ads
by: SI, Harini, et al.
Published: (2023)
by: SI, Harini, et al.
Published: (2023)
GUICourse: From General Vision Language Models to Versatile GUI Agents
by: Chen, Wentong, et al.
Published: (2024)
by: Chen, Wentong, et al.
Published: (2024)
EvoDiagram: Agentic Editable Diagram Creation via Design Expertise Evolution
by: Wang, Tianfu, et al.
Published: (2026)
by: Wang, Tianfu, et al.
Published: (2026)
True (VIS) Lies: Analyzing How Generative AI Recognizes Intentionality, Rhetoric, and Misleadingness in Visualization Lies
by: Blasilli, Graziano, et al.
Published: (2026)
by: Blasilli, Graziano, et al.
Published: (2026)
Towards Interactive Intelligence for Digital Humans
by: Cai, Yiyi, et al.
Published: (2025)
by: Cai, Yiyi, et al.
Published: (2025)
Ninja Codes: Neurally Generated Fiducial Markers for Stealthy 6-DoF Tracking
by: Takeuchi, Yuichiro, et al.
Published: (2025)
by: Takeuchi, Yuichiro, et al.
Published: (2025)
What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metric
by: Kerkouri, Mohamed Amine, et al.
Published: (2026)
by: Kerkouri, Mohamed Amine, et al.
Published: (2026)
A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation
by: Betala, Siddharth, et al.
Published: (2025)
by: Betala, Siddharth, et al.
Published: (2025)
A Review on Large Language Models for Visual Analytics
by: Agarwal, Navya Sonal, et al.
Published: (2025)
by: Agarwal, Navya Sonal, et al.
Published: (2025)
Code2World: A GUI World Model via Renderable Code Generation
by: Zheng, Yuhao, et al.
Published: (2026)
by: Zheng, Yuhao, et al.
Published: (2026)
UIClip: A Data-driven Model for Assessing User Interface Design
by: Wu, Jason, et al.
Published: (2024)
by: Wu, Jason, et al.
Published: (2024)
GPT-5 Model Corrected GPT-4V's Chart Reading Errors, Not Prompting
by: Yang, Kaichun, et al.
Published: (2025)
by: Yang, Kaichun, et al.
Published: (2025)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
by: Verma, Arnav, et al.
Published: (2025)
by: Verma, Arnav, et al.
Published: (2025)
AppCopilot: Toward General, Accurate, Long-Horizon, and Efficient Mobile Agent
by: Fan, Jingru, et al.
Published: (2025)
by: Fan, Jingru, et al.
Published: (2025)
GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
Deep Neural Encoder-Decoder Model to Relate fMRI Brain Activity with Naturalistic Stimuli
by: David, Florian, et al.
Published: (2025)
by: David, Florian, et al.
Published: (2025)
LLM4Brain: Training a Large Language Model for Brain Video Understanding
by: Zheng, Ruizhe, et al.
Published: (2024)
by: Zheng, Ruizhe, et al.
Published: (2024)
Resource-Efficient Gesture Recognition using Low-Resolution Thermal Camera via Spiking Neural Networks and Sparse Segmentation
by: Safa, Ali, et al.
Published: (2024)
by: Safa, Ali, et al.
Published: (2024)
Learning 6-DoF Fine-grained Grasp Detection Based on Part Affordance Grounding
by: Song, Yaoxian, et al.
Published: (2023)
by: Song, Yaoxian, et al.
Published: (2023)
Learning Multimodal Cues of Children's Uncertainty
by: Cheng, Qi, et al.
Published: (2024)
by: Cheng, Qi, et al.
Published: (2024)
GesGPT: Speech Gesture Synthesis With Text Parsing from ChatGPT
by: Gao, Nan, et al.
Published: (2023)
by: Gao, Nan, et al.
Published: (2023)
Morae: Proactively Pausing UI Agents for User Choices
by: Peng, Yi-Hao, et al.
Published: (2025)
by: Peng, Yi-Hao, et al.
Published: (2025)
What Color Scheme is More Effective in Assisting Readers to Locate Information in a Color-Coded Article?
by: Ng, Ho Yin, et al.
Published: (2024)
by: Ng, Ho Yin, et al.
Published: (2024)
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
by: Liu, Xinyi, et al.
Published: (2025)
by: Liu, Xinyi, et al.
Published: (2025)
Deciphering Emotions in Children Storybooks: A Comparative Analysis of Multimodal LLMs in Educational Applications
by: Asseri, Bushra, et al.
Published: (2025)
by: Asseri, Bushra, et al.
Published: (2025)
ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots
by: Hsiao, Yu-Chung, et al.
Published: (2022)
by: Hsiao, Yu-Chung, et al.
Published: (2022)
Fool Me Once? Contrasting Textual and Visual Explanations in a Clinical Decision-Support Setting
by: Kayser, Maxime, et al.
Published: (2024)
by: Kayser, Maxime, et al.
Published: (2024)
VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
by: Mazumdar, Amrita, et al.
Published: (2026)
by: Mazumdar, Amrita, et al.
Published: (2026)
Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
by: You, Keen, et al.
Published: (2024)
by: You, Keen, et al.
Published: (2024)
SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos
by: Huang, Xiyang, et al.
Published: (2026)
by: Huang, Xiyang, et al.
Published: (2026)
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
by: Xie, Tianbao, et al.
Published: (2025)
by: Xie, Tianbao, et al.
Published: (2025)
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
by: Zhao, Henry Hengyuan, et al.
Published: (2025)
by: Zhao, Henry Hengyuan, et al.
Published: (2025)
Similar Items
-
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
by: Wu, Zhiyong, et al.
Published: (2024) -
OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
by: Sun, Qiushi, et al.
Published: (2025) -
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
by: Wang, Siting, et al.
Published: (2025) -
AURORA: Navigating UI Tarpits via Automated Neural Screen Understanding
by: Khan, Safwat Ali, et al.
Published: (2024) -
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
by: Sun, Qiushi, et al.
Published: (2024)