OpenHands: An Open Platform for AI Software Developers as Generalist Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xingyao, Li, Boxuan, Song, Yufan, Xu, Frank F., Tang, Xiangru, Zhuge, Mingchen, Pan, Jiayi, Song, Yueqi, Li, Bowen, Singh, Jaskirat, Tran, Hoang H., Li, Fuqiang, Ma, Ren, Zheng, Mingzhang, Qian, Bill, Shao, Yanjun, Muennighoff, Niklas, Zhang, Yizhe, Hui, Binyuan, Lin, Junyang, Brennan, Robert, Peng, Hao, Ji, Heng, Neubig, Graham |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
by: Wang, Xingyao, et al.
Published: (2025)
by: Wang, Xingyao, et al.
Published: (2025)
OpenHands/software-agent-sdk: v1.21.0
by: Xingyao Wang, et al.
Published: (2026)
by: Xingyao Wang, et al.
Published: (2026)
OpenHands/software-agent-sdk: v1.19.1
by: Xingyao Wang, et al.
Published: (2026)
by: Xingyao Wang, et al.
Published: (2026)
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
by: Soni, Aditya Bharat, et al.
Published: (2025)
by: Soni, Aditya Bharat, et al.
Published: (2025)
What Is Missing in Multilingual Visual Reasoning and How to Fix It
by: Song, Yueqi, et al.
Published: (2024)
by: Song, Yueqi, et al.
Published: (2024)
Beyond Browsing: API-Based Web Agents
by: Song, Yueqi, et al.
Published: (2024)
by: Song, Yueqi, et al.
Published: (2024)
An image speaks a thousand words, but can everyone listen? On image transcreation for cultural relevance
by: Khanuja, Simran, et al.
Published: (2024)
by: Khanuja, Simran, et al.
Published: (2024)
Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds
by: Tan, Weihao, et al.
Published: (2025)
by: Tan, Weihao, et al.
Published: (2025)
VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge
by: Song, Yueqi, et al.
Published: (2025)
by: Song, Yueqi, et al.
Published: (2025)
TableLlama: Towards Open Large Generalist Models for Tables
by: Zhang, Tianshu, et al.
Published: (2023)
by: Zhang, Tianshu, et al.
Published: (2023)
Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
by: Yue, Xiang, et al.
Published: (2024)
by: Yue, Xiang, et al.
Published: (2024)
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
by: Nyandwi, Jean de Dieu, et al.
Published: (2025)
by: Nyandwi, Jean de Dieu, et al.
Published: (2025)
OctoPack: Instruction Tuning Code Large Language Models
by: Muennighoff, Niklas, et al.
Published: (2023)
by: Muennighoff, Niklas, et al.
Published: (2023)
Training Software Engineering Agents and Verifiers with SWE-Gym
by: Pan, Jiayi, et al.
Published: (2024)
by: Pan, Jiayi, et al.
Published: (2024)
A Rubric-Supervised Critic from Sparse Real-World Outcomes
by: Wang, Xingyao, et al.
Published: (2026)
by: Wang, Xingyao, et al.
Published: (2026)
Octo: An Open-Source Generalist Robot Policy
by: Octo Model Team, et al.
Published: (2024)
by: Octo Model Team, et al.
Published: (2024)
Open Government Data and Corporate Tax Avoidance: Evidence From Listed Companies of China
by: Hua Wang, et al.
Published: (2025)
by: Hua Wang, et al.
Published: (2025)
Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation
by: Yan, Haodong, et al.
Published: (2025)
by: Yan, Haodong, et al.
Published: (2025)
Trajectories of Depressive Symptom Among College Students in China During the COVID‐19 Pandemic: Association With Suicidal Ideation and Insomnia Symptoms
by: Binyuan Su, et al.
Published: (2025)
by: Binyuan Su, et al.
Published: (2025)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
by: Lei, Zhenxin, et al.
Published: (2025)
by: Lei, Zhenxin, et al.
Published: (2025)
Humanline: Online Alignment as Perceptual Loss
by: Liu, Sijia, et al.
Published: (2025)
by: Liu, Sijia, et al.
Published: (2025)
Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation
by: Ye, Chengyang, et al.
Published: (2024)
by: Ye, Chengyang, et al.
Published: (2024)
Copilot-Assisted Second-Thought Framework for Brain-to-Robot Hand Motion Decoding
by: Li, Yizhe, et al.
Published: (2026)
by: Li, Yizhe, et al.
Published: (2026)
Continual Hand-Eye Calibration for Open-world Robotic Manipulation
by: Li, Fazeng, et al.
Published: (2026)
by: Li, Fazeng, et al.
Published: (2026)
VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
by: Liu, Junpeng, et al.
Published: (2024)
by: Liu, Junpeng, et al.
Published: (2024)
UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding
by: Shi, Bowen, et al.
Published: (2024)
by: Shi, Bowen, et al.
Published: (2024)
NitroGen: An Open Foundation Model for Generalist Gaming Agents
by: Magne, Loïc, et al.
Published: (2026)
by: Magne, Loïc, et al.
Published: (2026)
OpenCAEPoro: A Parallel Simulation Framework for Multiphase and Multicomponent Porous Media Flows
by: Li, Shizhe, et al.
Published: (2024)
by: Li, Shizhe, et al.
Published: (2024)
NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
by: Hung, Chia-Yu, et al.
Published: (2025)
by: Hung, Chia-Yu, et al.
Published: (2025)
OpenChat: Advancing Open-source Language Models with Mixed-Quality Data
by: Wang, Guan, et al.
Published: (2023)
by: Wang, Guan, et al.
Published: (2023)
mmEgoHand: Egocentric Hand Pose Estimation and Gesture Recognition with Head-mounted Millimeter-wave Radar and IMU
by: Lv, Yizhe, et al.
Published: (2025)
by: Lv, Yizhe, et al.
Published: (2025)
R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
by: Jain, Naman, et al.
Published: (2025)
by: Jain, Naman, et al.
Published: (2025)
UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
by: Liang, Zhengyang, et al.
Published: (2025)
by: Liang, Zhengyang, et al.
Published: (2025)
All Beings Are Equal in Open Set Recognition
by: Li, Chaohua, et al.
Published: (2024)
by: Li, Chaohua, et al.
Published: (2024)
Finite-time blow-up in a quasilinear two-species chemotaxis system with two chemicals
by: Cai, Mingzhang, et al.
Published: (2026)
by: Cai, Mingzhang, et al.
Published: (2026)
OLMoE: Open Mixture-of-Experts Language Models
by: Muennighoff, Niklas, et al.
Published: (2024)
by: Muennighoff, Niklas, et al.
Published: (2024)
Curing with Code: The Intersection of AI, ML, and Medical Science
by: Jaskirat Singh Chawla
Published: (2025)
by: Jaskirat Singh Chawla
Published: (2025)
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Cognitive Kernel: An Open-source Agent System towards Generalist Autopilots
by: Zhang, Hongming, et al.
Published: (2024)
by: Zhang, Hongming, et al.
Published: (2024)
Hodge-theoretic Open/Closed Correspondence
by: Yu, Song
Published: (2025)
by: Yu, Song
Published: (2025)
Similar Items
-
The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
by: Wang, Xingyao, et al.
Published: (2025) -
OpenHands/software-agent-sdk: v1.21.0
by: Xingyao Wang, et al.
Published: (2026) -
OpenHands/software-agent-sdk: v1.19.1
by: Xingyao Wang, et al.
Published: (2026) -
Coding Agents with Multimodal Browsing are Generalist Problem Solvers
by: Soni, Aditya Bharat, et al.
Published: (2025) -
What Is Missing in Multilingual Visual Reasoning and How to Fix It
by: Song, Yueqi, et al.
Published: (2024)