NitroGen: An Open Foundation Model for Generalist Gaming Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Magne, Loïc, Awadalla, Anas, Wang, Guanzhi, Xu, Yinzhen, Belofsky, Joshua, Hu, Fengyuan, Kim, Joohwan, Schmidt, Ludwig, Gkioxari, Georgia, Kautz, Jan, Yue, Yisong, Choi, Yejin, Zhu, Yuke, Fan, Linxi "Jim" |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TOTEM: TOkenized Time Series EMbeddings for General Time Series Analysis
by: Talukder, Sabera, et al.
Published: (2024)
by: Talukder, Sabera, et al.
Published: (2024)
Find Any Part in 3D
by: Ma, Ziqi, et al.
Published: (2024)
by: Ma, Ziqi, et al.
Published: (2024)
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Visual Agentic AI for Spatial Reasoning with a Dynamic API
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
Feedforward 3D Editing via Text-Steerable Image-to-3D
by: Ma, Ziqi, et al.
Published: (2025)
by: Ma, Ziqi, et al.
Published: (2025)
DreamGen: Unlocking Generalization in Robot Learning through Video World Models
by: Jang, Joel, et al.
Published: (2025)
by: Jang, Joel, et al.
Published: (2025)
FLARE: Robot Learning with Implicit World Modeling
by: Zheng, Ruijie, et al.
Published: (2025)
by: Zheng, Ruijie, et al.
Published: (2025)
No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision
by: Sahoo, Aadarsh, et al.
Published: (2026)
by: Sahoo, Aadarsh, et al.
Published: (2026)
Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models
by: Ma, Ziqi, et al.
Published: (2026)
by: Ma, Ziqi, et al.
Published: (2026)
Aligning Text, Images, and 3D Structure Token-by-Token
by: Sahoo, Aadarsh, et al.
Published: (2025)
by: Sahoo, Aadarsh, et al.
Published: (2025)
Is This Tracker On? A Benchmark Protocol for Dynamic Tracking
by: Demler, Ilona, et al.
Published: (2025)
by: Demler, Ilona, et al.
Published: (2025)
MINT-1T: Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
by: Awadalla, Anas, et al.
Published: (2024)
by: Awadalla, Anas, et al.
Published: (2024)
AMAGO: Scalable In-Context Reinforcement Learning for Adaptive Agents
by: Grigsby, Jake, et al.
Published: (2023)
by: Grigsby, Jake, et al.
Published: (2023)
Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse
by: Zhang, Kuan, et al.
Published: (2026)
by: Zhang, Kuan, et al.
Published: (2026)
Reconstructing Hand-Held Objects in 3D from Images and Videos
by: Wu, Jane, et al.
Published: (2024)
by: Wu, Jane, et al.
Published: (2024)
Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
by: Kang, Raphi, et al.
Published: (2026)
by: Kang, Raphi, et al.
Published: (2026)
Is CLIP ideal? No. Can we fix it? Yes!
by: Kang, Raphi, et al.
Published: (2025)
by: Kang, Raphi, et al.
Published: (2025)
Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
Eureka: Human-Level Reward Design via Coding Large Language Models
by: Ma, Yecheng Jason, et al.
Published: (2023)
by: Ma, Yecheng Jason, et al.
Published: (2023)
DexMimicGen: Automated Data Generation for Bimanual Dexterous Manipulation via Imitation Learning
by: Jiang, Zhenyu, et al.
Published: (2024)
by: Jiang, Zhenyu, et al.
Published: (2024)
Same or Not? Enhancing Visual Perception in Vision-Language Models
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
by: Gao, Shenyuan, et al.
Published: (2026)
by: Gao, Shenyuan, et al.
Published: (2026)
From Generalist to Specialist Representation
by: Zheng, Yujia, et al.
Published: (2026)
by: Zheng, Yujia, et al.
Published: (2026)
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
by: Hatamizadeh, Ali, et al.
Published: (2026)
by: Hatamizadeh, Ali, et al.
Published: (2026)
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
by: Chandu, Khyathi Raghavi, et al.
Published: (2024)
by: Chandu, Khyathi Raghavi, et al.
Published: (2024)
Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
by: Xiao, Wenli, et al.
Published: (2025)
by: Xiao, Wenli, et al.
Published: (2025)
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
by: Zhao, Yue, et al.
Published: (2025)
by: Zhao, Yue, et al.
Published: (2025)
A Recommender System for NFT Collectibles with Item Feature
by: Choi, Minjoo, et al.
Published: (2024)
by: Choi, Minjoo, et al.
Published: (2024)
MonoTher-Depth: Enhancing Thermal Depth Estimation via Confidence-Aware Distillation
by: Zuo, Xingxing, et al.
Published: (2025)
by: Zuo, Xingxing, et al.
Published: (2025)
HumanoidMimicGen: Data Generation for Loco-Manipulation via Whole-Body Planning
by: Lin, Kevin, et al.
Published: (2026)
by: Lin, Kevin, et al.
Published: (2026)
Environmental Footprint of GenAI Research: Insights from the Moshi Foundation Model
by: López-Rauhut, Marta, et al.
Published: (2026)
by: López-Rauhut, Marta, et al.
Published: (2026)
Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids
by: Lin, Toru, et al.
Published: (2025)
by: Lin, Toru, et al.
Published: (2025)
RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots
by: Nasiriany, Soroush, et al.
Published: (2026)
by: Nasiriany, Soroush, et al.
Published: (2026)
Latent Target Score Matching, with an application to Simulation-Based Inference
by: Ko, Joohwan, et al.
Published: (2026)
by: Ko, Joohwan, et al.
Published: (2026)
Model-Informed Flows for Bayesian Inference
by: Ko, Joohwan, et al.
Published: (2025)
by: Ko, Joohwan, et al.
Published: (2025)
Amortized Factor Inference Networks for Posterior Inference
by: Ko, Joohwan, et al.
Published: (2026)
by: Ko, Joohwan, et al.
Published: (2026)
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions
by: Awadalla, Anas, et al.
Published: (2024)
by: Awadalla, Anas, et al.
Published: (2024)
GenMol: A Drug Discovery Generalist with Discrete Diffusion
by: Lee, Seul, et al.
Published: (2025)
by: Lee, Seul, et al.
Published: (2025)
A Tunable Reflection Surface with Independently Variable Phase and Slope
by: Abbas, Omran, et al.
Published: (2024)
by: Abbas, Omran, et al.
Published: (2024)
Similar Items
-
TOTEM: TOkenized Time Series EMbeddings for General Time Series Analysis
by: Talukder, Sabera, et al.
Published: (2024) -
Find Any Part in 3D
by: Ma, Ziqi, et al.
Published: (2024) -
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
by: NVIDIA, et al.
Published: (2025) -
Visual Agentic AI for Spatial Reasoning with a Dynamic API
by: Marsili, Damiano, et al.
Published: (2025) -
Feedforward 3D Editing via Text-Steerable Image-to-3D
by: Ma, Ziqi, et al.
Published: (2025)