Playing NetHack with LLMs: Potential & Limitations as Zero-Shot Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jeurissen, Dominik, Perez-Liebana, Diego, Gow, Jeremy, Cakmak, Duygu, Kwan, James |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LuckyMera: a Modular AI Framework for Building Hybrid NetHack Agents
von: Quarantiello, Luigi, et al.
Veröffentlicht: (2023)
von: Quarantiello, Luigi, et al.
Veröffentlicht: (2023)
From Code to Play: Benchmarking Program Search for Games Using Large Language Models
von: Eberhardinger, Manuel, et al.
Veröffentlicht: (2024)
von: Eberhardinger, Manuel, et al.
Veröffentlicht: (2024)
Strategy Game-Playing with Size-Constrained State Abstraction
von: Xu, Linjie, et al.
Veröffentlicht: (2024)
von: Xu, Linjie, et al.
Veröffentlicht: (2024)
Play Style Identification Using Low-Level Representations of Play Traces in MicroRTS
von: Xia, Ruizhe Yu, et al.
Veröffentlicht: (2025)
von: Xia, Ruizhe Yu, et al.
Veröffentlicht: (2025)
Seeding for Success: Skill and Stochasticity in Tabletop Games
von: Goodman, James, et al.
Veröffentlicht: (2025)
von: Goodman, James, et al.
Veröffentlicht: (2025)
Exploration with Foundation Models: Capabilities, Limitations, and Hybrid Approaches
von: Sasso, Remo, et al.
Veröffentlicht: (2025)
von: Sasso, Remo, et al.
Veröffentlicht: (2025)
Evaluating Environments Using Exploratory Agents
von: Khaleque, Bobby, et al.
Veröffentlicht: (2024)
von: Khaleque, Bobby, et al.
Veröffentlicht: (2024)
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
von: Taylor, Mia, et al.
Veröffentlicht: (2025)
von: Taylor, Mia, et al.
Veröffentlicht: (2025)
PyTAG: Tabletop Games for Multi-Agent Reinforcement Learning
von: Balla, Martin, et al.
Veröffentlicht: (2024)
von: Balla, Martin, et al.
Veröffentlicht: (2024)
JSON-Bag: A generic game trajectory representation
von: Nguyen, Dien, et al.
Veröffentlicht: (2025)
von: Nguyen, Dien, et al.
Veröffentlicht: (2025)
Hacking CTFs with Plain Agents
von: Turtayev, Rustem, et al.
Veröffentlicht: (2024)
von: Turtayev, Rustem, et al.
Veröffentlicht: (2024)
Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds
von: Sasso, Remo, et al.
Veröffentlicht: (2025)
von: Sasso, Remo, et al.
Veröffentlicht: (2025)
Plug-and-Play Emotion Graphs for Compositional Prompting in Zero-Shot Speech Emotion Recognition
von: Shi, Jiacheng, et al.
Veröffentlicht: (2025)
von: Shi, Jiacheng, et al.
Veröffentlicht: (2025)
Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition
von: Schulhoff, Sander, et al.
Veröffentlicht: (2023)
von: Schulhoff, Sander, et al.
Veröffentlicht: (2023)
The Multi-Agent Reinforcement Learning in MalmÖ (MARLÖ) Competition
von: Perez-Liebana, Diego, et al.
Veröffentlicht: (2019)
von: Perez-Liebana, Diego, et al.
Veröffentlicht: (2019)
Reasoning with LLMs for Zero-Shot Vulnerability Detection
von: Zibaeirad, Arastoo, et al.
Veröffentlicht: (2025)
von: Zibaeirad, Arastoo, et al.
Veröffentlicht: (2025)
RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents
von: Atinafu, Yonas, et al.
Veröffentlicht: (2026)
von: Atinafu, Yonas, et al.
Veröffentlicht: (2026)
LLM Agents can Autonomously Hack Websites
von: Fang, Richard, et al.
Veröffentlicht: (2024)
von: Fang, Richard, et al.
Veröffentlicht: (2024)
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
von: Reworr, et al.
Veröffentlicht: (2024)
von: Reworr, et al.
Veröffentlicht: (2024)
Zero-Shot Clinical Trial Patient Matching with LLMs
von: Wornow, Michael, et al.
Veröffentlicht: (2024)
von: Wornow, Michael, et al.
Veröffentlicht: (2024)
Are Stereotypes Leading LLMs' Zero-Shot Stance Detection ?
von: Dubreuil, Anthony, et al.
Veröffentlicht: (2025)
von: Dubreuil, Anthony, et al.
Veröffentlicht: (2025)
Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
von: Liang, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Liang, Zhiyuan, et al.
Veröffentlicht: (2025)
AI and the Net-Zero Journey: Energy Demand, Emissions, and the Potential for Transition
von: Devarakota, Pandu, et al.
Veröffentlicht: (2025)
von: Devarakota, Pandu, et al.
Veröffentlicht: (2025)
SmartPlay: A Benchmark for LLMs as Intelligent Agents
von: Wu, Yue, et al.
Veröffentlicht: (2023)
von: Wu, Yue, et al.
Veröffentlicht: (2023)
Higher Replay Ratio Empowers Sample-Efficient Multi-Agent Reinforcement Learning
von: Xu, Linjie, et al.
Veröffentlicht: (2024)
von: Xu, Linjie, et al.
Veröffentlicht: (2024)
Proof-of-Use: Mitigating Tool-Call Hacking in Deep Research Agents
von: Ma, SHengjie, et al.
Veröffentlicht: (2025)
von: Ma, SHengjie, et al.
Veröffentlicht: (2025)
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
von: Helff, Lukas, et al.
Veröffentlicht: (2026)
von: Helff, Lukas, et al.
Veröffentlicht: (2026)
Frog Soup: Zero-Shot, In-Context, and Sample-Efficient Frogger Agents
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
Wonderful Team: Zero-Shot Physical Task Planning with Visual LLMs
von: Wang, Zidan, et al.
Veröffentlicht: (2024)
von: Wang, Zidan, et al.
Veröffentlicht: (2024)
MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
von: Yuan, Huining, et al.
Veröffentlicht: (2025)
von: Yuan, Huining, et al.
Veröffentlicht: (2025)
Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
von: Pignatelli, Eduardo, et al.
Veröffentlicht: (2024)
von: Pignatelli, Eduardo, et al.
Veröffentlicht: (2024)
Zero-Shot Robustification of Zero-Shot Models
von: Adila, Dyah, et al.
Veröffentlicht: (2023)
von: Adila, Dyah, et al.
Veröffentlicht: (2023)
Zero-Shot Action Generalization with Limited Observations
von: Alchihabi, Abdullah, et al.
Veröffentlicht: (2025)
von: Alchihabi, Abdullah, et al.
Veröffentlicht: (2025)
ROAD: Reflective Optimization via Automated Debugging for Zero-Shot Agent Alignment
von: Temyingyong, Natchaya, et al.
Veröffentlicht: (2025)
von: Temyingyong, Natchaya, et al.
Veröffentlicht: (2025)
LLMs are not Zero-Shot Reasoners for Biomedical Information Extraction
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
Univariate to Multivariate: LLMs as Zero-Shot Predictors for Time-Series Forecasting
von: Madarasingha, Chamara, et al.
Veröffentlicht: (2025)
von: Madarasingha, Chamara, et al.
Veröffentlicht: (2025)
MultiVer: Zero-Shot Multi-Agent Vulnerability Detection
von: Rajan, Shreshth
Veröffentlicht: (2026)
von: Rajan, Shreshth
Veröffentlicht: (2026)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
von: Roth, Amit, et al.
Veröffentlicht: (2026)
von: Roth, Amit, et al.
Veröffentlicht: (2026)
Smooth Operators: LLMs Translating Imperfect Hints into Disfluency-Rich Transcripts
von: Altinok, Duygu
Veröffentlicht: (2025)
von: Altinok, Duygu
Veröffentlicht: (2025)
Enhancing Zero-Shot Time Series Forecasting in Off-the-Shelf LLMs via Noise Injection
von: Yin, Xingyou, et al.
Veröffentlicht: (2025)
von: Yin, Xingyou, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LuckyMera: a Modular AI Framework for Building Hybrid NetHack Agents
von: Quarantiello, Luigi, et al.
Veröffentlicht: (2023) -
From Code to Play: Benchmarking Program Search for Games Using Large Language Models
von: Eberhardinger, Manuel, et al.
Veröffentlicht: (2024) -
Strategy Game-Playing with Size-Constrained State Abstraction
von: Xu, Linjie, et al.
Veröffentlicht: (2024) -
Play Style Identification Using Low-Level Representations of Play Traces in MicroRTS
von: Xia, Ruizhe Yu, et al.
Veröffentlicht: (2025) -
Seeding for Success: Skill and Stochasticity in Tabletop Games
von: Goodman, James, et al.
Veröffentlicht: (2025)