TextQuests: How Good are LLMs at Text-Based Video Games?
Fuente:
arXiv
Saved in:
| Main Authors: | Phan, Long, Mazeika, Mantas, Zou, Andy, Hendrycks, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tamper-Resistant Safeguards for Open-Weight LLMs
by: Tamirisa, Rishub, et al.
Published: (2024)
by: Tamirisa, Rishub, et al.
Published: (2024)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
by: Mazeika, Mantas, et al.
Published: (2024)
by: Mazeika, Mantas, et al.
Published: (2024)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024)
by: Ren, Richard, et al.
Published: (2024)
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
by: Mazeika, Mantas, et al.
Published: (2025)
by: Mazeika, Mantas, et al.
Published: (2025)
GenQuest: An LLM-based Text Adventure Game for Language Learners
by: Wang, Qiao, et al.
Published: (2025)
by: Wang, Qiao, et al.
Published: (2025)
Aggressive Compression Enables LLM Weight Theft
by: Brown, Davis, et al.
Published: (2026)
by: Brown, Davis, et al.
Published: (2026)
Reducing Political Manipulation with Consistency Training
by: Phan, Long, et al.
Published: (2026)
by: Phan, Long, et al.
Published: (2026)
Are LLMs Good Text Diacritizers? An Arabic and Yoruba Case Study
by: Toyin, Hawau Olamide, et al.
Published: (2025)
by: Toyin, Hawau Olamide, et al.
Published: (2025)
TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning
by: Hudi, Frederikus, et al.
Published: (2025)
by: Hudi, Frederikus, et al.
Published: (2025)
HawkEye: Training Video-Text LLMs for Grounding Text in Videos
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
Improving Alignment and Robustness with Circuit Breakers
by: Zou, Andy, et al.
Published: (2024)
by: Zou, Andy, et al.
Published: (2024)
Commonsense Knowledge Editing Based on Free-Text in LLMs
by: Huang, Xiusheng, et al.
Published: (2024)
by: Huang, Xiusheng, et al.
Published: (2024)
LLMs Enable Bag-of-Texts Representations for Short-Text Clustering
by: Lin, I-Fan, et al.
Published: (2025)
by: Lin, I-Fan, et al.
Published: (2025)
Automatic Bug Detection in LLM-Powered Text-Based Games Using LLMs
by: Jin, Claire, et al.
Published: (2024)
by: Jin, Claire, et al.
Published: (2024)
Representation Engineering: A Top-Down Approach to AI Transparency
by: Zou, Andy, et al.
Published: (2023)
by: Zou, Andy, et al.
Published: (2023)
Playing With AI: How Do State-Of-The-Art Large Language Models Perform in the 1977 Text-Based Adventure Game Zork?
by: Gerrits, Berry
Published: (2026)
by: Gerrits, Berry
Published: (2026)
A Text-to-Game Engine for UGC-Based Role-Playing Games
by: Zhang, Lei, et al.
Published: (2024)
by: Zhang, Lei, et al.
Published: (2024)
The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
by: Ren, Richard, et al.
Published: (2025)
by: Ren, Richard, et al.
Published: (2025)
Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation
by: Phan, Phuc, et al.
Published: (2024)
by: Phan, Phuc, et al.
Published: (2024)
Persona Dynamics: Unveiling the Impact of Personality Traits on Agents in Text-Based Games
by: Lim, Seungwon, et al.
Published: (2025)
by: Lim, Seungwon, et al.
Published: (2025)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
by: Zhang, Ran, et al.
Published: (2024)
by: Zhang, Ran, et al.
Published: (2024)
How to Evaluate Coreference in Literary Texts?
by: Duron-Tejedor, Ana-Isabel, et al.
Published: (2023)
by: Duron-Tejedor, Ana-Isabel, et al.
Published: (2023)
Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs
by: Li, Yanhong, et al.
Published: (2025)
by: Li, Yanhong, et al.
Published: (2025)
Trustful LLMs: Customizing and Grounding Text Generation with Knowledge Bases and Dual Decoders
by: Zhu, Xiaofeng, et al.
Published: (2024)
by: Zhu, Xiaofeng, et al.
Published: (2024)
Conditioning LLMs to Generate Code-Switched Text
by: Heredia, Maite, et al.
Published: (2025)
by: Heredia, Maite, et al.
Published: (2025)
How Can We Effectively Expand the Vocabulary of LLMs with 0.01GB of Target Language Text?
by: Yamaguchi, Atsuki, et al.
Published: (2024)
by: Yamaguchi, Atsuki, et al.
Published: (2024)
How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game Environment
by: Ansell, Rebecca, et al.
Published: (2026)
by: Ansell, Rebecca, et al.
Published: (2026)
ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?
by: Zhang, Leixin, et al.
Published: (2024)
by: Zhang, Leixin, et al.
Published: (2024)
DuConTE: Dual-Granularity Text Encoder with Topology-Constrained Attention for Text-attributed Graphs
by: Liang, Lexuan, et al.
Published: (2026)
by: Liang, Lexuan, et al.
Published: (2026)
Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
Position: LLMs Can be Good Tutors in English Education
by: Ye, Jingheng, et al.
Published: (2025)
by: Ye, Jingheng, et al.
Published: (2025)
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
by: Wang, Boxin, et al.
Published: (2023)
by: Wang, Boxin, et al.
Published: (2023)
Estimating Causal Effects of Text Interventions Leveraging LLMs
by: Guo, Siyi, et al.
Published: (2024)
by: Guo, Siyi, et al.
Published: (2024)
Difficulty Estimation and Simplification of French Text Using LLMs
by: Jamet, Henri, et al.
Published: (2024)
by: Jamet, Henri, et al.
Published: (2024)
Can LLMs Follow Simple Rules?
by: Mu, Norman, et al.
Published: (2023)
by: Mu, Norman, et al.
Published: (2023)
QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentation
by: Li, Jiazheng, et al.
Published: (2025)
by: Li, Jiazheng, et al.
Published: (2025)
Rethinking the Role of Text Complexity in Language Model Pretraining
by: Velasco, Dan John, et al.
Published: (2025)
by: Velasco, Dan John, et al.
Published: (2025)
Can LLMs Narrate Tabular Data? An Evaluation Framework for Natural Language Representations of Text-to-SQL System Outputs
by: Singh, Jyotika, et al.
Published: (2025)
by: Singh, Jyotika, et al.
Published: (2025)
Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
by: Islam, Tunazzina
Published: (2026)
by: Islam, Tunazzina
Published: (2026)
Similar Items
-
Tamper-Resistant Safeguards for Open-Weight LLMs
by: Tamirisa, Rishub, et al.
Published: (2024) -
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
by: Mazeika, Mantas, et al.
Published: (2024) -
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024) -
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
by: Mazeika, Mantas, et al.
Published: (2025) -
GenQuest: An LLM-based Text Adventure Game for Language Learners
by: Wang, Qiao, et al.
Published: (2025)