How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game Environment
Fuente:
arXiv
Guardado en:
| Autores principales: | Ansell, Rebecca, Toney-Wails, Autumn |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ReaGeo: Reasoning-Enhanced End-to-End Geocoding with LLMs
por: Cui, Jian, et al.
Publicado: (2026)
por: Cui, Jian, et al.
Publicado: (2026)
Critical Insights into Leading Conversational AI Models
por: Kohli, Urja, et al.
Publicado: (2025)
por: Kohli, Urja, et al.
Publicado: (2025)
ChatGPT4PCG Competition: Character-like Level Generation for Science Birds
por: Taveekitworachai, Pittawat, et al.
Publicado: (2023)
por: Taveekitworachai, Pittawat, et al.
Publicado: (2023)
HELEA: Hard-Negative Benchmark and LLM-based Reranking for Robust Entity Alignment
por: Jang, Yoonjin, et al.
Publicado: (2026)
por: Jang, Yoonjin, et al.
Publicado: (2026)
TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination
por: Li, Xu, et al.
Publicado: (2026)
por: Li, Xu, et al.
Publicado: (2026)
ChatGPT4PCG 2 Competition: Prompt Engineering for Science Birds Level Generation
por: Taveekitworachai, Pittawat, et al.
Publicado: (2024)
por: Taveekitworachai, Pittawat, et al.
Publicado: (2024)
GraphWalk: Enabling Reasoning in Large Language Models through Tool-Based Graph Navigation
por: Ghandi, Taraneh, et al.
Publicado: (2026)
por: Ghandi, Taraneh, et al.
Publicado: (2026)
GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
por: Zhang, Peiyan, et al.
Publicado: (2025)
por: Zhang, Peiyan, et al.
Publicado: (2025)
Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
por: Nakamura, Mason, et al.
Publicado: (2025)
por: Nakamura, Mason, et al.
Publicado: (2025)
Both Ends Count! Just How Good are LLM Agents at "Text-to-Big SQL"?
por: Eizaguirre, Germán T., et al.
Publicado: (2026)
por: Eizaguirre, Germán T., et al.
Publicado: (2026)
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
por: Steele, Brady, et al.
Publicado: (2026)
por: Steele, Brady, et al.
Publicado: (2026)
Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation
por: Koc, Vincent
Publicado: (2025)
por: Koc, Vincent
Publicado: (2025)
When Words Change the Model: Sensitivity of LLMs for Constraint Programming Modelling
por: Pellegrino, Alessio, et al.
Publicado: (2025)
por: Pellegrino, Alessio, et al.
Publicado: (2025)
Can AI Assist in Olympiad Coding
por: Ren, Samuel
Publicado: (2025)
por: Ren, Samuel
Publicado: (2025)
ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
por: Zhou, Yue, et al.
Publicado: (2025)
por: Zhou, Yue, et al.
Publicado: (2025)
Game of Thought: Robust Information Seeking with Large Language Models Using Game Theory
por: Cui, Langyuan, et al.
Publicado: (2026)
por: Cui, Langyuan, et al.
Publicado: (2026)
Automated Theorem Provers Help Improve Large Language Model Reasoning
por: McGinness, Lachlan, et al.
Publicado: (2024)
por: McGinness, Lachlan, et al.
Publicado: (2024)
Retrieval and Augmentation of Domain Knowledge for Text-to-SQL Semantic Parsing
por: Patwardhan, Manasi, et al.
Publicado: (2025)
por: Patwardhan, Manasi, et al.
Publicado: (2025)
Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents
por: Chua, Jaymari, et al.
Publicado: (2025)
por: Chua, Jaymari, et al.
Publicado: (2025)
Open-TI: Open Traffic Intelligence with Augmented Language Model
por: Da, Longchao, et al.
Publicado: (2023)
por: Da, Longchao, et al.
Publicado: (2023)
FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation
por: Hildebrand, Samuel, et al.
Publicado: (2025)
por: Hildebrand, Samuel, et al.
Publicado: (2025)
REPOT: Recoverable Program-of-Thought via Checkpoint Repair
por: Mazaheri, Parsa
Publicado: (2026)
por: Mazaheri, Parsa
Publicado: (2026)
Assisting humans in complex comparisons: automated information comparison at scale
por: Yuen, Truman, et al.
Publicado: (2024)
por: Yuen, Truman, et al.
Publicado: (2024)
Reinforced Language Models for Sequential Decision Making
por: Dilkes, Jim, et al.
Publicado: (2025)
por: Dilkes, Jim, et al.
Publicado: (2025)
REVOLVE: Optimizing AI Systems by Tracking Response Evolution in Textual Optimization
por: Zhang, Peiyan, et al.
Publicado: (2024)
por: Zhang, Peiyan, et al.
Publicado: (2024)
Language Models, Graph Searching, and Supervision Adulteration: When More Supervision is Less and How to Make More More
por: Frydenlund, Arvid
Publicado: (2025)
por: Frydenlund, Arvid
Publicado: (2025)
Improving Existing Optimization Algorithms with LLMs
por: Sartori, Camilo Chacón, et al.
Publicado: (2025)
por: Sartori, Camilo Chacón, et al.
Publicado: (2025)
The Reasoning-Creativity Trade-off: Toward Creativity-Driven Problem Solving
por: Luyten, Max Ruiz, et al.
Publicado: (2026)
por: Luyten, Max Ruiz, et al.
Publicado: (2026)
SCULPT: Constraint-Guided Pruned MCTS that Carves Efficient Paths for Mathematical Reasoning
por: Fang, Qitong, et al.
Publicado: (2026)
por: Fang, Qitong, et al.
Publicado: (2026)
From Search to Reasoning: A Five-Level RAG Capability Framework for Enterprise Data
por: Gill, Gurbinder, et al.
Publicado: (2025)
por: Gill, Gurbinder, et al.
Publicado: (2025)
Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy
por: He, Langzhou, et al.
Publicado: (2026)
por: He, Langzhou, et al.
Publicado: (2026)
Dynamic Policy Induction for Adaptive Prompt Optimization: Bridging the Efficiency-Accuracy Gap via Lightweight Reinforcement Learning
por: Xu, Jiexi
Publicado: (2025)
por: Xu, Jiexi
Publicado: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
por: Estevanell-Valladares, Ernesto L., et al.
Publicado: (2025)
por: Estevanell-Valladares, Ernesto L., et al.
Publicado: (2025)
Learning, Fast and Slow: Towards LLMs That Adapt Continually
por: Tiwari, Rishabh, et al.
Publicado: (2026)
por: Tiwari, Rishabh, et al.
Publicado: (2026)
GraphEval36K: Benchmarking Coding and Reasoning Capabilities of Large Language Models on Graph Datasets
por: Wu, Qiming, et al.
Publicado: (2024)
por: Wu, Qiming, et al.
Publicado: (2024)
ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
por: Chen, Hao, et al.
Publicado: (2025)
por: Chen, Hao, et al.
Publicado: (2025)
TRIZ Agents: A Multi-Agent LLM Approach for TRIZ-Based Innovation
por: Szczepanik, Kamil, et al.
Publicado: (2025)
por: Szczepanik, Kamil, et al.
Publicado: (2025)
On the Limits of Learned Importance Scoring for KV Cache Compression
por: Steele, Brady
Publicado: (2026)
por: Steele, Brady
Publicado: (2026)
Comparison of Unsupervised Metrics for Evaluating Judicial Decision Extraction
por: Litvak, Ivan Leonidovich, et al.
Publicado: (2025)
por: Litvak, Ivan Leonidovich, et al.
Publicado: (2025)
Survey Transfer Learning: Recycling Data with Silicon Responses
por: Amini, Ali
Publicado: (2025)
por: Amini, Ali
Publicado: (2025)
Ejemplares similares
-
ReaGeo: Reasoning-Enhanced End-to-End Geocoding with LLMs
por: Cui, Jian, et al.
Publicado: (2026) -
Critical Insights into Leading Conversational AI Models
por: Kohli, Urja, et al.
Publicado: (2025) -
ChatGPT4PCG Competition: Character-like Level Generation for Science Birds
por: Taveekitworachai, Pittawat, et al.
Publicado: (2023) -
HELEA: Hard-Negative Benchmark and LLM-based Reranking for Robust Entity Alignment
por: Jang, Yoonjin, et al.
Publicado: (2026) -
TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination
por: Li, Xu, et al.
Publicado: (2026)