Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kaesberg, Lars Benedikt, Yang, Tianyu, Bauer, Niklas, Ruas, Terry, Wahle, Jan Philip, Gipp, Bela |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SPaRC: A Spatial Pathfinding Reasoning Challenge
von: Kaesberg, Lars Benedikt, et al.
Veröffentlicht: (2025)
von: Kaesberg, Lars Benedikt, et al.
Veröffentlicht: (2025)
CiteAssist: A System for Automated Preprint Citation and BibTeX Generation
von: Kaesberg, Lars Benedikt, et al.
Veröffentlicht: (2024)
von: Kaesberg, Lars Benedikt, et al.
Veröffentlicht: (2024)
MALLM: Multi-Agent Large Language Models Framework
von: Becker, Jonas, et al.
Veröffentlicht: (2025)
von: Becker, Jonas, et al.
Veröffentlicht: (2025)
Voting or Consensus? Decision-Making in Multi-Agent Debate
von: Kaesberg, Lars Benedikt, et al.
Veröffentlicht: (2025)
von: Kaesberg, Lars Benedikt, et al.
Veröffentlicht: (2025)
Stay Focused: Problem Drift in Multi-Agent Debate
von: Becker, Jonas, et al.
Veröffentlicht: (2025)
von: Becker, Jonas, et al.
Veröffentlicht: (2025)
Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling
von: Wunderlich, Florian Valentin, et al.
Veröffentlicht: (2026)
von: Wunderlich, Florian Valentin, et al.
Veröffentlicht: (2026)
Paraphrase Types for Generation and Detection
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2023)
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2023)
Paraphrase Types Elicit Prompt Engineering Capabilities
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2024)
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2024)
Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges
von: Becker, Jonas, et al.
Veröffentlicht: (2024)
von: Becker, Jonas, et al.
Veröffentlicht: (2024)
Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains
von: Schmidt, Finn, et al.
Veröffentlicht: (2026)
von: Schmidt, Finn, et al.
Veröffentlicht: (2026)
Piecing Together Cross-Document Coreference Resolution Datasets: Systematic Dataset Analysis and Unification
von: Zhukova, Anastasia, et al.
Veröffentlicht: (2026)
von: Zhukova, Anastasia, et al.
Veröffentlicht: (2026)
What's under the hood: Investigating Automatic Metrics on Meeting Summarization
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
CADS: A Systematic Literature Review on the Challenges of Abstractive Dialogue Summarization
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
How Large Language Models are Transforming Machine-Paraphrased Plagiarism
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2022)
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2022)
Towards Human Understanding of Paraphrase Types in Large Language Models
von: Meier, Dominik, et al.
Veröffentlicht: (2024)
von: Meier, Dominik, et al.
Veröffentlicht: (2024)
TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent
von: Meier, Dominik, et al.
Veröffentlicht: (2025)
von: Meier, Dominik, et al.
Veröffentlicht: (2025)
You need to MIMIC to get FAME: Solving Meeting Transcript Scarcity with a Multi-Agent Conversations
von: Kirstein, Frederic, et al.
Veröffentlicht: (2025)
von: Kirstein, Frederic, et al.
Veröffentlicht: (2025)
D3: A Massive Dataset of Scholarly Metadata for Analyzing the State of Computer Science Research
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2022)
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2022)
We are Who We Cite: Bridges of Influence Between Natural Language Processing and Other Academic Fields
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2023)
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2023)
Big Tech-Funded AI Papers Have Higher Citation Impact, Greater Insularity, and Larger Recency Bias
von: Gnewuch, Max Martin, et al.
Veröffentlicht: (2025)
von: Gnewuch, Max Martin, et al.
Veröffentlicht: (2025)
Citation Amnesia: On The Recency Bias of NLP and Other Academic Fields
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2024)
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2024)
Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
Testing the Generalization of Neural Language Models for COVID-19 Misinformation Detection
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2021)
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2021)
What's Wrong? Refining Meeting Summaries with LLM Feedback
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents
von: Yang, Tianyu, et al.
Veröffentlicht: (2025)
von: Yang, Tianyu, et al.
Veröffentlicht: (2025)
Overview of the Plagiarism Detection Task at PAN 2025
von: Greiner-Petter, André, et al.
Veröffentlicht: (2025)
von: Greiner-Petter, André, et al.
Veröffentlicht: (2025)
Tell me what I need to know: Exploring LLM-based (Personalized) Abstractive Multi-Source Meeting Summarization
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
von: Kirstein, Frederic, et al.
Veröffentlicht: (2024)
Re-FRAME the Meeting Summarization SCOPE: Fact-Based Summarization and Personalization via Questions
von: Kirstein, Frederic, et al.
Veröffentlicht: (2025)
von: Kirstein, Frederic, et al.
Veröffentlicht: (2025)
Affect, Body, Cognition, Demographics, and Emotion: The ABCDE of Text Features for Computational Affective Science
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2025)
von: Wahle, Jan Philip, et al.
Veröffentlicht: (2025)
What's in the News? Towards Identification of Bias by Commission, Omission, and Source Selection (COSS)
von: Zhukova, Anastasia, et al.
Veröffentlicht: (2025)
von: Zhukova, Anastasia, et al.
Veröffentlicht: (2025)
MAGPIE: Multi-Task Media-Bias Analysis Generalization for Pre-Trained Identification of Expressions
von: Horych, Tomáš, et al.
Veröffentlicht: (2024)
von: Horych, Tomáš, et al.
Veröffentlicht: (2024)
The Media Bias Taxonomy: A Systematic Literature Review on the Forms and Automated Detection of Media Bias
von: Spinde, Timo, et al.
Veröffentlicht: (2023)
von: Spinde, Timo, et al.
Veröffentlicht: (2023)
The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection
von: Horych, Tomas, et al.
Veröffentlicht: (2024)
von: Horych, Tomas, et al.
Veröffentlicht: (2024)
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
von: Shen, Yujiong, et al.
Veröffentlicht: (2026)
von: Shen, Yujiong, et al.
Veröffentlicht: (2026)
Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025)
von: Stogiannidis, Ilias, et al.
Veröffentlicht: (2025)
Automated Collection of Evaluation Dataset for Semantic Search in Low-Resource Domain Language
von: Zhukova, Anastasia, et al.
Veröffentlicht: (2024)
von: Zhukova, Anastasia, et al.
Veröffentlicht: (2024)
Language Modeling and Understanding Through Paraphrase Generation and Detection
von: Wahle, Jan Philip
Veröffentlicht: (2026)
von: Wahle, Jan Philip
Veröffentlicht: (2026)
The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing Research
von: Abdalla, Mohamed, et al.
Veröffentlicht: (2023)
von: Abdalla, Mohamed, et al.
Veröffentlicht: (2023)
Evaluating Step-by-Step Reasoning through Symbolic Verification
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2022)
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2022)
LogSigma at SemEval-2026 Task 3: Uncertainty-Weighted Multitask Learning for Dimensional Aspect-Based Sentiment Analysis
von: Hikal, Baraa, et al.
Veröffentlicht: (2026)
von: Hikal, Baraa, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SPaRC: A Spatial Pathfinding Reasoning Challenge
von: Kaesberg, Lars Benedikt, et al.
Veröffentlicht: (2025) -
CiteAssist: A System for Automated Preprint Citation and BibTeX Generation
von: Kaesberg, Lars Benedikt, et al.
Veröffentlicht: (2024) -
MALLM: Multi-Agent Large Language Models Framework
von: Becker, Jonas, et al.
Veröffentlicht: (2025) -
Voting or Consensus? Decision-Making in Multi-Agent Debate
von: Kaesberg, Lars Benedikt, et al.
Veröffentlicht: (2025) -
Stay Focused: Problem Drift in Multi-Agent Debate
von: Becker, Jonas, et al.
Veröffentlicht: (2025)