Gespeichert in:
| Hauptverfasser: | Hakimov, Sherzod, Bernard, Roland, Leiber, Tim, Osswald, Karl, Richert, Kristina, Yang, Ruilin, Bernardi, Raffaella, Schlangen, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2510.08098 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2026)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2026)
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
Ad-hoc Concept Forming in the Game Codenames as a Means for Evaluating Large Language Models
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2025)
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2025)
Sharing the Cost of Success: A Game for Evaluating and Learning Collaborative Multi-Agent Instruction Giving and Following Policies
von: Sadler, Philipp, et al.
Veröffentlicht: (2024)
von: Sadler, Philipp, et al.
Veröffentlicht: (2024)
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2024)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2024)
Plant in Cupboard, Orange on Rably, Inat Aphone. Benchmarking Incremental Learning of Situation and Language Model using a Text-Simulated Situated Environment
von: Jordan, Jonathan, et al.
Veröffentlicht: (2025)
von: Jordan, Jonathan, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2024)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2024)
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
Learning Communication Policies for Different Follower Behaviors in a Collaborative Reference Game
von: Sadler, Philipp, et al.
Veröffentlicht: (2024)
von: Sadler, Philipp, et al.
Veröffentlicht: (2024)
How Many Parameters Does it Take to Change a Light Bulb? Evaluating Performance in Self-Play of Conversational Games as a Function of Model Characteristics
von: Bhavsar, Nidhir, et al.
Veröffentlicht: (2024)
von: Bhavsar, Nidhir, et al.
Veröffentlicht: (2024)
TurkicNLP: An NLP Toolkit for Turkic Languages
von: Hakimov, Sherzod
Veröffentlicht: (2026)
von: Hakimov, Sherzod
Veröffentlicht: (2026)
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
von: Beyer, Anne, et al.
Veröffentlicht: (2024)
von: Beyer, Anne, et al.
Veröffentlicht: (2024)
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2026)
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2026)
Unveiling Global Narratives: A Multilingual Twitter Dataset of News Media on the Russo-Ukrainian Conflict
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2023)
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2023)
M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets
von: Thakkar, Gaurish, et al.
Veröffentlicht: (2024)
von: Thakkar, Gaurish, et al.
Veröffentlicht: (2024)
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
von: Schlangen, David, et al.
Veröffentlicht: (2025)
von: Schlangen, David, et al.
Veröffentlicht: (2025)
Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2024)
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2024)
A Systematic Analysis of Large Language Models as Soft Reasoners: The Case of Syllogistic Inferences
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2024)
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2024)
Free-text Rationale Generation under Readability Level Control
von: Hsu, Yi-Sheng, et al.
Veröffentlicht: (2024)
von: Hsu, Yi-Sheng, et al.
Veröffentlicht: (2024)
Playpen: An Environment for Exploring Learning Through Conversational Interaction
von: Horst, Nicola, et al.
Veröffentlicht: (2025)
von: Horst, Nicola, et al.
Veröffentlicht: (2025)
What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025
von: Yang, Jing, et al.
Veröffentlicht: (2026)
von: Yang, Jing, et al.
Veröffentlicht: (2026)
How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
Strategic Responses to Personalized Pricing and Demand for Privacy: An Experiment
von: Bó, Inácio, et al.
Veröffentlicht: (2023)
von: Bó, Inácio, et al.
Veröffentlicht: (2023)
Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes
von: Li, Meng, et al.
Veröffentlicht: (2025)
von: Li, Meng, et al.
Veröffentlicht: (2025)
The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
von: Bertolazzi, Leonardo, et al.
Veröffentlicht: (2025)
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests
von: Momentè, Filippo, et al.
Veröffentlicht: (2025)
von: Momentè, Filippo, et al.
Veröffentlicht: (2025)
Taking Action Towards Graceful Interaction: The Effects of Performing Actions on Modelling Policies for Instruction Clarification Requests
von: Madureira, Brielen, et al.
Veröffentlicht: (2024)
von: Madureira, Brielen, et al.
Veröffentlicht: (2024)
LLMs as Function Approximators: Terminology, Taxonomy, and Questions for Evaluation
von: Schlangen, David
Veröffentlicht: (2024)
von: Schlangen, David
Veröffentlicht: (2024)
SoT: Structured-of-Thought Prompting Guides Multilingual Reasoning in Large Language Models
von: Qi, Rui, et al.
Veröffentlicht: (2025)
von: Qi, Rui, et al.
Veröffentlicht: (2025)
The Price of a Second Thought: On the Evaluation of Reasoning Efficiency in Large Language Models
von: Fan, Siqi, et al.
Veröffentlicht: (2025)
von: Fan, Siqi, et al.
Veröffentlicht: (2025)
Diastereoselective Synthesis of Pyridone ribo‐C‐Nucleosides via Heck Reaction and Oxidation
von: Tim Gniech, et al.
Veröffentlicht: (2024)
von: Tim Gniech, et al.
Veröffentlicht: (2024)
Front Cover: Diastereoselective Synthesis of Pyridone ribo‐C‐Nucleosides via Heck Reaction and Oxidation (Eur. J. Org. Chem. 28/2024)
von: Tim Gniech, et al.
Veröffentlicht: (2024)
von: Tim Gniech, et al.
Veröffentlicht: (2024)
Medicina e arte uma ressonância
von: Walter Osswald
Veröffentlicht: (2002)
von: Walter Osswald
Veröffentlicht: (2002)
Aspectos de autoridad y poder en las ceremonias de canonización de Ignacio de Loyola y Francisco Javier en Portugal
von: Cristina Osswald
Veröffentlicht: (2013)
von: Cristina Osswald
Veröffentlicht: (2013)
On Christian Martyrdom in Japan (1597-1658)
von: Cristina Osswald
Veröffentlicht: (2021)
von: Cristina Osswald
Veröffentlicht: (2021)
EL OLVIDO EN LA FENOMENOLOGÍA DE HUSSERL. DOS FENÓMENOS LÍMITE
von: Andrés Osswald
Veröffentlicht: (2017)
von: Andrés Osswald
Veröffentlicht: (2017)
A perceção dos jesuítas no mundo português: entre o trato de e o gosto por orientalia (sécs. XVI-XVII)
von: Cristina Osswald
Veröffentlicht: (2018)
von: Cristina Osswald
Veröffentlicht: (2018)
Sete notas sobre a cura pelo nada
von: Walter Osswald
Veröffentlicht: (2003)
von: Walter Osswald
Veröffentlicht: (2003)
Las narrativas de "las pioneras". Cuestiones de género y moralidades en el desarrollo de la danza moderna en la Argentina (1940-1960)
von: Denise Osswald
Veröffentlicht: (2010)
von: Denise Osswald
Veröffentlicht: (2010)
SR-FoT: A Syllogistic-Reasoning Framework of Thought for Large Language Models Tackling Knowledge-based Reasoning Tasks
von: Wan, Wentao, et al.
Veröffentlicht: (2025)
von: Wan, Wentao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2026) -
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025) -
Ad-hoc Concept Forming in the Game Codenames as a Means for Evaluating Large Language Models
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2025) -
Sharing the Cost of Success: A Game for Evaluating and Learning Collaborative Multi-Agent Instruction Giving and Following Policies
von: Sadler, Philipp, et al.
Veröffentlicht: (2024) -
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2024)