Reasoning Gets Harder for LLMs Inside A Dialogue
Fuente:
arXiv
Guardado en:
| Autores principales: | Kartáč, Ivan, Lango, Mateusz, Dušek, Ondřej |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
por: Kartáč, Ivan, et al.
Publicado: (2025)
por: Kartáč, Ivan, et al.
Publicado: (2025)
UFAL-CUNI at SemEval-2026 Task 11: An Efficient Modular Neuro-symbolic Method for Syllogistic Reasoning
por: Kartáč, Ivan, et al.
Publicado: (2026)
por: Kartáč, Ivan, et al.
Publicado: (2026)
LLM Agents Implement an NLG System from Scratch: Building Interpretable Rule-Based RDF-to-Text Generators
por: Lango, Mateusz, et al.
Publicado: (2025)
por: Lango, Mateusz, et al.
Publicado: (2025)
SRS-Stories: Vocabulary-constrained multilingual story generation for language learning
por: Kamzela, Wiktor, et al.
Publicado: (2025)
por: Kamzela, Wiktor, et al.
Publicado: (2025)
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
por: Wojciechowski, Adam, et al.
Publicado: (2024)
por: Wojciechowski, Adam, et al.
Publicado: (2024)
Leveraging Large Language Models for Building Interpretable Rule-Based Data-to-Text Systems
por: Warczyński, Jędrzej, et al.
Publicado: (2025)
por: Warczyński, Jędrzej, et al.
Publicado: (2025)
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
por: Balloccu, Simone, et al.
Publicado: (2024)
por: Balloccu, Simone, et al.
Publicado: (2024)
A Survey of Text Style Transfer: Applications and Ethical Implications
por: Mukherjee, Sourabrata, et al.
Publicado: (2024)
por: Mukherjee, Sourabrata, et al.
Publicado: (2024)
LLMs as Span Annotators: A Comparative Study of LLMs and Humans
por: Kasner, Zdeněk, et al.
Publicado: (2025)
por: Kasner, Zdeněk, et al.
Publicado: (2025)
LEEETs-Dial: Linguistic Entrainment in End-to-End Task-oriented Dialogue systems
por: Kumar, Nalin, et al.
Publicado: (2023)
por: Kumar, Nalin, et al.
Publicado: (2023)
AnimatedLLM: Explaining LLMs with Interactive Visualizations
por: Kasner, Zdeněk, et al.
Publicado: (2025)
por: Kasner, Zdeněk, et al.
Publicado: (2025)
Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text Generation
por: Kasner, Zdeněk, et al.
Publicado: (2024)
por: Kasner, Zdeněk, et al.
Publicado: (2024)
Generating clickbait spoilers with an ensemble of large language models
por: Woźny, Mateusz, et al.
Publicado: (2024)
por: Woźny, Mateusz, et al.
Publicado: (2024)
Text Style Transfer: An Introductory Overview
por: Mukherjee, Sourabrata, et al.
Publicado: (2024)
por: Mukherjee, Sourabrata, et al.
Publicado: (2024)
Polish-ASTE: Aspect-Sentiment Triplet Extraction Datasets for Polish
por: Lango, Marta, et al.
Publicado: (2025)
por: Lango, Marta, et al.
Publicado: (2025)
ASTE Transformer Modelling Dependencies in Aspect-Sentiment Triplet Extraction
por: Naglik, Iwo, et al.
Publicado: (2024)
por: Naglik, Iwo, et al.
Publicado: (2024)
FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation
por: Onderková, Kristýna, et al.
Publicado: (2025)
por: Onderková, Kristýna, et al.
Publicado: (2025)
Strategies for Span Labeling with Large Language Models
por: Semin, Danil, et al.
Publicado: (2026)
por: Semin, Danil, et al.
Publicado: (2026)
Real-World Summarization: When Evaluation Reaches Its Limits
por: Schmidtová, Patrícia, et al.
Publicado: (2025)
por: Schmidtová, Patrícia, et al.
Publicado: (2025)
Exploring ReAct Prompting for Task-Oriented Dialogue: Insights and Shortcomings
por: Elizabeth, Michelle, et al.
Publicado: (2024)
por: Elizabeth, Michelle, et al.
Publicado: (2024)
factgenie: A Framework for Span-based Evaluation of Generated Texts
por: Kasner, Zdeněk, et al.
Publicado: (2024)
por: Kasner, Zdeněk, et al.
Publicado: (2024)
Are Large Language Models Actually Good at Text Style Transfer?
por: Mukherjee, Sourabrata, et al.
Publicado: (2024)
por: Mukherjee, Sourabrata, et al.
Publicado: (2024)
Teaching LLMs at Charles University: Assignments and Activities
por: Helcl, Jindřich, et al.
Publicado: (2024)
por: Helcl, Jindřich, et al.
Publicado: (2024)
The Problem of Coherence in Natural Language Explanations of Recommendations
por: Raczyński, Jakub, et al.
Publicado: (2023)
por: Raczyński, Jakub, et al.
Publicado: (2023)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
por: Jin, Keyan, et al.
Publicado: (2025)
por: Jin, Keyan, et al.
Publicado: (2025)
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?
por: Mukherjee, Sourabrata, et al.
Publicado: (2025)
por: Mukherjee, Sourabrata, et al.
Publicado: (2025)
The Harder The Better: Maintaining Supervised Fine-tuning Generalization with Less but Harder Data
por: Shang, Zhaoyang, et al.
Publicado: (2025)
por: Shang, Zhaoyang, et al.
Publicado: (2025)
Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices
por: Schmidtová, Patrícia, et al.
Publicado: (2024)
por: Schmidtová, Patrícia, et al.
Publicado: (2024)
Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
por: Badola, Kartikeya, et al.
Publicado: (2025)
por: Badola, Kartikeya, et al.
Publicado: (2025)
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
por: Kim, Eunsu, et al.
Publicado: (2025)
por: Kim, Eunsu, et al.
Publicado: (2025)
Simpler becomes Harder: Do LLMs Exhibit a Coherent Behavior on Simplified Corpora?
por: Anschütz, Miriam, et al.
Publicado: (2024)
por: Anschütz, Miriam, et al.
Publicado: (2024)
Think Multilingual, Not Harder: A Data-Efficient Framework for Teaching Reasoning Models to Code-Switch
por: Lin, Eleanor M., et al.
Publicado: (2026)
por: Lin, Eleanor M., et al.
Publicado: (2026)
Text Detoxification as Style Transfer in English and Hindi
por: Mukherjee, Sourabrata, et al.
Publicado: (2024)
por: Mukherjee, Sourabrata, et al.
Publicado: (2024)
On the Universal Truthfulness Hyperplane Inside LLMs
por: Liu, Junteng, et al.
Publicado: (2024)
por: Liu, Junteng, et al.
Publicado: (2024)
Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded Reasoning
por: Zhu, Jiayuan, et al.
Publicado: (2025)
por: Zhu, Jiayuan, et al.
Publicado: (2025)
Are LLMs Robust for Spoken Dialogues?
por: Mousavi, Seyed Mahed, et al.
Publicado: (2024)
por: Mousavi, Seyed Mahed, et al.
Publicado: (2024)
Understanding the role of FFNs in driving multilingual behaviour in LLMs
por: Bhattacharya, Sunit, et al.
Publicado: (2024)
por: Bhattacharya, Sunit, et al.
Publicado: (2024)
Multilingual Text Style Transfer: Datasets & Models for Indian Languages
por: Mukherjee, Sourabrata, et al.
Publicado: (2024)
por: Mukherjee, Sourabrata, et al.
Publicado: (2024)
Inside-Out: Hidden Factual Knowledge in LLMs
por: Gekhman, Zorik, et al.
Publicado: (2025)
por: Gekhman, Zorik, et al.
Publicado: (2025)
Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench
por: Narad, Reuben, et al.
Publicado: (2025)
por: Narad, Reuben, et al.
Publicado: (2025)
Ejemplares similares
-
OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs
por: Kartáč, Ivan, et al.
Publicado: (2025) -
UFAL-CUNI at SemEval-2026 Task 11: An Efficient Modular Neuro-symbolic Method for Syllogistic Reasoning
por: Kartáč, Ivan, et al.
Publicado: (2026) -
LLM Agents Implement an NLG System from Scratch: Building Interpretable Rule-Based RDF-to-Text Generators
por: Lango, Mateusz, et al.
Publicado: (2025) -
SRS-Stories: Vocabulary-constrained multilingual story generation for language learning
por: Kamzela, Wiktor, et al.
Publicado: (2025) -
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
por: Wojciechowski, Adam, et al.
Publicado: (2024)