Is Length Really A Liability? An Evaluation of Multi-turn LLM Conversations using BoolQ
Fuente:
arXiv
Saved in:
| Main Authors: | Neergaard, Karl, Qiu, Le, Chersoni, Emmanuele |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StockGenChaR: A Study on the Evaluation of Large Vision-Language Models on Stock Chart Captioning
by: Qiu, Le, et al.
Published: (2024)
by: Qiu, Le, et al.
Published: (2024)
Empirical Sufficiency Lower Bounds for Language Modeling with Locally-Bootstrapped Semantic Structures
by: Prange, Jakob, et al.
Published: (2023)
by: Prange, Jakob, et al.
Published: (2023)
Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from Demonstratives
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
Composing or Not Composing? Towards Distributional Construction Grammars
by: Blache, Philippe, et al.
Published: (2024)
by: Blache, Philippe, et al.
Published: (2024)
Learning to Look at the Other Side: A Semantic Probing Study of Word Embeddings in LLMs with Enabled Bidirectional Attention
by: Feng, Zhaoxin, et al.
Published: (2025)
by: Feng, Zhaoxin, et al.
Published: (2025)
Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy
by: Feng, Zhaoxin, et al.
Published: (2026)
by: Feng, Zhaoxin, et al.
Published: (2026)
From BERT to LLMs: Comparing and Understanding Chinese Classifier Prediction in Language Models
by: Zhang, Ziqi, et al.
Published: (2025)
by: Zhang, Ziqi, et al.
Published: (2025)
Log Probabilities Are a Reliable Estimate of Semantic Plausibility in Base and Instruction-Tuned Language Models
by: Kauf, Carina, et al.
Published: (2024)
by: Kauf, Carina, et al.
Published: (2024)
Sparse Brains are Also Adaptive Brains: Cognitive-Load-Aware Dynamic Activation for LLMs
by: Yang, Yiheng, et al.
Published: (2025)
by: Yang, Yiheng, et al.
Published: (2025)
Token Statistics Reveal Conversational Drift in Multi-turn LLM Interaction
by: Hafez, Wael, et al.
Published: (2026)
by: Hafez, Wael, et al.
Published: (2026)
MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety
by: Song, Jialin, et al.
Published: (2026)
by: Song, Jialin, et al.
Published: (2026)
ExpliCa: Evaluating Explicit Causal Reasoning in Large Language Models
by: Miliani, Martina, et al.
Published: (2025)
by: Miliani, Martina, et al.
Published: (2025)
Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detection
by: Kumarage, Tharindu, et al.
Published: (2025)
by: Kumarage, Tharindu, et al.
Published: (2025)
REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM
by: Jindal, Madhur, et al.
Published: (2025)
by: Jindal, Madhur, et al.
Published: (2025)
On the Multi-turn Instruction Following for Conversational Web Agents
by: Deng, Yang, et al.
Published: (2024)
by: Deng, Yang, et al.
Published: (2024)
SAGE: A Top-Down Bottom-Up Knowledge-Grounded User Simulator for Multi-turn AGent Evaluation
by: Shea, Ryan, et al.
Published: (2025)
by: Shea, Ryan, et al.
Published: (2025)
Stop Listening to Me! How Multi-turn Conversations Can Degrade LLM Reliability
by: Guo, Kevin H., et al.
Published: (2026)
by: Guo, Kevin H., et al.
Published: (2026)
Pixel-Level Reasoning Segmentation via Multi-turn Conversations
by: Cai, Dexian, et al.
Published: (2025)
by: Cai, Dexian, et al.
Published: (2025)
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation
by: Cheng, Yiruo, et al.
Published: (2024)
by: Cheng, Yiruo, et al.
Published: (2024)
FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs
by: Fan, Zhiting, et al.
Published: (2024)
by: Fan, Zhiting, et al.
Published: (2024)
SMILE: Single-turn to Multi-turn Inclusive Language Expansion via ChatGPT for Mental Health Support
by: Qiu, Huachuan, et al.
Published: (2023)
by: Qiu, Huachuan, et al.
Published: (2023)
MAC: A Multi-Agent Framework for Interactive User Clarification in Multi-turn Conversations
by: Acikgoz, Emre Can, et al.
Published: (2025)
by: Acikgoz, Emre Can, et al.
Published: (2025)
A Decade-Scale Benchmark Evaluating LLMs' Clinical Practice Guidelines Detection and Adherence in Multi-turn Conversations
by: Tan, Andong, et al.
Published: (2026)
by: Tan, Andong, et al.
Published: (2026)
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
by: Ma, Chang, et al.
Published: (2024)
by: Ma, Chang, et al.
Published: (2024)
NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation
by: Wang, Xiaoyang, et al.
Published: (2021)
by: Wang, Xiaoyang, et al.
Published: (2021)
Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model
by: Niu, Fuqiang, et al.
Published: (2024)
by: Niu, Fuqiang, et al.
Published: (2024)
Source-primed Multi-turn Conversation Helps Large Language Models Translate Documents
by: Hu, Hanxu, et al.
Published: (2025)
by: Hu, Hanxu, et al.
Published: (2025)
Mining Beyond the Bools: Learning Data Transformations and Temporal Specifications
by: Kouteili, Sam Nicholas, et al.
Published: (2026)
by: Kouteili, Sam Nicholas, et al.
Published: (2026)
BoolQuestions: Does Dense Retrieval Understand Boolean Logic in Language?
by: Zhang, Zongmeng, et al.
Published: (2024)
by: Zhang, Zongmeng, et al.
Published: (2024)
Asymmetric Actor-Critic for Multi-turn LLM Agents
by: Jiang, Shuli, et al.
Published: (2026)
by: Jiang, Shuli, et al.
Published: (2026)
Is Your LLM Really Mastering the Concept? A Multi-Agent Benchmark
by: Xu, Shuhang, et al.
Published: (2025)
by: Xu, Shuhang, et al.
Published: (2025)
SAPIENT: Mastering Multi-turn Conversational Recommendation with Strategic Planning and Monte Carlo Tree Search
by: Du, Hanwen, et al.
Published: (2024)
by: Du, Hanwen, et al.
Published: (2024)
Defensive M2S: Training Guardrail Models on Compressed Multi-turn Conversations
by: Kim, Hyunjun
Published: (2026)
by: Kim, Hyunjun
Published: (2026)
Magnet: Multi-turn Tool-use Data Synthesis and Distillation via Graph Translation
by: Yin, Fan, et al.
Published: (2025)
by: Yin, Fan, et al.
Published: (2025)
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
by: Roytburg, Dani, et al.
Published: (2026)
by: Roytburg, Dani, et al.
Published: (2026)
DRInQ: Evaluating Conversational Implicature with Controlled Context Variation
by: Arai, Hirona Jacqueline, et al.
Published: (2026)
by: Arai, Hirona Jacqueline, et al.
Published: (2026)
Broaden your SCOPE! Efficient Multi-turn Conversation Planning for LLMs with Semantic Space
by: Chen, Zhiliang, et al.
Published: (2025)
by: Chen, Zhiliang, et al.
Published: (2025)
Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
by: Gao, Bin, et al.
Published: (2024)
by: Gao, Bin, et al.
Published: (2024)
Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
by: Ibrahim, Lujain, et al.
Published: (2025)
by: Ibrahim, Lujain, et al.
Published: (2025)
A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems
by: Yi, Zihao, et al.
Published: (2024)
by: Yi, Zihao, et al.
Published: (2024)
Similar Items
-
StockGenChaR: A Study on the Evaluation of Large Vision-Language Models on Stock Chart Captioning
by: Qiu, Le, et al.
Published: (2024) -
Empirical Sufficiency Lower Bounds for Language Modeling with Locally-Bootstrapped Semantic Structures
by: Prange, Jakob, et al.
Published: (2023) -
Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from Demonstratives
by: Wang, Yu, et al.
Published: (2026) -
Composing or Not Composing? Towards Distributional Construction Grammars
by: Blache, Philippe, et al.
Published: (2024) -
Learning to Look at the Other Side: A Semantic Probing Study of Word Embeddings in LLMs with Enabled Bidirectional Attention
by: Feng, Zhaoxin, et al.
Published: (2025)