Do Chatbot LLMs Talk Too Much? The YapBench Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Borisov, Vadim, Gröger, Michael, Mikhael, Mina, Schreiber, Richard H. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Open Artificial Knowledge
von: Borisov, Vadim, et al.
Veröffentlicht: (2024)
von: Borisov, Vadim, et al.
Veröffentlicht: (2024)
What Ails Generative Structure-based Drug Design: Expressivity is Too Little or Too Much?
von: Karczewski, Rafał, et al.
Veröffentlicht: (2024)
von: Karczewski, Rafał, et al.
Veröffentlicht: (2024)
How Much Is Too Much? Adaptive, Context-Aware Risk Detection in Naturalistic Driving
von: Kalantari, Amir Hossein, et al.
Veröffentlicht: (2025)
von: Kalantari, Amir Hossein, et al.
Veröffentlicht: (2025)
How Much is Too Much? Exploring LoRA Rank Trade-offs for Retaining Knowledge and Domain Robustness
von: Rathore, Darshita, et al.
Veröffentlicht: (2025)
von: Rathore, Darshita, et al.
Veröffentlicht: (2025)
How Much Reasoning Do Retrieval-Augmented Models Add beyond LLMs? A Benchmarking Framework for Multi-Hop Inference over Hybrid Knowledge
von: Lin, Junhong, et al.
Veröffentlicht: (2026)
von: Lin, Junhong, et al.
Veröffentlicht: (2026)
Interpreting Microbiome Relative Abundance Data Using Symbolic Regression
von: Haldar, Swagatam, et al.
Veröffentlicht: (2024)
von: Haldar, Swagatam, et al.
Veröffentlicht: (2024)
DNR Bench: Benchmarking Over-Reasoning in Reasoning LLMs
von: Hashemi, Masoud, et al.
Veröffentlicht: (2025)
von: Hashemi, Masoud, et al.
Veröffentlicht: (2025)
OSS-Bench: Benchmark Generator for Coding LLMs
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yuancheng, et al.
Veröffentlicht: (2025)
MergeBench: A Benchmark for Merging Domain-Specialized LLMs
von: He, Yifei, et al.
Veröffentlicht: (2025)
von: He, Yifei, et al.
Veröffentlicht: (2025)
MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness
von: Hathidara, Ashutosh, et al.
Veröffentlicht: (2026)
von: Hathidara, Ashutosh, et al.
Veröffentlicht: (2026)
ThrowBench: Benchmarking LLMs by Predicting Runtime Exceptions
von: Prenner, Julian Aron, et al.
Veröffentlicht: (2025)
von: Prenner, Julian Aron, et al.
Veröffentlicht: (2025)
CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs
von: Zhou, Yu, et al.
Veröffentlicht: (2024)
von: Zhou, Yu, et al.
Veröffentlicht: (2024)
SortBench: Benchmarking LLMs based on their ability to sort lists
von: Herbold, Steffen
Veröffentlicht: (2025)
von: Herbold, Steffen
Veröffentlicht: (2025)
'Hello, World!': Making GNNs Talk with LLMs
von: Kim, Sunwoo, et al.
Veröffentlicht: (2025)
von: Kim, Sunwoo, et al.
Veröffentlicht: (2025)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
von: Wang, Erchi, et al.
Veröffentlicht: (2026)
von: Wang, Erchi, et al.
Veröffentlicht: (2026)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
von: Lin, Zicheng, et al.
Veröffentlicht: (2024)
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
von: Zhou, Yujun, et al.
Veröffentlicht: (2024)
XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
Do Large Language Models Know How Much They Know?
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
SensorBench: Benchmarking LLMs in Coding-Based Sensor Processing
von: Quan, Pengrui, et al.
Veröffentlicht: (2024)
von: Quan, Pengrui, et al.
Veröffentlicht: (2024)
YRC-Bench: A Benchmark for Learning to Coordinate with Experts
von: Danesh, Mohamad H., et al.
Veröffentlicht: (2025)
von: Danesh, Mohamad H., et al.
Veröffentlicht: (2025)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
von: Li, Xin, et al.
Veröffentlicht: (2025)
von: Li, Xin, et al.
Veröffentlicht: (2025)
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents
von: De Brouwer, Edward, et al.
Veröffentlicht: (2026)
von: De Brouwer, Edward, et al.
Veröffentlicht: (2026)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
DM-Bench: Benchmarking LLMs for Personalized Decision Making in Diabetes Management
von: Cardei, Maria Ana, et al.
Veröffentlicht: (2025)
von: Cardei, Maria Ana, et al.
Veröffentlicht: (2025)
V4FinBench: Benchmarking Tabular Foundation Models, LLMs, and Standard Methods on Corporate Bankruptcy Prediction
von: Kostrzewa, Marcin, et al.
Veröffentlicht: (2026)
von: Kostrzewa, Marcin, et al.
Veröffentlicht: (2026)
SemBench: A Benchmark for Semantic Query Processing Engines
von: Lao, Jiale, et al.
Veröffentlicht: (2025)
von: Lao, Jiale, et al.
Veröffentlicht: (2025)
seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
von: Ramezanali, Mohammad, et al.
Veröffentlicht: (2025)
von: Ramezanali, Mohammad, et al.
Veröffentlicht: (2025)
BertaQA: How Much Do Language Models Know About Local Culture?
von: Etxaniz, Julen, et al.
Veröffentlicht: (2024)
von: Etxaniz, Julen, et al.
Veröffentlicht: (2024)
TopoBench: A Framework for Benchmarking Topological Deep Learning
von: Telyatnikov, Lev, et al.
Veröffentlicht: (2024)
von: Telyatnikov, Lev, et al.
Veröffentlicht: (2024)
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
von: Sharifloo, Amir Molzam, et al.
Veröffentlicht: (2025)
TaoBench: Do Automated Theorem Prover LLMs Generalize Beyond MathLib?
von: Taylor, Alexander K, et al.
Veröffentlicht: (2026)
von: Taylor, Alexander K, et al.
Veröffentlicht: (2026)
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
von: Bi, Zhen, et al.
Veröffentlicht: (2026)
von: Bi, Zhen, et al.
Veröffentlicht: (2026)
LikeBench: Evaluating Subjective Likability in LLMs for Personalization
von: Rahman, Md Awsafur, et al.
Veröffentlicht: (2025)
von: Rahman, Md Awsafur, et al.
Veröffentlicht: (2025)
Language Game: Talking to Non-Human Systems
von: Zhang, Yanbo, et al.
Veröffentlicht: (2026)
von: Zhang, Yanbo, et al.
Veröffentlicht: (2026)
LJ-Bench: Ontology-Based Benchmark for U.S. Crime
von: Tseng, Hung Yun, et al.
Veröffentlicht: (2026)
von: Tseng, Hung Yun, et al.
Veröffentlicht: (2026)
Calibrating LLMs with Information-Theoretic Evidential Deep Learning
von: Li, Yawei, et al.
Veröffentlicht: (2025)
von: Li, Yawei, et al.
Veröffentlicht: (2025)
ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
von: Yu, Han, et al.
Veröffentlicht: (2025)
von: Yu, Han, et al.
Veröffentlicht: (2025)
GC-Bench: An Open and Unified Benchmark for Graph Condensation
von: Sun, Qingyun, et al.
Veröffentlicht: (2024)
von: Sun, Qingyun, et al.
Veröffentlicht: (2024)
On the Universality of Transformer Architectures; How Much Attention Is Enough?
von: Abbasi, Amirreza, et al.
Veröffentlicht: (2025)
von: Abbasi, Amirreza, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Open Artificial Knowledge
von: Borisov, Vadim, et al.
Veröffentlicht: (2024) -
What Ails Generative Structure-based Drug Design: Expressivity is Too Little or Too Much?
von: Karczewski, Rafał, et al.
Veröffentlicht: (2024) -
How Much Is Too Much? Adaptive, Context-Aware Risk Detection in Naturalistic Driving
von: Kalantari, Amir Hossein, et al.
Veröffentlicht: (2025) -
How Much is Too Much? Exploring LoRA Rank Trade-offs for Retaining Knowledge and Domain Robustness
von: Rathore, Darshita, et al.
Veröffentlicht: (2025) -
How Much Reasoning Do Retrieval-Augmented Models Add beyond LLMs? A Benchmarking Framework for Multi-Hop Inference over Hybrid Knowledge
von: Lin, Junhong, et al.
Veröffentlicht: (2026)