Gespeichert in:
| Hauptverfasser: | Schmidt, David Maria, Schubert, Raoul, Cimiano, Philipp |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2507.21257 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lexicalization Is All You Need: Examining the Impact of Lexical Knowledge in a Compositional QALD System
von: Schmidt, David Maria, et al.
Veröffentlicht: (2024)
von: Schmidt, David Maria, et al.
Veröffentlicht: (2024)
Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
von: Fleisig, Eve, et al.
Veröffentlicht: (2025)
von: Fleisig, Eve, et al.
Veröffentlicht: (2025)
From Argumentation to Deliberation: Perspectivized Stance Vectors for Fine-grained (Dis)agreement Analysis
von: Plenz, Moritz, et al.
Veröffentlicht: (2025)
von: Plenz, Moritz, et al.
Veröffentlicht: (2025)
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
von: Kim, Jiwan, et al.
Veröffentlicht: (2025)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
von: Liu, Yijun, et al.
Veröffentlicht: (2024)
von: Liu, Yijun, et al.
Veröffentlicht: (2024)
Question Answering with LLMs and Learning from Answer Sets
von: Borroto, Manuel, et al.
Veröffentlicht: (2025)
von: Borroto, Manuel, et al.
Veröffentlicht: (2025)
Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
von: Hong, Zijin, et al.
Veröffentlicht: (2025)
von: Hong, Zijin, et al.
Veröffentlicht: (2025)
RepLiQA: A Question-Answering Dataset for Benchmarking LLMs on Unseen Reference Content
von: Monteiro, Joao, et al.
Veröffentlicht: (2024)
von: Monteiro, Joao, et al.
Veröffentlicht: (2024)
I've got the "Answer"! Interpretation of LLMs Hidden States in Question Answering
von: Goloviznina, Valeriya, et al.
Veröffentlicht: (2024)
von: Goloviznina, Valeriya, et al.
Veröffentlicht: (2024)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
von: Yang, Siwei, et al.
Veröffentlicht: (2024)
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
von: Zhou, Wenrui, et al.
Veröffentlicht: (2025)
von: Zhou, Wenrui, et al.
Veröffentlicht: (2025)
MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2024)
von: Li, Shuyue Stella, et al.
Veröffentlicht: (2024)
Crafting Interpretable Embeddings by Asking LLMs Questions
von: Benara, Vinamra, et al.
Veröffentlicht: (2024)
von: Benara, Vinamra, et al.
Veröffentlicht: (2024)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
von: Huang, Kaixuan, et al.
Veröffentlicht: (2025)
von: Huang, Kaixuan, et al.
Veröffentlicht: (2025)
LLMs as Function Approximators: Terminology, Taxonomy, and Questions for Evaluation
von: Schlangen, David
Veröffentlicht: (2024)
von: Schlangen, David
Veröffentlicht: (2024)
Diversity of Thought Improves Reasoning Abilities of LLMs
von: Naik, Ranjita, et al.
Veröffentlicht: (2023)
von: Naik, Ranjita, et al.
Veröffentlicht: (2023)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
von: Liu, Chaoqun, et al.
Veröffentlicht: (2025)
von: Liu, Chaoqun, et al.
Veröffentlicht: (2025)
Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education
von: Henkel, Owen, et al.
Veröffentlicht: (2024)
von: Henkel, Owen, et al.
Veröffentlicht: (2024)
Enhancing Reasoning Abilities of Small LLMs with Cognitive Alignment
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
Beyond Surface Structure: A Causal Assessment of LLMs' Comprehension Ability
von: Han, Yujin, et al.
Veröffentlicht: (2024)
von: Han, Yujin, et al.
Veröffentlicht: (2024)
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects
von: Mustapha, Ahmad, et al.
Veröffentlicht: (2024)
von: Mustapha, Ahmad, et al.
Veröffentlicht: (2024)
Rethinking the Understanding Ability across LLMs through Mutual Information
von: Wang, Shaojie, et al.
Veröffentlicht: (2025)
von: Wang, Shaojie, et al.
Veröffentlicht: (2025)
DORA Explorer: Improving the Exploration Ability of LLMs Without Training
von: Gurjar, Priya, et al.
Veröffentlicht: (2026)
von: Gurjar, Priya, et al.
Veröffentlicht: (2026)
PANDA: Preference Adaptation for Enhancing Domain-Specific Abilities of LLMs
von: Liu, An, et al.
Veröffentlicht: (2024)
von: Liu, An, et al.
Veröffentlicht: (2024)
Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines
von: Seo, Jean, et al.
Veröffentlicht: (2026)
von: Seo, Jean, et al.
Veröffentlicht: (2026)
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
von: Beyer, Anne, et al.
Veröffentlicht: (2024)
von: Beyer, Anne, et al.
Veröffentlicht: (2024)
Perceived Political Bias in LLMs Reduces Persuasive Abilities
von: DiGiuseppe, Matthew, et al.
Veröffentlicht: (2026)
von: DiGiuseppe, Matthew, et al.
Veröffentlicht: (2026)
ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks
von: Schmidt, Jan-Philipp
Veröffentlicht: (2026)
von: Schmidt, Jan-Philipp
Veröffentlicht: (2026)
Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2024)
von: Myrzakhan, Aidar, et al.
Veröffentlicht: (2024)
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs
von: Kim, Minsang, et al.
Veröffentlicht: (2024)
von: Kim, Minsang, et al.
Veröffentlicht: (2024)
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
von: Zhou, Chengliang, et al.
Veröffentlicht: (2025)
von: Zhou, Chengliang, et al.
Veröffentlicht: (2025)
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
von: Fei, Zhaoye, et al.
Veröffentlicht: (2025)
von: Fei, Zhaoye, et al.
Veröffentlicht: (2025)
AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs
von: Fathullah, Yassir, et al.
Veröffentlicht: (2023)
von: Fathullah, Yassir, et al.
Veröffentlicht: (2023)
Can LLMs Ask Good Questions?
von: Zhang, Yueheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yueheng, et al.
Veröffentlicht: (2025)
AbsPyramid: Benchmarking the Abstraction Ability of Language Models with a Unified Entailment Graph
von: Wang, Zhaowei, et al.
Veröffentlicht: (2023)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2023)
Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language Understanding
von: Schmidt, Fabian David, et al.
Veröffentlicht: (2025)
von: Schmidt, Fabian David, et al.
Veröffentlicht: (2025)
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
von: Huang, Jen-tse, et al.
Veröffentlicht: (2024)
von: Huang, Jen-tse, et al.
Veröffentlicht: (2024)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
LLMs in Interpreting Legal Documents
von: Corbo, Simone
Veröffentlicht: (2025)
von: Corbo, Simone
Veröffentlicht: (2025)
Ähnliche Einträge
-
Lexicalization Is All You Need: Examining the Impact of Lexical Knowledge in a Compositional QALD System
von: Schmidt, David Maria, et al.
Veröffentlicht: (2024) -
Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
von: Fleisig, Eve, et al.
Veröffentlicht: (2025) -
From Argumentation to Deliberation: Perspectivized Stance Vectors for Fine-grained (Dis)agreement Analysis
von: Plenz, Moritz, et al.
Veröffentlicht: (2025) -
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
von: Kim, Jiwan, et al.
Veröffentlicht: (2025) -
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
von: Liu, Yijun, et al.
Veröffentlicht: (2024)