Gespeichert in:
| Hauptverfasser: | Meyer, Gérôme, Breuer, Philip, Fürst, Jonathan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2409.18596 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Focusing on Students, not Machines: Grounded Question Generation and Automated Answer Grading
von: Meyer, Gérôme, et al.
Veröffentlicht: (2025)
von: Meyer, Gérôme, et al.
Veröffentlicht: (2025)
Statistical Comparative Analysis of Semantic Similarities and Model Transferability Across Datasets for Short Answer Grading
von: Bonthu, Sridevi, et al.
Veröffentlicht: (2025)
von: Bonthu, Sridevi, et al.
Veröffentlicht: (2025)
Bench360: Benchmarking Local LLM Inference from 360 Degrees
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
von: Yao, Siyang, et al.
Veröffentlicht: (2026)
von: Yao, Siyang, et al.
Veröffentlicht: (2026)
Proving that Cryptic Crossword Clue Answers are Correct
von: Andrews, Martin, et al.
Veröffentlicht: (2024)
von: Andrews, Martin, et al.
Veröffentlicht: (2024)
When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMs
von: Zagribelnyy, Bogdan, et al.
Veröffentlicht: (2026)
von: Zagribelnyy, Bogdan, et al.
Veröffentlicht: (2026)
Measuring and Reducing LLM Hallucination without Gold-Standard Answers
von: Wei, Jiaheng, et al.
Veröffentlicht: (2024)
von: Wei, Jiaheng, et al.
Veröffentlicht: (2024)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
von: Chandak, Nikhil, et al.
Veröffentlicht: (2025)
Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
von: Baan, Joris, et al.
Veröffentlicht: (2026)
von: Baan, Joris, et al.
Veröffentlicht: (2026)
SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
von: Kirchhof, Michael, et al.
Veröffentlicht: (2025)
von: Kirchhof, Michael, et al.
Veröffentlicht: (2025)
Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
von: Bennion, Jonathan, et al.
Veröffentlicht: (2025)
von: Bennion, Jonathan, et al.
Veröffentlicht: (2025)
Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation
von: Qi, Jirui, et al.
Veröffentlicht: (2024)
von: Qi, Jirui, et al.
Veröffentlicht: (2024)
Rewarding Intellectual Humility Learning When Not To Answer In Large Language Models
von: Jha, Abha, et al.
Veröffentlicht: (2026)
von: Jha, Abha, et al.
Veröffentlicht: (2026)
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
von: Hu, Zhengyu, et al.
Veröffentlicht: (2026)
von: Hu, Zhengyu, et al.
Veröffentlicht: (2026)
Train Once, Answer All: Many Pretraining Experiments for the Cost of One
von: Bordt, Sebastian, et al.
Veröffentlicht: (2025)
von: Bordt, Sebastian, et al.
Veröffentlicht: (2025)
MultiQ&A: An Analysis in Measuring Robustness via Automated Crowdsourcing of Question Perturbations and Answers
von: Cho, Nicole, et al.
Veröffentlicht: (2025)
von: Cho, Nicole, et al.
Veröffentlicht: (2025)
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions
von: Li, Ruizhe, et al.
Veröffentlicht: (2024)
von: Li, Ruizhe, et al.
Veröffentlicht: (2024)
Explicit Diversity Conditions for Effective Question Answer Generation with Large Language Models
von: Yadav, Vikas, et al.
Veröffentlicht: (2024)
von: Yadav, Vikas, et al.
Veröffentlicht: (2024)
The Challenge of Achieving Attributability in Multilingual Table-to-Text Generation with Question-Answer Blueprints
von: Haussmann, Aden
Veröffentlicht: (2025)
von: Haussmann, Aden
Veröffentlicht: (2025)
Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2025)
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2025)
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
von: Li, Kenneth, et al.
Veröffentlicht: (2023)
von: Li, Kenneth, et al.
Veröffentlicht: (2023)
ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
von: Wang, Zhilin, et al.
Veröffentlicht: (2025)
When Does a Language Model Commit? A Finite-Answer Theory of Pre-Verbalization Commitment
von: Zhang, Long, et al.
Veröffentlicht: (2026)
von: Zhang, Long, et al.
Veröffentlicht: (2026)
SuRe: Summarizing Retrievals using Answer Candidates for Open-domain QA of LLMs
von: Kim, Jaehyung, et al.
Veröffentlicht: (2024)
von: Kim, Jaehyung, et al.
Veröffentlicht: (2024)
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2025)
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
von: Yan, Tianyi Lorena, et al.
Veröffentlicht: (2025)
von: Yan, Tianyi Lorena, et al.
Veröffentlicht: (2025)
ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
von: Karger, Ezra, et al.
Veröffentlicht: (2024)
von: Karger, Ezra, et al.
Veröffentlicht: (2024)
A Careful Examination of Large Language Model Performance on Grade School Arithmetic
von: Zhang, Hugh, et al.
Veröffentlicht: (2024)
von: Zhang, Hugh, et al.
Veröffentlicht: (2024)
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
von: Tian, Yijun, et al.
Veröffentlicht: (2024)
von: Tian, Yijun, et al.
Veröffentlicht: (2024)
MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters
von: Dada, Amin, et al.
Veröffentlicht: (2025)
von: Dada, Amin, et al.
Veröffentlicht: (2025)
From Flat to Structural: Enhancing Automated Short Answer Grading with GraphRAG
von: Chu, Yucheng, et al.
Veröffentlicht: (2026)
von: Chu, Yucheng, et al.
Veröffentlicht: (2026)
Crystal-KV: Efficient KV Cache Management for Chain-of-Thought LLMs via Answer-First Principle
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
von: Zhao, Shuai, et al.
Veröffentlicht: (2025)
CORI: CJKV Benchmark with Romanization Integration -- A step towards Cross-lingual Transfer Beyond Textual Scripts
von: Nguyen, Hoang H., et al.
Veröffentlicht: (2024)
von: Nguyen, Hoang H., et al.
Veröffentlicht: (2024)
Towards A Unified View of Answer Calibration for Multi-Step Reasoning
von: Deng, Shumin, et al.
Veröffentlicht: (2023)
von: Deng, Shumin, et al.
Veröffentlicht: (2023)
Implicit Probabilistic Reasoning Does Not Reflect Explicit Answers in Large Language Models
von: Mondal, Manuel, et al.
Veröffentlicht: (2024)
von: Mondal, Manuel, et al.
Veröffentlicht: (2024)
Fantastic Bugs and Where to Find Them in AI Benchmarks
von: Truong, Sang, et al.
Veröffentlicht: (2025)
von: Truong, Sang, et al.
Veröffentlicht: (2025)
Short-form Text Rewriting with Phi Silica
von: Tadimeti, Divya, et al.
Veröffentlicht: (2026)
von: Tadimeti, Divya, et al.
Veröffentlicht: (2026)
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
von: Li, Haoming, et al.
Veröffentlicht: (2025)
von: Li, Haoming, et al.
Veröffentlicht: (2025)
DiffuSpeech: Silent Thought, Spoken Answer via Unified Speech-Text Diffusion
von: Lou, Yuxuan, et al.
Veröffentlicht: (2026)
von: Lou, Yuxuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Focusing on Students, not Machines: Grounded Question Generation and Automated Answer Grading
von: Meyer, Gérôme, et al.
Veröffentlicht: (2025) -
Statistical Comparative Analysis of Semantic Similarities and Model Transferability Across Datasets for Short Answer Grading
von: Bonthu, Sridevi, et al.
Veröffentlicht: (2025) -
Bench360: Benchmarking Local LLM Inference from 360 Degrees
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025) -
When Answers Stray from Questions: Hallucination Detection via Question-Answer Orthogonal Decomposition
von: Yao, Siyang, et al.
Veröffentlicht: (2026) -
Proving that Cryptic Crossword Clue Answers are Correct
von: Andrews, Martin, et al.
Veröffentlicht: (2024)