Metric assessment protocol in the context of answer fluctuation on MCQ tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Goliakova, Ekaterina, Renard, Xavier, Lesot, Marie-Jeanne, Laugel, Thibault, Marsala, Christophe, Detyniecki, Marcin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Alignment Reduces Expressed but Not Encoded Gender Bias: A Unified Framework and Study
by: Bouchouchi, Nour, et al.
Published: (2026)
by: Bouchouchi, Nour, et al.
Published: (2026)
Understanding Prediction Discrepancies in Machine Learning Classifiers
by: Renard, Xavier, et al.
Published: (2021)
by: Renard, Xavier, et al.
Published: (2021)
Dynamic Interpretability for Model Comparison via Decision Rules
by: Rida, Adam, et al.
Published: (2023)
by: Rida, Adam, et al.
Published: (2023)
Why do explanations fail? A typology and discussion on failures in XAI
by: Bove, Clara, et al.
Published: (2024)
by: Bove, Clara, et al.
Published: (2024)
Post-processing fairness with minimal changes
by: Di Gennaro, Federico, et al.
Published: (2024)
by: Di Gennaro, Federico, et al.
Published: (2024)
SAKE: Steering Activations for Knowledge Editing
by: Scialanga, Marco, et al.
Published: (2025)
by: Scialanga, Marco, et al.
Published: (2025)
ACT: Agentic Classification Tree
by: Grari, Vincent, et al.
Published: (2025)
by: Grari, Vincent, et al.
Published: (2025)
Aggregative Semantics for Quantitative Bipolar Argumentation Frameworks
by: Munro, Yann, et al.
Published: (2026)
by: Munro, Yann, et al.
Published: (2026)
OptiGrad: A Fair and more Efficient Price Elasticity Optimization via a Gradient Based Learning
by: Grari, Vincent, et al.
Published: (2024)
by: Grari, Vincent, et al.
Published: (2024)
Controlled Model Debiasing through Minimal and Interpretable Updates
by: Di Gennaro, Federico, et al.
Published: (2025)
by: Di Gennaro, Federico, et al.
Published: (2025)
When mitigating bias is unfair: multiplicity and arbitrariness in algorithmic group fairness
by: Krco, Natasa, et al.
Published: (2023)
by: Krco, Natasa, et al.
Published: (2023)
An action language-based formalisation of an abstract argumentation framework
by: Munro, Yann, et al.
Published: (2024)
by: Munro, Yann, et al.
Published: (2024)
Reasoning and Sampling-Augmented MCQ Difficulty Prediction via LLMs
by: Feng, Wanyong, et al.
Published: (2025)
by: Feng, Wanyong, et al.
Published: (2025)
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ
by: Allard, Marc-Antoine, et al.
Published: (2024)
by: Allard, Marc-Antoine, et al.
Published: (2024)
Agentic Adversarial QA for Improving Domain-Specific LLMs
by: Grari, Vincent, et al.
Published: (2026)
by: Grari, Vincent, et al.
Published: (2026)
Buggy rule diagnosis for combined steps through final answer evaluation in stepwise tasks
by: van der Hoek, Gerben, et al.
Published: (2025)
by: van der Hoek, Gerben, et al.
Published: (2025)
From Demographics to Survey Anchors: Evaluating LLM Agents for Modeling Retirement Attitudes
by: Garzón, Rubén, et al.
Published: (2026)
by: Garzón, Rubén, et al.
Published: (2026)
AutoMCQ -- Automatically Generate Code Comprehension Questions using GenAI
by: Goodfellow, Martin, et al.
Published: (2025)
by: Goodfellow, Martin, et al.
Published: (2025)
Extract, Match, and Score: An Evaluation Paradigm for Long Question-context-answer Triplets in Financial Analysis
by: Hu, Bo, et al.
Published: (2025)
by: Hu, Bo, et al.
Published: (2025)
When Numbers Tell Half the Story: Human-Metric Alignment in Topic Model Evaluation
by: Prouteau, Thibault, et al.
Published: (2026)
by: Prouteau, Thibault, et al.
Published: (2026)
Integrating Temporality and Causality into Acyclic Argumentation Frameworks using a Transition System
by: Munro, Y., et al.
Published: (2023)
by: Munro, Y., et al.
Published: (2023)
MCQ Difficulty Prediction via Modeling Learner Heterogeneity Using Data-Driven Cognitive Profiling
by: Krishnan, Dhriti, et al.
Published: (2026)
by: Krishnan, Dhriti, et al.
Published: (2026)
CMI-MTL: Cross-Mamba interaction based multi-task learning for medical visual question answering
by: Jin, Qiangguo, et al.
Published: (2025)
by: Jin, Qiangguo, et al.
Published: (2025)
What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
by: Maar, Jim, et al.
Published: (2026)
by: Maar, Jim, et al.
Published: (2026)
Does quantization affect models' performance on long-context tasks?
by: Mekala, Anmol, et al.
Published: (2025)
by: Mekala, Anmol, et al.
Published: (2025)
Five questions and answers about artificial intelligence
by: Prieto, Alberto, et al.
Published: (2024)
by: Prieto, Alberto, et al.
Published: (2024)
Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
by: Bhatti, Hunzalah Hassan, et al.
Published: (2025)
by: Bhatti, Hunzalah Hassan, et al.
Published: (2025)
Orchestrating LLM Agents for Scientific Research: A Pilot Study of Multiple Choice Question (MCQ) Generation and Evaluation
by: An, Yuan
Published: (2026)
by: An, Yuan
Published: (2026)
Diagnostics of cognitive failures in multi-agent expert systems using dynamic evaluation protocols and subsequent mutation of the processing context
by: Sorstkins, Andrejs, et al.
Published: (2025)
by: Sorstkins, Andrejs, et al.
Published: (2025)
Enhancing Knowledge Graph Construction: Evaluating with Emphasis on Hallucination, Omission, and Graph Similarity Metrics
by: Ghanem, Hussam, et al.
Published: (2025)
by: Ghanem, Hussam, et al.
Published: (2025)
What do your logits know? (The answer may surprise you!)
by: Fedzechkina, Masha, et al.
Published: (2026)
by: Fedzechkina, Masha, et al.
Published: (2026)
FLIRT: Feedback Loop In-context Red Teaming
by: Mehrabi, Ninareh, et al.
Published: (2023)
by: Mehrabi, Ninareh, et al.
Published: (2023)
Online inductive learning from answer sets for efficient reinforcement learning exploration
by: Veronese, Celeste, et al.
Published: (2025)
by: Veronese, Celeste, et al.
Published: (2025)
Can MLLMs generate human-like feedback in grading multimodal short answers?
by: Sil, Pritam, et al.
Published: (2024)
by: Sil, Pritam, et al.
Published: (2024)
Hypothetical answers to continuous queries over data streams
by: Cruz-Filipe, Luís, et al.
Published: (2019)
by: Cruz-Filipe, Luís, et al.
Published: (2019)
Cross-Modal Transferable Image-to-Video Attack on Video Quality Metrics
by: Gotin, Georgii, et al.
Published: (2025)
by: Gotin, Georgii, et al.
Published: (2025)
Cross-Entropy Games and Frost Training
by: Renard, Arthur, et al.
Published: (2026)
by: Renard, Arthur, et al.
Published: (2026)
Agribot: agriculture-specific question answer system
by: Jain, Naman, et al.
Published: (2025)
by: Jain, Naman, et al.
Published: (2025)
Why does in-context learning fail sometimes? Evaluating in-context learning on open and closed questions
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Detecting Sexual Content at the Sentence Level in First Millennium Latin Texts
by: Clérice, Thibault
Published: (2023)
by: Clérice, Thibault
Published: (2023)
Similar Items
-
Alignment Reduces Expressed but Not Encoded Gender Bias: A Unified Framework and Study
by: Bouchouchi, Nour, et al.
Published: (2026) -
Understanding Prediction Discrepancies in Machine Learning Classifiers
by: Renard, Xavier, et al.
Published: (2021) -
Dynamic Interpretability for Model Comparison via Decision Rules
by: Rida, Adam, et al.
Published: (2023) -
Why do explanations fail? A typology and discussion on failures in XAI
by: Bove, Clara, et al.
Published: (2024) -
Post-processing fairness with minimal changes
by: Di Gennaro, Federico, et al.
Published: (2024)