Who can we trust? LLM-as-a-jury for Comparative Assessment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qian, Mengjie, Sun, Guangzhi, Gales, Mark J. F., Knill, Kate M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Natural Language-based Assessment of L2 Oral Proficiency using LLMs
von: Bannò, Stefano, et al.
Veröffentlicht: (2025)
von: Bannò, Stefano, et al.
Veröffentlicht: (2025)
Exploiting the English Vocabulary Profile for L2 word-level vocabulary assessment with LLMs
von: Bannò, Stefano, et al.
Veröffentlicht: (2025)
von: Bannò, Stefano, et al.
Veröffentlicht: (2025)
Grammatical Error Feedback: An Implicit Evaluation Approach
von: Bannò, Stefano, et al.
Veröffentlicht: (2024)
von: Bannò, Stefano, et al.
Veröffentlicht: (2024)
Data Augmentation for Spoken Grammatical Error Correction
von: Karanasou, Penny, et al.
Veröffentlicht: (2025)
von: Karanasou, Penny, et al.
Veröffentlicht: (2025)
Towards End-to-End Spoken Grammatical Error Correction
von: Bannò, Stefano, et al.
Veröffentlicht: (2023)
von: Bannò, Stefano, et al.
Veröffentlicht: (2023)
Learn and Don't Forget: Adding a New Language to ASR Foundation Models
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
End-to-End Spoken Grammatical Error Correction
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
Can we trust the evaluation on ChatGPT?
von: Aiyappa, Rachith, et al.
Veröffentlicht: (2023)
von: Aiyappa, Rachith, et al.
Veröffentlicht: (2023)
Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge
von: Fathullah, Yassir, et al.
Veröffentlicht: (2025)
von: Fathullah, Yassir, et al.
Veröffentlicht: (2025)
SkillAggregation: Reference-free LLM-Dependent Aggregation
von: Sun, Guangzhi, et al.
Veröffentlicht: (2024)
von: Sun, Guangzhi, et al.
Veröffentlicht: (2024)
Speak & Improve Corpus 2025: an L2 English Speech Corpus for Language Assessment and Feedback
von: Knill, Kate, et al.
Veröffentlicht: (2024)
von: Knill, Kate, et al.
Veröffentlicht: (2024)
Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
von: Sun, Guangzhi, et al.
Veröffentlicht: (2025)
von: Sun, Guangzhi, et al.
Veröffentlicht: (2025)
ASR Error Correction using Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2024)
von: Ma, Rao, et al.
Veröffentlicht: (2024)
Speaker Retrieval in the Wild: Challenges, Effectiveness and Robustness
von: Loweimi, Erfan, et al.
Veröffentlicht: (2025)
von: Loweimi, Erfan, et al.
Veröffentlicht: (2025)
Large Language Models in Fire Engineering: An Examination of Technical Questions Against Domain Knowledge
von: Hostetter, Haley, et al.
Veröffentlicht: (2024)
von: Hostetter, Haley, et al.
Veröffentlicht: (2024)
Towards Self-Referential Analytic Assessment: A Profile-Based Approach to L2 Writing Evaluation with LLMs
von: Bannò, Stefano, et al.
Veröffentlicht: (2026)
von: Bannò, Stefano, et al.
Veröffentlicht: (2026)
StructEval: Deepen and Broaden Large Language Model Assessment via Structured Evaluation
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
Bayesian WeakS-to-Strong from Text Classification to Generation
von: Cui, Ziyun, et al.
Veröffentlicht: (2024)
von: Cui, Ziyun, et al.
Veröffentlicht: (2024)
How can we trust opaque systems? Criteria for robust explanations in XAI
von: Boge, Florian J., et al.
Veröffentlicht: (2025)
von: Boge, Florian J., et al.
Veröffentlicht: (2025)
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
von: Tsujimura, Hikaru, et al.
Veröffentlicht: (2025)
von: Tsujimura, Hikaru, et al.
Veröffentlicht: (2025)
The Boy Who Survived: Removing Harry Potter from an LLM is harder than reported
von: Shostack, Adam
Veröffentlicht: (2024)
von: Shostack, Adam
Veröffentlicht: (2024)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
Assessment of L2 Oral Proficiency using Speech Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
Who Benefits From Sinus Surgery? Comparing Generative AI and Supervised Machine Learning for Predicting Surgical Outcomes in Chronic Rhinosinusitis
von: Chowdhury, Sayeed Shafayet, et al.
Veröffentlicht: (2026)
von: Chowdhury, Sayeed Shafayet, et al.
Veröffentlicht: (2026)
Cross-Lingual Transfer Learning for Speech Translation
von: Ma, Rao, et al.
Veröffentlicht: (2024)
von: Ma, Rao, et al.
Veröffentlicht: (2024)
An LLM Feature-based Framework for Dialogue Constructiveness Assessment
von: Zhou, Lexin, et al.
Veröffentlicht: (2024)
von: Zhou, Lexin, et al.
Veröffentlicht: (2024)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
von: Islam, Tunazzina
Veröffentlicht: (2026)
von: Islam, Tunazzina
Veröffentlicht: (2026)
Comparative Analysis of Demonstration Selection Algorithms for LLM In-Context Learning
von: Shu, Dong, et al.
Veröffentlicht: (2024)
von: Shu, Dong, et al.
Veröffentlicht: (2024)
Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
Aligning Backchannel and Dialogue Context Representations via Contrastive LLM Fine-Tuning
von: Qian, Livia, et al.
Veröffentlicht: (2026)
von: Qian, Livia, et al.
Veröffentlicht: (2026)
Question Difficulty Ranking for Multiple-Choice Reading Comprehension
von: Raina, Vatsal, et al.
Veröffentlicht: (2024)
von: Raina, Vatsal, et al.
Veröffentlicht: (2024)
Efficient Sample-Specific Encoder Perturbations
von: Fathullah, Yassir, et al.
Veröffentlicht: (2024)
von: Fathullah, Yassir, et al.
Veröffentlicht: (2024)
Flow-of-Options: Diversified and Improved LLM Reasoning by Thinking Through Options
von: Nair, Lakshmi, et al.
Veröffentlicht: (2025)
von: Nair, Lakshmi, et al.
Veröffentlicht: (2025)
When can transformers reason with abstract symbols?
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2023)
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2023)
Self-Improving LLM Agents at Test-Time
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2025)
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2025)
Enhancing LLM Evaluations: The Garbling Trick
von: Bradley, William F.
Veröffentlicht: (2024)
von: Bradley, William F.
Veröffentlicht: (2024)
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2026)
Enhancing Marker Scoring Accuracy through Ordinal Confidence Modelling in Educational Assessments
von: Chakravarty, Abhirup, et al.
Veröffentlicht: (2025)
von: Chakravarty, Abhirup, et al.
Veröffentlicht: (2025)
Large Language Models can be Strong Self-Detoxifiers
von: Ko, Ching-Yun, et al.
Veröffentlicht: (2024)
von: Ko, Ching-Yun, et al.
Veröffentlicht: (2024)
New Encoders for German Trained from Scratch: Comparing ModernGBERT with Converted LLM2Vec Models
von: Wunderle, Julia, et al.
Veröffentlicht: (2025)
von: Wunderle, Julia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Natural Language-based Assessment of L2 Oral Proficiency using LLMs
von: Bannò, Stefano, et al.
Veröffentlicht: (2025) -
Exploiting the English Vocabulary Profile for L2 word-level vocabulary assessment with LLMs
von: Bannò, Stefano, et al.
Veröffentlicht: (2025) -
Grammatical Error Feedback: An Implicit Evaluation Approach
von: Bannò, Stefano, et al.
Veröffentlicht: (2024) -
Data Augmentation for Spoken Grammatical Error Correction
von: Karanasou, Penny, et al.
Veröffentlicht: (2025) -
Towards End-to-End Spoken Grammatical Error Correction
von: Bannò, Stefano, et al.
Veröffentlicht: (2023)