Finetuning LLMs for Comparative Assessment Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Raina, Vatsal, Liusie, Adian, Gales, Mark |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons
von: Liusie, Adian, et al.
Veröffentlicht: (2024)
von: Liusie, Adian, et al.
Veröffentlicht: (2024)
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
von: Liusie, Adian, et al.
Veröffentlicht: (2023)
von: Liusie, Adian, et al.
Veröffentlicht: (2023)
WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models
von: Molenda, Piotr, et al.
Veröffentlicht: (2024)
von: Molenda, Piotr, et al.
Veröffentlicht: (2024)
Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models
von: Liusie, Adian, et al.
Veröffentlicht: (2024)
von: Liusie, Adian, et al.
Veröffentlicht: (2024)
Question-Based Retrieval using Atomic Units for Enterprise RAG
von: Raina, Vatsal, et al.
Veröffentlicht: (2024)
von: Raina, Vatsal, et al.
Veröffentlicht: (2024)
Question Difficulty Ranking for Multiple-Choice Reading Comprehension
von: Raina, Vatsal, et al.
Veröffentlicht: (2024)
von: Raina, Vatsal, et al.
Veröffentlicht: (2024)
An Information-Theoretic Approach to Analyze NLP Classification Tasks
von: Wang, Luran, et al.
Veröffentlicht: (2024)
von: Wang, Luran, et al.
Veröffentlicht: (2024)
Investigating the Emergent Audio Classification Ability of ASR Foundation Models
von: Ma, Rao, et al.
Veröffentlicht: (2023)
von: Ma, Rao, et al.
Veröffentlicht: (2023)
Is it Possible to Modify Text to a Target Readability Level? An Initial Investigation Using Zero-Shot Large Language Models
von: Farajidizaji, Asma, et al.
Veröffentlicht: (2023)
von: Farajidizaji, Asma, et al.
Veröffentlicht: (2023)
CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models
von: Sun, Guangzhi, et al.
Veröffentlicht: (2024)
von: Sun, Guangzhi, et al.
Veröffentlicht: (2024)
Zero-shot Audio Topic Reranking using Large Language Models
von: Qian, Mengjie, et al.
Veröffentlicht: (2023)
von: Qian, Mengjie, et al.
Veröffentlicht: (2023)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM
von: Lu, Xiaoding, et al.
Veröffentlicht: (2024)
von: Lu, Xiaoding, et al.
Veröffentlicht: (2024)
LLM Task Interference: An Initial Study on the Impact of Task-Switch in Conversational History
von: Gupta, Akash, et al.
Veröffentlicht: (2024)
von: Gupta, Akash, et al.
Veröffentlicht: (2024)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
Embedding the Teacher: Distilling vLLM Preferences for Scalable Image Retrieval
von: He, Eric, et al.
Veröffentlicht: (2025)
von: He, Eric, et al.
Veröffentlicht: (2025)
Towards Self-Referential Analytic Assessment: A Profile-Based Approach to L2 Writing Evaluation with LLMs
von: Bannò, Stefano, et al.
Veröffentlicht: (2026)
von: Bannò, Stefano, et al.
Veröffentlicht: (2026)
Probing the Limits of Stylistic Alignment in Vision-Language Models
von: Farajidizaji, Asma, et al.
Veröffentlicht: (2025)
von: Farajidizaji, Asma, et al.
Veröffentlicht: (2025)
Exploiting the English Vocabulary Profile for L2 word-level vocabulary assessment with LLMs
von: Bannò, Stefano, et al.
Veröffentlicht: (2025)
von: Bannò, Stefano, et al.
Veröffentlicht: (2025)
Exploiting the English Grammar Profile for L2 grammatical analysis with LLMs
von: Bannò, Stefano, et al.
Veröffentlicht: (2026)
von: Bannò, Stefano, et al.
Veröffentlicht: (2026)
Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
von: Ruan, Zhiwen, et al.
Veröffentlicht: (2025)
Who can we trust? LLM-as-a-jury for Comparative Assessment
von: Qian, Mengjie, et al.
Veröffentlicht: (2026)
von: Qian, Mengjie, et al.
Veröffentlicht: (2026)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
Natural Language-based Assessment of L2 Oral Proficiency using LLMs
von: Bannò, Stefano, et al.
Veröffentlicht: (2025)
von: Bannò, Stefano, et al.
Veröffentlicht: (2025)
Improving Task Diversity in Label Efficient Supervised Finetuning of LLMs
von: Arabelly, Abhinav, et al.
Veröffentlicht: (2025)
von: Arabelly, Abhinav, et al.
Veröffentlicht: (2025)
Speak & Improve Corpus 2025: an L2 English Speech Corpus for Language Assessment and Feedback
von: Knill, Kate, et al.
Veröffentlicht: (2024)
von: Knill, Kate, et al.
Veröffentlicht: (2024)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2026)
von: Vasudeva, Bhavya, et al.
Veröffentlicht: (2026)
Locking Down the Finetuned LLMs Safety
von: Zhu, Minjun, et al.
Veröffentlicht: (2024)
von: Zhu, Minjun, et al.
Veröffentlicht: (2024)
Understanding the Effects of Domain Finetuning on LLMs
von: Tanwar, Eshaan, et al.
Veröffentlicht: (2025)
von: Tanwar, Eshaan, et al.
Veröffentlicht: (2025)
Efficient Sample-Specific Encoder Perturbations
von: Fathullah, Yassir, et al.
Veröffentlicht: (2024)
von: Fathullah, Yassir, et al.
Veröffentlicht: (2024)
A Survey of Prompt Engineering Methods in Large Language Models for Different NLP Tasks
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
von: Vatsal, Shubham, et al.
Veröffentlicht: (2024)
Speak & Improve Challenge 2025: Tasks and Baseline Systems
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
mdok-style at SemEval-2026 Task 10: Finetuning LLMs for Conspiracy Detection
von: Macko, Dominik
Veröffentlicht: (2026)
von: Macko, Dominik
Veröffentlicht: (2026)
Synthesizing Privacy-Preserving Text Data via Finetuning without Finetuning Billion-Scale LLMs
von: Tan, Bowen, et al.
Veröffentlicht: (2025)
von: Tan, Bowen, et al.
Veröffentlicht: (2025)
mdok-style at SemEval-2026 Task 9: Finetuning LLMs for Multilingual Polarization Detection
von: Macko, Dominik, et al.
Veröffentlicht: (2026)
von: Macko, Dominik, et al.
Veröffentlicht: (2026)
Evaluation of Finetuned LLMs in AMR Parsing
von: Ho, Shu Han
Veröffentlicht: (2025)
von: Ho, Shu Han
Veröffentlicht: (2025)
The Impact of Editorial Intervention on Detecting Native Language Traces
von: Uluslu, Ahmet Yavuz, et al.
Veröffentlicht: (2026)
von: Uluslu, Ahmet Yavuz, et al.
Veröffentlicht: (2026)
Grammatical Error Feedback: An Implicit Evaluation Approach
von: Bannò, Stefano, et al.
Veröffentlicht: (2024)
von: Bannò, Stefano, et al.
Veröffentlicht: (2024)
MemLLM: Finetuning LLMs to Use An Explicit Read-Write Memory
von: Modarressi, Ali, et al.
Veröffentlicht: (2024)
von: Modarressi, Ali, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons
von: Liusie, Adian, et al.
Veröffentlicht: (2024) -
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
von: Raina, Vyas, et al.
Veröffentlicht: (2024) -
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
von: Liusie, Adian, et al.
Veröffentlicht: (2023) -
WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models
von: Molenda, Piotr, et al.
Veröffentlicht: (2024) -
Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models
von: Liusie, Adian, et al.
Veröffentlicht: (2024)