Refine-n-Judge: Curating High-Quality Preference Chains for LLM-Fine-Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Cayir, Derin, Tao, Renjie, Rungta, Rashi, Sun, Kai, Chen, Sean, Khan, Haidar, Kim, Minseok, Reinspach, Julia, Liu, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment
by: Rad, Melissa Kazemi, et al.
Published: (2025)
by: Rad, Melissa Kazemi, et al.
Published: (2025)
Augmenting Security and Privacy in the Virtual Realm: An Analysis of Extended Reality Devices
by: Cayir, Derin, et al.
Published: (2024)
by: Cayir, Derin, et al.
Published: (2024)
Auditory Environments to Physical Environments
by: Ivaana Rungta
Published: (2026)
by: Ivaana Rungta
Published: (2026)
MCTS-Refined CoT: High-Quality Fine-Tuning Data for LLM-Based Repository Issue Resolution
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
by: Yang, Jinming, et al.
Published: (2026)
by: Yang, Jinming, et al.
Published: (2026)
CPJ: Explainable Agricultural Pest Diagnosis via Caption-Prompt-Judge with LLM-Judged Refinement
by: Zhang, Wentao, et al.
Published: (2025)
by: Zhang, Wentao, et al.
Published: (2025)
MallowsPO: Fine-Tune Your LLM with Preference Dispersions
by: Chen, Haoxian, et al.
Published: (2024)
by: Chen, Haoxian, et al.
Published: (2024)
Mind, Meaning and Architecture
by: Ivaana Rungta, Ivaana
Published: (2026)
by: Ivaana Rungta, Ivaana
Published: (2026)
Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMs
by: Puerto, Haritz, et al.
Published: (2024)
by: Puerto, Haritz, et al.
Published: (2024)
Chapter Power, Patronage, and Confessionalism
by: Terzioglu, Derin
Published: (2021)
by: Terzioglu, Derin
Published: (2021)
Automated Data Curation for Robust Language Model Fine-Tuning
by: Chen, Jiuhai, et al.
Published: (2024)
by: Chen, Jiuhai, et al.
Published: (2024)
LLM-as-a-Fuzzy-Judge: Fine-Tuning Large Language Models as a Clinical Evaluation Judge with Fuzzy Logic
by: Zheng, Weibing, et al.
Published: (2025)
by: Zheng, Weibing, et al.
Published: (2025)
Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges
by: Li, Yanran
Published: (2026)
by: Li, Yanran
Published: (2026)
Self-Preference Bias in LLM-as-a-Judge
by: Wataoka, Koki, et al.
Published: (2024)
by: Wataoka, Koki, et al.
Published: (2024)
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
by: Qin, Chongli, et al.
Published: (2025)
by: Qin, Chongli, et al.
Published: (2025)
TRACT: Regression-Aware Fine-tuning Meets Chain-of-Thought Reasoning for LLM-as-a-Judge
by: Chiang, Cheng-Han, et al.
Published: (2025)
by: Chiang, Cheng-Han, et al.
Published: (2025)
Performance Improvement of Time-Balance Radar Schedulers Through Decision Policies (Extended Version)
by: Çayır, Ömer, et al.
Published: (2017)
by: Çayır, Ömer, et al.
Published: (2017)
On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization
by: Singh, Janvijay, et al.
Published: (2025)
by: Singh, Janvijay, et al.
Published: (2025)
Attention or Convolution: Transformer Encoders in Audio Language Models for Inference Efficiency
by: Jeon, Sungho, et al.
Published: (2023)
by: Jeon, Sungho, et al.
Published: (2023)
ConQRet: Benchmarking Fine-Grained Evaluation of Retrieval Augmented Argumentation with LLM Judges
by: Dhole, Kaustubh D., et al.
Published: (2024)
by: Dhole, Kaustubh D., et al.
Published: (2024)
Increasing the p-Selmer rank by twisting
by: Kim, Minseok
Published: (2025)
by: Kim, Minseok
Published: (2025)
$p$-twisted Selmer near-companion curves
by: Kim, Minseok
Published: (2025)
by: Kim, Minseok
Published: (2025)
CodeLutra: Boosting LLM Code Generation via Preference-Guided Refinement
by: Tao, Leitian, et al.
Published: (2024)
by: Tao, Leitian, et al.
Published: (2024)
Enhancing Reinforcement Learning Fine-Tuning with an Online Refiner
by: Ma, Hao, et al.
Published: (2026)
by: Ma, Hao, et al.
Published: (2026)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
by: Ding, Ruomeng, et al.
Published: (2026)
by: Ding, Ruomeng, et al.
Published: (2026)
Automated Concept Discovery for LLM-as-a-Judge Preference Analysis
by: Wedgwood, James, et al.
Published: (2026)
by: Wedgwood, James, et al.
Published: (2026)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
by: Fernando, Heshan, et al.
Published: (2024)
by: Fernando, Heshan, et al.
Published: (2024)
The Fine-Tuning Paradox: Boosting Translation Quality Without Sacrificing LLM Abilities
by: Stap, David, et al.
Published: (2024)
by: Stap, David, et al.
Published: (2024)
On quotients of derivatives of $L$-functions inside the critical strip
by: Lunia, Rashi
Published: (2024)
by: Lunia, Rashi
Published: (2024)
From Classroom to Marketplace: English as a Tool for Business, Monetization, and Global Careers in a Neoliberal Market
by: Iftikhar Khan, et al.
Published: (2026)
by: Iftikhar Khan, et al.
Published: (2026)
Selection of LLM Fine-Tuning Data based on Orthogonal Rules
by: Li, Xiaomin, et al.
Published: (2024)
by: Li, Xiaomin, et al.
Published: (2024)
On the Impact of Fine-Tuning on Chain-of-Thought Reasoning
by: Lobo, Elita, et al.
Published: (2024)
by: Lobo, Elita, et al.
Published: (2024)
Aligning Paralinguistic Understanding and Generation in Speech LLMs via Multi-Task Reinforcement Learning
by: Chen, Jingxiang, et al.
Published: (2026)
by: Chen, Jingxiang, et al.
Published: (2026)
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
by: Zheng, Ruobing, et al.
Published: (2026)
by: Zheng, Ruobing, et al.
Published: (2026)
Historicizing Sunni Islam in the Ottoman Empire, c. 1450-c. 1750
by: Krstić, Tijana, et al.
Published: (2020)
by: Krstić, Tijana, et al.
Published: (2020)
Experimental support for a sustainable treatment strategy of acrylic fiber dyeing wastewater
by: Işık Kabdaşlı, et al.
Published: (2024)
by: Işık Kabdaşlı, et al.
Published: (2024)
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents
by: Huang, Jiawei, et al.
Published: (2026)
by: Huang, Jiawei, et al.
Published: (2026)
Fine‐Tuning Lightweight LLMs With Human‐Curated Data on Electrical Circuit Fundamentals for E‐Learning
by: André Rocha, et al.
Published: (2026)
by: André Rocha, et al.
Published: (2026)
Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judge
by: Liu, Zhuo, et al.
Published: (2025)
by: Liu, Zhuo, et al.
Published: (2025)
Velázquez, Painter & Curator
by: Vázquez, Julia
Published: (2025)
by: Vázquez, Julia
Published: (2025)
Similar Items
-
Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment
by: Rad, Melissa Kazemi, et al.
Published: (2025) -
Augmenting Security and Privacy in the Virtual Realm: An Analysis of Extended Reality Devices
by: Cayir, Derin, et al.
Published: (2024) -
Auditory Environments to Physical Environments
by: Ivaana Rungta
Published: (2026) -
MCTS-Refined CoT: High-Quality Fine-Tuning Data for LLM-Based Repository Issue Resolution
by: Wang, Yibo, et al.
Published: (2025) -
Quantifying and Mitigating Self-Preference Bias of LLM Judges
by: Yang, Jinming, et al.
Published: (2026)