Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | Rad, Melissa Kazemi, Nghiem, Huy, Luo, Andy, Wadhwa, Sahil, Sorower, Mohammad, Rawls, Stephen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
STELP: Secure Transpilation and Execution of LLM-Generated Programs
por: Shinde, Swapnil, et al.
Publicado: (2026)
por: Shinde, Swapnil, et al.
Publicado: (2026)
Building Safe GenAI Applications: An End-to-End Overview of Red Teaming for Large Language Models
por: Purpura, Alberto, et al.
Publicado: (2025)
por: Purpura, Alberto, et al.
Publicado: (2025)
GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
por: Rad, Melissa Kazemi, et al.
Publicado: (2025)
por: Rad, Melissa Kazemi, et al.
Publicado: (2025)
Refine-n-Judge: Curating High-Quality Preference Chains for LLM-Fine-Tuning
por: Cayir, Derin, et al.
Publicado: (2025)
por: Cayir, Derin, et al.
Publicado: (2025)
Adaptive Instruction Composition for Automated LLM Red-Teaming
por: Zymet, Jesse, et al.
Publicado: (2026)
por: Zymet, Jesse, et al.
Publicado: (2026)
TRACT: Regression-Aware Fine-tuning Meets Chain-of-Thought Reasoning for LLM-as-a-Judge
por: Chiang, Cheng-Han, et al.
Publicado: (2025)
por: Chiang, Cheng-Han, et al.
Publicado: (2025)
On the Impact of Fine-Tuning on Chain-of-Thought Reasoning
por: Lobo, Elita, et al.
Publicado: (2024)
por: Lobo, Elita, et al.
Publicado: (2024)
Long-Document QA with Chain-of-Structured-Thought and Fine-Tuned SLMs
por: Liang, Zhuowen, et al.
Publicado: (2026)
por: Liang, Zhuowen, et al.
Publicado: (2026)
EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
por: Mohammadi, Hadi, et al.
Publicado: (2025)
por: Mohammadi, Hadi, et al.
Publicado: (2025)
Fine-Tuned Thoughts: Leveraging Chain-of-Thought Reasoning for Industrial Asset Health Monitoring
por: Lin, Shuxin, et al.
Publicado: (2025)
por: Lin, Shuxin, et al.
Publicado: (2025)
Northeastern Uni at Multilingual Counterspeech Generation: Enhancing Counter Speech Generation with LLM Alignment through Direct Preference Optimization
por: Wadhwa, Sahil, et al.
Publicado: (2024)
por: Wadhwa, Sahil, et al.
Publicado: (2024)
Alignment Dynamics in LLM Fine-Tuning
por: Huang, Yuhan, et al.
Publicado: (2026)
por: Huang, Yuhan, et al.
Publicado: (2026)
Learning to Refine with Fine-Grained Natural Language Feedback
por: Wadhwa, Manya, et al.
Publicado: (2024)
por: Wadhwa, Manya, et al.
Publicado: (2024)
HoT: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs
por: Nguyen, Tin, et al.
Publicado: (2025)
por: Nguyen, Tin, et al.
Publicado: (2025)
Focused Chain-of-Thought: Efficient LLM Reasoning via Structured Input Information
por: Struppek, Lukas, et al.
Publicado: (2025)
por: Struppek, Lukas, et al.
Publicado: (2025)
No Free Lunch with Guardrails
por: Kumar, Divyanshu, et al.
Publicado: (2025)
por: Kumar, Divyanshu, et al.
Publicado: (2025)
AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
por: Zhu, Zihao, et al.
Publicado: (2025)
por: Zhu, Zihao, et al.
Publicado: (2025)
CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuning
por: Zhu, Wenqiao, et al.
Publicado: (2025)
por: Zhu, Wenqiao, et al.
Publicado: (2025)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
por: Hsiung, Lei, et al.
Publicado: (2025)
por: Hsiung, Lei, et al.
Publicado: (2025)
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
por: Li, Miao, et al.
Publicado: (2026)
por: Li, Miao, et al.
Publicado: (2026)
HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models
por: Nghiem, Huy, et al.
Publicado: (2024)
por: Nghiem, Huy, et al.
Publicado: (2024)
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
por: Wang, Yibin, et al.
Publicado: (2025)
por: Wang, Yibin, et al.
Publicado: (2025)
Beyond Fine-Tuning: In-Context Learning and Chain-of-Thought for Reasoned Distractor Generation
por: Alhazmi, Elaf, et al.
Publicado: (2026)
por: Alhazmi, Elaf, et al.
Publicado: (2026)
Cause-Aware Empathetic Response Generation via Chain-of-Thought Fine-Tuning
por: Chen, Xinhao, et al.
Publicado: (2024)
por: Chen, Xinhao, et al.
Publicado: (2024)
Refined Quantum Algorithms for Principal Component Analysis and Solving Linear System
por: Nghiem, Nhat A.
Publicado: (2025)
por: Nghiem, Nhat A.
Publicado: (2025)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
por: Nghiem, Huy, et al.
Publicado: (2025)
por: Nghiem, Huy, et al.
Publicado: (2025)
Guardrails in Logit Space: Safety Token Regularization for LLM Alignment
por: Bach, Thong, et al.
Publicado: (2026)
por: Bach, Thong, et al.
Publicado: (2026)
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM
por: Kiruluta, Andrew, et al.
Publicado: (2025)
por: Kiruluta, Andrew, et al.
Publicado: (2025)
Hierarchical Chain-of-Thought Prompting: Enhancing LLM Reasoning Performance and Efficiency
por: Huang, Xingshuai, et al.
Publicado: (2026)
por: Huang, Xingshuai, et al.
Publicado: (2026)
Fine-grained Analysis of Brain-LLM Alignment through Input Attribution
por: Proietti, Michela, et al.
Publicado: (2025)
por: Proietti, Michela, et al.
Publicado: (2025)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
por: Mittal, Avni, et al.
Publicado: (2026)
por: Mittal, Avni, et al.
Publicado: (2026)
LIFT: Improving Long Context Understanding Through Long Input Fine-Tuning
por: Mao, Yansheng, et al.
Publicado: (2024)
por: Mao, Yansheng, et al.
Publicado: (2024)
On the Role of Reasoning Patterns in the Generalization Discrepancy of Long Chain-of-Thought Supervised Fine-Tuning
por: Li, Zhaoyi, et al.
Publicado: (2026)
por: Li, Zhaoyi, et al.
Publicado: (2026)
Dual-Channel Closed Loop Supply Chain Competition: A Stackelberg--Nash Approach
por: Wadhwa, Gurkirat
Publicado: (2026)
por: Wadhwa, Gurkirat
Publicado: (2026)
Rethinking Code Refinement: Learning to Judge Code Efficiency
por: Seo, Minju, et al.
Publicado: (2024)
por: Seo, Minju, et al.
Publicado: (2024)
Fine-Grained Alignment and Noise Refinement for Compositional Text-to-Image Generation
por: Izadi, Amir Mohammad, et al.
Publicado: (2025)
por: Izadi, Amir Mohammad, et al.
Publicado: (2025)
ARES: Alternating Reinforcement Learning and Supervised Fine-Tuning for Enhanced Multi-Modal Chain-of-Thought Reasoning Through Diverse AI Feedback
por: Byun, Ju-Seung, et al.
Publicado: (2024)
por: Byun, Ju-Seung, et al.
Publicado: (2024)
"Define Your Terms" : Enhancing Efficient Offensive Speech Classification with Definition
por: Nghiem, Huy, et al.
Publicado: (2024)
por: Nghiem, Huy, et al.
Publicado: (2024)
Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMs
por: Puerto, Haritz, et al.
Publicado: (2024)
por: Puerto, Haritz, et al.
Publicado: (2024)
Parameter Efficiency Is Not Memory Efficiency: Rethinking Fine-Tuning for On-Device LLM Adaptation
por: Tenison, Irene, et al.
Publicado: (2026)
por: Tenison, Irene, et al.
Publicado: (2026)
Ejemplares similares
-
STELP: Secure Transpilation and Execution of LLM-Generated Programs
por: Shinde, Swapnil, et al.
Publicado: (2026) -
Building Safe GenAI Applications: An End-to-End Overview of Red Teaming for Large Language Models
por: Purpura, Alberto, et al.
Publicado: (2025) -
GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
por: Rad, Melissa Kazemi, et al.
Publicado: (2025) -
Refine-n-Judge: Curating High-Quality Preference Chains for LLM-Fine-Tuning
por: Cayir, Derin, et al.
Publicado: (2025) -
Adaptive Instruction Composition for Automated LLM Red-Teaming
por: Zymet, Jesse, et al.
Publicado: (2026)