Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Rad, Melissa Kazemi, Nghiem, Huy, Luo, Andy, Wadhwa, Sahil, Sorower, Mohammad, Rawls, Stephen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STELP: Secure Transpilation and Execution of LLM-Generated Programs
by: Shinde, Swapnil, et al.
Published: (2026)
by: Shinde, Swapnil, et al.
Published: (2026)
Building Safe GenAI Applications: An End-to-End Overview of Red Teaming for Large Language Models
by: Purpura, Alberto, et al.
Published: (2025)
by: Purpura, Alberto, et al.
Published: (2025)
GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
by: Rad, Melissa Kazemi, et al.
Published: (2025)
by: Rad, Melissa Kazemi, et al.
Published: (2025)
Refine-n-Judge: Curating High-Quality Preference Chains for LLM-Fine-Tuning
by: Cayir, Derin, et al.
Published: (2025)
by: Cayir, Derin, et al.
Published: (2025)
Adaptive Instruction Composition for Automated LLM Red-Teaming
by: Zymet, Jesse, et al.
Published: (2026)
by: Zymet, Jesse, et al.
Published: (2026)
TRACT: Regression-Aware Fine-tuning Meets Chain-of-Thought Reasoning for LLM-as-a-Judge
by: Chiang, Cheng-Han, et al.
Published: (2025)
by: Chiang, Cheng-Han, et al.
Published: (2025)
On the Impact of Fine-Tuning on Chain-of-Thought Reasoning
by: Lobo, Elita, et al.
Published: (2024)
by: Lobo, Elita, et al.
Published: (2024)
Long-Document QA with Chain-of-Structured-Thought and Fine-Tuned SLMs
by: Liang, Zhuowen, et al.
Published: (2026)
by: Liang, Zhuowen, et al.
Published: (2026)
EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
by: Mohammadi, Hadi, et al.
Published: (2025)
by: Mohammadi, Hadi, et al.
Published: (2025)
Fine-Tuned Thoughts: Leveraging Chain-of-Thought Reasoning for Industrial Asset Health Monitoring
by: Lin, Shuxin, et al.
Published: (2025)
by: Lin, Shuxin, et al.
Published: (2025)
Northeastern Uni at Multilingual Counterspeech Generation: Enhancing Counter Speech Generation with LLM Alignment through Direct Preference Optimization
by: Wadhwa, Sahil, et al.
Published: (2024)
by: Wadhwa, Sahil, et al.
Published: (2024)
Alignment Dynamics in LLM Fine-Tuning
by: Huang, Yuhan, et al.
Published: (2026)
by: Huang, Yuhan, et al.
Published: (2026)
Learning to Refine with Fine-Grained Natural Language Feedback
by: Wadhwa, Manya, et al.
Published: (2024)
by: Wadhwa, Manya, et al.
Published: (2024)
HoT: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs
by: Nguyen, Tin, et al.
Published: (2025)
by: Nguyen, Tin, et al.
Published: (2025)
Focused Chain-of-Thought: Efficient LLM Reasoning via Structured Input Information
by: Struppek, Lukas, et al.
Published: (2025)
by: Struppek, Lukas, et al.
Published: (2025)
No Free Lunch with Guardrails
by: Kumar, Divyanshu, et al.
Published: (2025)
by: Kumar, Divyanshu, et al.
Published: (2025)
AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
by: Zhu, Zihao, et al.
Published: (2025)
by: Zhu, Zihao, et al.
Published: (2025)
CARFT: Boosting LLM Reasoning via Contrastive Learning with Annotated Chain-of-Thought-based Reinforced Fine-Tuning
by: Zhu, Wenqiao, et al.
Published: (2025)
by: Zhu, Wenqiao, et al.
Published: (2025)
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
by: Hsiung, Lei, et al.
Published: (2025)
by: Hsiung, Lei, et al.
Published: (2025)
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
by: Li, Miao, et al.
Published: (2026)
by: Li, Miao, et al.
Published: (2026)
HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models
by: Nghiem, Huy, et al.
Published: (2024)
by: Nghiem, Huy, et al.
Published: (2024)
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
Beyond Fine-Tuning: In-Context Learning and Chain-of-Thought for Reasoned Distractor Generation
by: Alhazmi, Elaf, et al.
Published: (2026)
by: Alhazmi, Elaf, et al.
Published: (2026)
Cause-Aware Empathetic Response Generation via Chain-of-Thought Fine-Tuning
by: Chen, Xinhao, et al.
Published: (2024)
by: Chen, Xinhao, et al.
Published: (2024)
Refined Quantum Algorithms for Principal Component Analysis and Solving Linear System
by: Nghiem, Nhat A.
Published: (2025)
by: Nghiem, Nhat A.
Published: (2025)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
by: Nghiem, Huy, et al.
Published: (2025)
by: Nghiem, Huy, et al.
Published: (2025)
Guardrails in Logit Space: Safety Token Regularization for LLM Alignment
by: Bach, Thong, et al.
Published: (2026)
by: Bach, Thong, et al.
Published: (2026)
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM
by: Kiruluta, Andrew, et al.
Published: (2025)
by: Kiruluta, Andrew, et al.
Published: (2025)
Hierarchical Chain-of-Thought Prompting: Enhancing LLM Reasoning Performance and Efficiency
by: Huang, Xingshuai, et al.
Published: (2026)
by: Huang, Xingshuai, et al.
Published: (2026)
Fine-grained Analysis of Brain-LLM Alignment through Input Attribution
by: Proietti, Michela, et al.
Published: (2025)
by: Proietti, Michela, et al.
Published: (2025)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
by: Mittal, Avni, et al.
Published: (2026)
by: Mittal, Avni, et al.
Published: (2026)
LIFT: Improving Long Context Understanding Through Long Input Fine-Tuning
by: Mao, Yansheng, et al.
Published: (2024)
by: Mao, Yansheng, et al.
Published: (2024)
On the Role of Reasoning Patterns in the Generalization Discrepancy of Long Chain-of-Thought Supervised Fine-Tuning
by: Li, Zhaoyi, et al.
Published: (2026)
by: Li, Zhaoyi, et al.
Published: (2026)
Dual-Channel Closed Loop Supply Chain Competition: A Stackelberg--Nash Approach
by: Wadhwa, Gurkirat
Published: (2026)
by: Wadhwa, Gurkirat
Published: (2026)
Rethinking Code Refinement: Learning to Judge Code Efficiency
by: Seo, Minju, et al.
Published: (2024)
by: Seo, Minju, et al.
Published: (2024)
Fine-Grained Alignment and Noise Refinement for Compositional Text-to-Image Generation
by: Izadi, Amir Mohammad, et al.
Published: (2025)
by: Izadi, Amir Mohammad, et al.
Published: (2025)
ARES: Alternating Reinforcement Learning and Supervised Fine-Tuning for Enhanced Multi-Modal Chain-of-Thought Reasoning Through Diverse AI Feedback
by: Byun, Ju-Seung, et al.
Published: (2024)
by: Byun, Ju-Seung, et al.
Published: (2024)
"Define Your Terms" : Enhancing Efficient Offensive Speech Classification with Definition
by: Nghiem, Huy, et al.
Published: (2024)
by: Nghiem, Huy, et al.
Published: (2024)
Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMs
by: Puerto, Haritz, et al.
Published: (2024)
by: Puerto, Haritz, et al.
Published: (2024)
Parameter Efficiency Is Not Memory Efficiency: Rethinking Fine-Tuning for On-Device LLM Adaptation
by: Tenison, Irene, et al.
Published: (2026)
by: Tenison, Irene, et al.
Published: (2026)
Similar Items
-
STELP: Secure Transpilation and Execution of LLM-Generated Programs
by: Shinde, Swapnil, et al.
Published: (2026) -
Building Safe GenAI Applications: An End-to-End Overview of Red Teaming for Large Language Models
by: Purpura, Alberto, et al.
Published: (2025) -
GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
by: Rad, Melissa Kazemi, et al.
Published: (2025) -
Refine-n-Judge: Curating High-Quality Preference Chains for LLM-Fine-Tuning
by: Cayir, Derin, et al.
Published: (2025) -
Adaptive Instruction Composition for Automated LLM Red-Teaming
by: Zymet, Jesse, et al.
Published: (2026)