Saved in:
| Main Authors: | Garg, Sahil, Cheezum, Brad, Dutta, Sridhar, Agarwal, Vishal |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.22269 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Road of Adaptive AI for Precision in Cybersecurity
by: Garg, Sahil
Published: (2025)
by: Garg, Sahil
Published: (2025)
What is in a name? Mitigating Name Bias in Text Embeddings via Anonymization
by: Manchanda, Sahil, et al.
Published: (2025)
by: Manchanda, Sahil, et al.
Published: (2025)
Build, Judge, Optimize: A Blueprint for Continuous Improvement of Multi-Agent Consumer Assistants
by: Herrera, Alejandro Breen, et al.
Published: (2026)
by: Herrera, Alejandro Breen, et al.
Published: (2026)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
by: Duo, Jiangshan, et al.
Published: (2026)
by: Duo, Jiangshan, et al.
Published: (2026)
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
by: Qu, Yuxiao, et al.
Published: (2024)
by: Qu, Yuxiao, et al.
Published: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
ELSA: A Style Aligned Dataset for Emotionally Intelligent Language Generation
by: Gandhi, Vishal, et al.
Published: (2025)
by: Gandhi, Vishal, et al.
Published: (2025)
Local Prompt Optimization
by: Jain, Yash, et al.
Published: (2025)
by: Jain, Yash, et al.
Published: (2025)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
by: Chow, Yinlam, et al.
Published: (2024)
by: Chow, Yinlam, et al.
Published: (2024)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
by: Xu, Austin, et al.
Published: (2025)
by: Xu, Austin, et al.
Published: (2025)
VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
by: Jain, Dhruv, et al.
Published: (2025)
by: Jain, Dhruv, et al.
Published: (2025)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
by: Li, Zhuochun, et al.
Published: (2026)
by: Li, Zhuochun, et al.
Published: (2026)
CLewR: Curriculum Learning with Restarts for Machine Translation Preference Learning
by: Dragomir, Alexandra, et al.
Published: (2026)
by: Dragomir, Alexandra, et al.
Published: (2026)
Sparse Shift Autoencoders for Identifying Concepts from Large Language Model Activations
by: Joshi, Shruti, et al.
Published: (2025)
by: Joshi, Shruti, et al.
Published: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025)
by: Liu, Yixin, et al.
Published: (2025)
The Perfect Blend: Redefining RLHF with Mixture of Judges
by: Xu, Tengyu, et al.
Published: (2024)
by: Xu, Tengyu, et al.
Published: (2024)
Investigating Non-Transitivity in LLM-as-a-Judge
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
ORPO-Distill: Mixed-Policy Preference Optimization for Cross-Architecture LLM Distillation
by: Singh, Aasheesh, et al.
Published: (2025)
by: Singh, Aasheesh, et al.
Published: (2025)
WebArena: A Realistic Web Environment for Building Autonomous Agents
by: Zhou, Shuyan, et al.
Published: (2023)
by: Zhou, Shuyan, et al.
Published: (2023)
Can Agents Judge Systematic Reviews Like Humans? Evaluating SLRs with LLM-based Multi-Agent System
by: Mushtaq, Abdullah, et al.
Published: (2025)
by: Mushtaq, Abdullah, et al.
Published: (2025)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
by: Yang, Jinming, et al.
Published: (2026)
by: Yang, Jinming, et al.
Published: (2026)
Debate Helps Weak Judges Reward Stronger Models
by: Elasky, Ethan, et al.
Published: (2026)
by: Elasky, Ethan, et al.
Published: (2026)
TabDistill: Distilling Transformers into Neural Nets for Few-Shot Tabular Classification
by: Dissanayake, Pasan, et al.
Published: (2025)
by: Dissanayake, Pasan, et al.
Published: (2025)
Languages are Modalities: Cross-Lingual Alignment via Encoder Injection
by: Agarwal, Rajan, et al.
Published: (2025)
by: Agarwal, Rajan, et al.
Published: (2025)
Correlated Errors in Large Language Models
by: Kim, Elliot, et al.
Published: (2025)
by: Kim, Elliot, et al.
Published: (2025)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
by: Ballon, Marthe, et al.
Published: (2026)
by: Ballon, Marthe, et al.
Published: (2026)
Context Over Content: Exposing Evaluation Faking in Automated Judges
by: Gupta, Manan, et al.
Published: (2026)
by: Gupta, Manan, et al.
Published: (2026)
JuStRank: Benchmarking LLM Judges for System Ranking
by: Gera, Ariel, et al.
Published: (2024)
by: Gera, Ariel, et al.
Published: (2024)
Becoming Experienced Judges: Selective Test-Time Learning for Evaluators
by: Jwa, Seungyeon, et al.
Published: (2025)
by: Jwa, Seungyeon, et al.
Published: (2025)
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
by: Salinas, David, et al.
Published: (2025)
by: Salinas, David, et al.
Published: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
by: Liu, Yixin, et al.
Published: (2026)
by: Liu, Yixin, et al.
Published: (2026)
GNN-as-Judge: Unleashing the Power of LLMs for Graph Learning with GNN Feedback
by: Xu, Ruiyao, et al.
Published: (2026)
by: Xu, Ruiyao, et al.
Published: (2026)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
by: Gupta, Manan, et al.
Published: (2026)
by: Gupta, Manan, et al.
Published: (2026)
Neural Networks for Learnable and Scalable Influence Estimation of Instruction Fine-Tuning Data
by: Agarwal, Ishika, et al.
Published: (2025)
by: Agarwal, Ishika, et al.
Published: (2025)
Approximating Human Preferences Using a Multi-Judge Learned System
by: Sprejer, Eitán, et al.
Published: (2025)
by: Sprejer, Eitán, et al.
Published: (2025)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
by: Hu, Renjun, et al.
Published: (2025)
by: Hu, Renjun, et al.
Published: (2025)
Enabling Weak LLMs to Judge Response Reliability via Meta Ranking
by: Liu, Zijun, et al.
Published: (2024)
by: Liu, Zijun, et al.
Published: (2024)
ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
by: Chen, Ziru, et al.
Published: (2024)
by: Chen, Ziru, et al.
Published: (2024)
Mitigating Hallucinated Translations in Large Language Models with Hallucination-focused Preference Optimization
by: Tang, Zilu, et al.
Published: (2025)
by: Tang, Zilu, et al.
Published: (2025)
Similar Items
-
The Road of Adaptive AI for Precision in Cybersecurity
by: Garg, Sahil
Published: (2025) -
What is in a name? Mitigating Name Bias in Text Embeddings via Anonymization
by: Manchanda, Sahil, et al.
Published: (2025) -
Build, Judge, Optimize: A Blueprint for Continuous Improvement of Multi-Agent Consumer Assistants
by: Herrera, Alejandro Breen, et al.
Published: (2026) -
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
by: Duo, Jiangshan, et al.
Published: (2026) -
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
by: Qu, Yuxiao, et al.
Published: (2024)