Rubric-Conditioned LLM Grading: Alignment, Uncertainty, and Robustness
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deng, Haotian, Farber, Chris, Lee, Jiyoon, Tang, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
R3: Robust Rubric-Agnostic Reward Models
von: Anugraha, David, et al.
Veröffentlicht: (2025)
von: Anugraha, David, et al.
Veröffentlicht: (2025)
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
von: Li, Gaotang, et al.
Veröffentlicht: (2026)
von: Li, Gaotang, et al.
Veröffentlicht: (2026)
Benchmarking Uncertainty Calibration in Large Language Model Long-Form Question Answering
von: Müller, Philip, et al.
Veröffentlicht: (2026)
von: Müller, Philip, et al.
Veröffentlicht: (2026)
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
von: Zhao, Jiale, et al.
Veröffentlicht: (2026)
von: Zhao, Jiale, et al.
Veröffentlicht: (2026)
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
von: Anugraha, David, et al.
Veröffentlicht: (2025)
von: Anugraha, David, et al.
Veröffentlicht: (2025)
HawkesLLM: Semantic Uncertainty Propagation in Agentic Text Simulation
von: Deng, Zewei, et al.
Veröffentlicht: (2026)
von: Deng, Zewei, et al.
Veröffentlicht: (2026)
Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training
von: Xu, Ran, et al.
Veröffentlicht: (2026)
von: Xu, Ran, et al.
Veröffentlicht: (2026)
RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains
von: Jiang, Haoxiang, et al.
Veröffentlicht: (2026)
von: Jiang, Haoxiang, et al.
Veröffentlicht: (2026)
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
von: Sharma, Manasi, et al.
Veröffentlicht: (2025)
von: Sharma, Manasi, et al.
Veröffentlicht: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
On the Robustness of Reward Models for Language Model Alignment
von: Hong, Jiwoo, et al.
Veröffentlicht: (2025)
von: Hong, Jiwoo, et al.
Veröffentlicht: (2025)
Learning to Judge: LLMs Designing and Applying Evaluation Rubrics
von: Siro, Clemencia, et al.
Veröffentlicht: (2026)
von: Siro, Clemencia, et al.
Veröffentlicht: (2026)
Long-Short Alignment for Effective Long-Context Modeling in LLMs
von: Du, Tianqi, et al.
Veröffentlicht: (2025)
von: Du, Tianqi, et al.
Veröffentlicht: (2025)
AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning
von: Wu, Peilin, et al.
Veröffentlicht: (2026)
von: Wu, Peilin, et al.
Veröffentlicht: (2026)
PoeTone: A Framework for Constrained Generation of Structured Chinese Songci with LLMs
von: Qu, Zhan, et al.
Veröffentlicht: (2025)
von: Qu, Zhan, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Rubric Anchors
von: Huang, Zenan, et al.
Veröffentlicht: (2025)
von: Huang, Zenan, et al.
Veröffentlicht: (2025)
LASER: An LLM-based ASR Scoring and Evaluation Rubric
von: Parulekar, Amruta, et al.
Veröffentlicht: (2025)
von: Parulekar, Amruta, et al.
Veröffentlicht: (2025)
Balancing the Reasoning Load: Difficulty-Differentiated Policy Optimization with Length Redistribution for Efficient and Robust Reinforcement Learning
von: Xia, Yinan, et al.
Veröffentlicht: (2026)
von: Xia, Yinan, et al.
Veröffentlicht: (2026)
C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences
von: Kawabata, Akira, et al.
Veröffentlicht: (2026)
von: Kawabata, Akira, et al.
Veröffentlicht: (2026)
LLM Processes: Numerical Predictive Distributions Conditioned on Natural Language
von: Requeima, James, et al.
Veröffentlicht: (2024)
von: Requeima, James, et al.
Veröffentlicht: (2024)
Online Rubrics Elicitation from Pairwise Comparisons
von: Rezaei, MohammadHossein, et al.
Veröffentlicht: (2025)
von: Rezaei, MohammadHossein, et al.
Veröffentlicht: (2025)
Enhance the Robustness of Text-Centric Multimodal Alignments
von: Yen, Ting-Yu, et al.
Veröffentlicht: (2024)
von: Yen, Ting-Yu, et al.
Veröffentlicht: (2024)
Align-then-Unlearn: Embedding Alignment for LLM Unlearning
von: Spohn, Philipp, et al.
Veröffentlicht: (2025)
von: Spohn, Philipp, et al.
Veröffentlicht: (2025)
Single Character Perturbations Break LLM Alignment
von: Lin, Leon, et al.
Veröffentlicht: (2024)
von: Lin, Leon, et al.
Veröffentlicht: (2024)
PolyNorm: Few-Shot LLM-Based Text Normalization for Text-to-Speech
von: Wong, Michel, et al.
Veröffentlicht: (2025)
von: Wong, Michel, et al.
Veröffentlicht: (2025)
Training AI Co-Scientists Using Rubric Rewards
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
von: Shao, Rulin, et al.
Veröffentlicht: (2025)
von: Shao, Rulin, et al.
Veröffentlicht: (2025)
Robust Multi-Objective Preference Alignment with Online DPO
von: Gupta, Raghav, et al.
Veröffentlicht: (2025)
von: Gupta, Raghav, et al.
Veröffentlicht: (2025)
Estimating Semantic Alphabet Size for LLM Uncertainty Quantification
von: McCabe, Lucas H., et al.
Veröffentlicht: (2025)
von: McCabe, Lucas H., et al.
Veröffentlicht: (2025)
SAUP: Situation Awareness Uncertainty Propagation on LLM Agent
von: Zhao, Qiwei, et al.
Veröffentlicht: (2024)
von: Zhao, Qiwei, et al.
Veröffentlicht: (2024)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
von: Gunjal, Anisha, et al.
Veröffentlicht: (2025)
von: Gunjal, Anisha, et al.
Veröffentlicht: (2025)
Advancing LLM Safe Alignment with Safety Representation Ranking
von: Du, Tianqi, et al.
Veröffentlicht: (2025)
von: Du, Tianqi, et al.
Veröffentlicht: (2025)
Transfer Q Star: Principled Decoding for LLM Alignment
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2024)
Less is More: Improving LLM Alignment via Preference Data Selection
von: Deng, Xun, et al.
Veröffentlicht: (2025)
von: Deng, Xun, et al.
Veröffentlicht: (2025)
Who's Asking? Evaluating LLM Robustness to Inquiry Personas in Factual Question Answering
von: Akpinar, Nil-Jana, et al.
Veröffentlicht: (2025)
von: Akpinar, Nil-Jana, et al.
Veröffentlicht: (2025)
Deep Learning-based Method for Expressing Knowledge Boundary of Black-Box LLM
von: Sheng, Haotian, et al.
Veröffentlicht: (2026)
von: Sheng, Haotian, et al.
Veröffentlicht: (2026)
LLM Flow Processes for Text-Conditioned Regression
von: Biggs, Felix, et al.
Veröffentlicht: (2026)
von: Biggs, Felix, et al.
Veröffentlicht: (2026)
Energy-Based Reward Models for Robust Language Model Alignment
von: Lochab, Anamika, et al.
Veröffentlicht: (2025)
von: Lochab, Anamika, et al.
Veröffentlicht: (2025)
UCCI: Calibrated Uncertainty for Cost-Optimal LLM Cascade Routing
von: Kotte, Varun
Veröffentlicht: (2026)
von: Kotte, Varun
Veröffentlicht: (2026)
PluralLLM: Pluralistic Alignment in LLMs via Federated Learning
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2025)
von: Srewa, Mahmoud, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
R3: Robust Rubric-Agnostic Reward Models
von: Anugraha, David, et al.
Veröffentlicht: (2025) -
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
von: Li, Gaotang, et al.
Veröffentlicht: (2026) -
Benchmarking Uncertainty Calibration in Large Language Model Long-Form Question Answering
von: Müller, Philip, et al.
Veröffentlicht: (2026) -
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
von: Zhao, Jiale, et al.
Veröffentlicht: (2026) -
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
von: Anugraha, David, et al.
Veröffentlicht: (2025)