Feedback-to-Rubrics: Can We Learn Expert Criteria from Inline Comments?
Fuente:
arXiv
Saved in:
| Main Authors: | Yoshida, Kotaro, Kuroki, So, Imajuku, Yuki, Nakamura, Taishi, Iwai, Ryunosuke, Goda, Haruki, Akiba, Takuya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search
by: Inoue, Yuichi, et al.
Published: (2025)
by: Inoue, Yuichi, et al.
Published: (2025)
Agent Skill Acquisition for Large Language Models via CycleQD
by: Kuroki, So, et al.
Published: (2024)
by: Kuroki, So, et al.
Published: (2024)
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
by: Nakamura, Taishi, et al.
Published: (2025)
by: Nakamura, Taishi, et al.
Published: (2025)
ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution
by: Lange, Robert Tjarko, et al.
Published: (2025)
by: Lange, Robert Tjarko, et al.
Published: (2025)
Robust Invariant Representation Learning by Distribution Extrapolation
by: Yoshida, Kotaro, et al.
Published: (2025)
by: Yoshida, Kotaro, et al.
Published: (2025)
How Does Preconditioning Guide Feature Learning in Deep Neural Networks?
by: Yoshida, Kotaro, et al.
Published: (2025)
by: Yoshida, Kotaro, et al.
Published: (2025)
Multi-Play Combinatorial Semi-Bandit Problem
by: Nakamura, Shintaro, et al.
Published: (2025)
by: Nakamura, Shintaro, et al.
Published: (2025)
Balancing Speed and Stability: The Trade-offs of FP8 vs. BF16 Training in LLMs
by: Fujii, Kazuki, et al.
Published: (2024)
by: Fujii, Kazuki, et al.
Published: (2024)
Towards Understanding Variants of Invariant Risk Minimization through the Lens of Calibration
by: Yoshida, Kotaro, et al.
Published: (2024)
by: Yoshida, Kotaro, et al.
Published: (2024)
Sudoku-Bench: Evaluating creative reasoning with Sudoku variants
by: Seely, Jeffrey, et al.
Published: (2025)
by: Seely, Jeffrey, et al.
Published: (2025)
UnMaskFork: Test-Time Scaling for Masked Diffusion via Deterministic Action Branching
by: Misaki, Kou, et al.
Published: (2026)
by: Misaki, Kou, et al.
Published: (2026)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
by: Nakamura, Taishi, et al.
Published: (2025)
by: Nakamura, Taishi, et al.
Published: (2025)
On the Optimal Reasoning Length for RL-Trained Language Models
by: Nohara, Daisuke, et al.
Published: (2026)
by: Nohara, Daisuke, et al.
Published: (2026)
Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
by: Xie, Lipeng, et al.
Published: (2025)
by: Xie, Lipeng, et al.
Published: (2025)
KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI
by: Kuroki, So, et al.
Published: (2025)
by: Kuroki, So, et al.
Published: (2025)
DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation
by: Shing, Makoto, et al.
Published: (2025)
by: Shing, Makoto, et al.
Published: (2025)
Dynamic Structure Estimation from Bandit Feedback using Nonvanishing Exponential Sums
by: Ohnishi, Motoya, et al.
Published: (2022)
by: Ohnishi, Motoya, et al.
Published: (2022)
Sequence-Aware Inline Measurement Attribution for Good-Bad Wafer Diagnosis
by: Miyaguchi, Kohei, et al.
Published: (2025)
by: Miyaguchi, Kohei, et al.
Published: (2025)
Can We Detect Failures Without Failure Data? Uncertainty-Aware Runtime Failure Detection for Imitation Learning Policies
by: Xu, Chen, et al.
Published: (2025)
by: Xu, Chen, et al.
Published: (2025)
On What We Can Learn from Low-Resolution Data
by: Frehr, Theresa Dahl, et al.
Published: (2026)
by: Frehr, Theresa Dahl, et al.
Published: (2026)
What Can We Learn From MIMO Graph Convolutions?
by: Roth, Andreas, et al.
Published: (2025)
by: Roth, Andreas, et al.
Published: (2025)
Reinforcement Learning with Rubric Anchors
by: Huang, Zenan, et al.
Published: (2025)
by: Huang, Zenan, et al.
Published: (2025)
Test-Time Alignment of LLMs via Sampling-Based Optimal Control in pre-logit space
by: Kanai, Sekitoshi, et al.
Published: (2025)
by: Kanai, Sekitoshi, et al.
Published: (2025)
AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning
by: Wu, Peilin, et al.
Published: (2026)
by: Wu, Peilin, et al.
Published: (2026)
Sequential Decision-Making for Inline Text Autocomplete
by: Chitnis, Rohan, et al.
Published: (2024)
by: Chitnis, Rohan, et al.
Published: (2024)
Off-Policy Evaluation and Learning for Matching Markets
by: Hayashi, Yudai, et al.
Published: (2025)
by: Hayashi, Yudai, et al.
Published: (2025)
What Can We Learn from State Space Models for Machine Learning on Graphs?
by: Huang, Yinan, et al.
Published: (2024)
by: Huang, Yinan, et al.
Published: (2024)
ODIA: Oriented Distillation for Inline Acceleration of LLM-based Function Calling
by: Zhang, Hanlong, et al.
Published: (2025)
by: Zhang, Hanlong, et al.
Published: (2025)
Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering
by: Imajuku, Yuki, et al.
Published: (2025)
by: Imajuku, Yuki, et al.
Published: (2025)
Learning to Judge: LLMs Designing and Applying Evaluation Rubrics
by: Siro, Clemencia, et al.
Published: (2026)
by: Siro, Clemencia, et al.
Published: (2026)
PCM Selector: Penalized Covariate-Mediator Selection Operator for Evaluating Linear Causal Effects
by: Nanmo, Hisayoshi, et al.
Published: (2024)
by: Nanmo, Hisayoshi, et al.
Published: (2024)
Bayesian Meta-Learning with Expert Feedback for Task-Shift Adaptation through Causal Embeddings
by: Mäkinen, Lotta, et al.
Published: (2026)
by: Mäkinen, Lotta, et al.
Published: (2026)
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
by: Li, Gaotang, et al.
Published: (2026)
by: Li, Gaotang, et al.
Published: (2026)
An Empirical Study of Pre-trained Model Selection for Out-of-Distribution Generalization and Calibration
by: Naganuma, Hiroki, et al.
Published: (2023)
by: Naganuma, Hiroki, et al.
Published: (2023)
DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
by: Yoshida, Kotaro, et al.
Published: (2025)
by: Yoshida, Kotaro, et al.
Published: (2025)
Can We Theoretically Quantify the Impacts of Local Updates on the Generalization Performance of Federated Learning?
by: Ju, Peizhong, et al.
Published: (2024)
by: Ju, Peizhong, et al.
Published: (2024)
Step-wise Rubric Rewards for LLM Reasoning
by: Xie, Weichu, et al.
Published: (2026)
by: Xie, Weichu, et al.
Published: (2026)
Agentic Rubrics as Contextual Verifiers for SWE Agents
by: Raghavendra, Mohit, et al.
Published: (2026)
by: Raghavendra, Mohit, et al.
Published: (2026)
Robust Reward Modeling via Causal Rubrics
by: Srivastava, Pragya, et al.
Published: (2025)
by: Srivastava, Pragya, et al.
Published: (2025)
Similar Items
-
Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search
by: Inoue, Yuichi, et al.
Published: (2025) -
Agent Skill Acquisition for Large Language Models via CycleQD
by: Kuroki, So, et al.
Published: (2024) -
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization
by: Nakamura, Taishi, et al.
Published: (2025) -
ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution
by: Lange, Robert Tjarko, et al.
Published: (2025) -
Robust Invariant Representation Learning by Distribution Extrapolation
by: Yoshida, Kotaro, et al.
Published: (2025)