ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yinjie, Yang, Ling, Li, Guohao, Wang, Mengdi, Aragam, Bryon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Co-Evolving LLM Coder and Unit Tester via Reinforcement Learning
by: Wang, Yinjie, et al.
Published: (2025)
by: Wang, Yinjie, et al.
Published: (2025)
OpenClaw-RL: Train Any Agent Simply by Talking
by: Wang, Yinjie, et al.
Published: (2026)
by: Wang, Yinjie, et al.
Published: (2026)
GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators
by: Guo, Jiacheng, et al.
Published: (2025)
by: Guo, Jiacheng, et al.
Published: (2025)
Do LLMs dream of elephants (when told not to)? Latent concept association and associative memory in transformers
by: Jiang, Yibo, et al.
Published: (2024)
by: Jiang, Yibo, et al.
Published: (2024)
Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models
by: Wang, Yinjie, et al.
Published: (2025)
by: Wang, Yinjie, et al.
Published: (2025)
FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
by: Xiao, Ruixuan, et al.
Published: (2024)
by: Xiao, Ruixuan, et al.
Published: (2024)
Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows
by: Chen, Hardy, et al.
Published: (2026)
by: Chen, Hardy, et al.
Published: (2026)
RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System
by: Wang, Yinjie, et al.
Published: (2026)
by: Wang, Yinjie, et al.
Published: (2026)
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
by: Wang, Yun, et al.
Published: (2026)
by: Wang, Yun, et al.
Published: (2026)
FlowCompile: An Optimizing Compiler for Structured LLM Workflows
by: Li, Junyan, et al.
Published: (2026)
by: Li, Junyan, et al.
Published: (2026)
On the Origins of Linear Representations in Large Language Models
by: Jiang, Yibo, et al.
Published: (2024)
by: Jiang, Yibo, et al.
Published: (2024)
Offloading Score: Measuring AI Reliance Through Counterfactual Workflows
by: Padmakumar, Vishakh, et al.
Published: (2026)
by: Padmakumar, Vishakh, et al.
Published: (2026)
Greedy equivalence search for nonparametric graphical models
by: Aragam, Bryon
Published: (2024)
by: Aragam, Bryon
Published: (2024)
On the Role of Preference Variance in Preference Optimization
by: Guo, Jiacheng, et al.
Published: (2025)
by: Guo, Jiacheng, et al.
Published: (2025)
Model-free Estimation of Latent Structure via Multiscale Nonparametric Maximum Likelihood
by: Aragam, Bryon, et al.
Published: (2024)
by: Aragam, Bryon, et al.
Published: (2024)
Calibrating LLMs with Preference Optimization on Thought Trees for Generating Rationale in Science Question Scoring
by: Li, Jiazheng, et al.
Published: (2024)
by: Li, Jiazheng, et al.
Published: (2024)
ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates
by: Yang, Ling, et al.
Published: (2025)
by: Yang, Ling, et al.
Published: (2025)
Expectation Confirmation Preference Optimization for Multi-Turn Conversational Recommendation Agent
by: Feng, Xueyang, et al.
Published: (2025)
by: Feng, Xueyang, et al.
Published: (2025)
Trait-Aware Policy Optimization for Autoregressive Multi-Trait Essay Scoring
by: Wang, Zhengyang, et al.
Published: (2026)
by: Wang, Zhengyang, et al.
Published: (2026)
Evaluating Scoring Bias in LLM-as-a-Judge
by: Li, Qingquan, et al.
Published: (2025)
by: Li, Qingquan, et al.
Published: (2025)
Improve LLM-based Automatic Essay Scoring with Linguistic Features
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
by: Hou, Zhaoyi Joey, et al.
Published: (2025)
An Adversary-Resistant Multi-Agent LLM System via Credibility Scoring
by: Ebrahimi, Sana, et al.
Published: (2025)
by: Ebrahimi, Sana, et al.
Published: (2025)
GEMA-Score: Granular Explainable Multi-Agent Scoring Framework for Radiology Report Evaluation
by: Zhang, Zhenxuan, et al.
Published: (2025)
by: Zhang, Zhenxuan, et al.
Published: (2025)
LLM Agents at the Roundtable: A Multi-Perspective and Dialectical Reasoning Framework for Essay Scoring
by: Jang, Jinhee, et al.
Published: (2025)
by: Jang, Jinhee, et al.
Published: (2025)
LLM-AutoDiff: Auto-Differentiate Any LLM Workflow
by: Yin, Li, et al.
Published: (2025)
by: Yin, Li, et al.
Published: (2025)
From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring
by: Nguyen, Minh Hoang, et al.
Published: (2026)
by: Nguyen, Minh Hoang, et al.
Published: (2026)
Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
by: Yang, Yunqiao, et al.
Published: (2025)
by: Yang, Yunqiao, et al.
Published: (2025)
FlowBot: Inducing LLM Workflows with Bilevel Optimization and Textual Gradients
by: Yu, Hongyeon, et al.
Published: (2026)
by: Yu, Hongyeon, et al.
Published: (2026)
From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
by: Yue, Ling, et al.
Published: (2026)
by: Yue, Ling, et al.
Published: (2026)
SAS: Simulated Attention Score
by: Zheng, Chuanyang, et al.
Published: (2025)
by: Zheng, Chuanyang, et al.
Published: (2025)
LLM-POTUS Score: A Framework of Analyzing Presidential Debates with Large Language Models
by: Liu, Zhengliang, et al.
Published: (2024)
by: Liu, Zhengliang, et al.
Published: (2024)
STELLA: Self-Evolving LLM Agent for Biomedical Research
by: Jin, Ruofan, et al.
Published: (2025)
by: Jin, Ruofan, et al.
Published: (2025)
WISE-Flow: Workflow-Induced Structured Experience for Self-Evolving Conversational Service Agents
by: Zhou, Yuqing, et al.
Published: (2026)
by: Zhou, Yuqing, et al.
Published: (2026)
Agent Workflow Memory
by: Wang, Zora Zhiruo, et al.
Published: (2024)
by: Wang, Zora Zhiruo, et al.
Published: (2024)
Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores
by: Blackwell, Robert E., et al.
Published: (2024)
by: Blackwell, Robert E., et al.
Published: (2024)
Rank-Then-Score: Enhancing Large Language Models for Automated Essay Scoring
by: Cai, Yida, et al.
Published: (2025)
by: Cai, Yida, et al.
Published: (2025)
MindFlow: Revolutionizing E-commerce Customer Support with Multimodal LLM Agents
by: Gong, Ming, et al.
Published: (2025)
by: Gong, Ming, et al.
Published: (2025)
DEPO: Dual-Efficiency Preference Optimization for LLM Agents
by: Chen, Sirui, et al.
Published: (2025)
by: Chen, Sirui, et al.
Published: (2025)
Pair2Score: Pairwise-to-Absolute Transfer for LLM-Based Essay Scoring
by: Hallaç, İbrahim Rıza, et al.
Published: (2026)
by: Hallaç, İbrahim Rıza, et al.
Published: (2026)
SPECTRA: Revealing the Full Spectrum of User Preferences via Distributional LLM Inference
by: Zhang, Luyang, et al.
Published: (2025)
by: Zhang, Luyang, et al.
Published: (2025)
Similar Items
-
Co-Evolving LLM Coder and Unit Tester via Reinforcement Learning
by: Wang, Yinjie, et al.
Published: (2025) -
OpenClaw-RL: Train Any Agent Simply by Talking
by: Wang, Yinjie, et al.
Published: (2026) -
GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators
by: Guo, Jiacheng, et al.
Published: (2025) -
Do LLMs dream of elephants (when told not to)? Latent concept association and associative memory in transformers
by: Jiang, Yibo, et al.
Published: (2024) -
Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models
by: Wang, Yinjie, et al.
Published: (2025)