Toward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Chuang, Yun-Shiuan, Kulkarni, Chaitanya, Chiu, Alec, Thangali, Avinash, Pan, Zijie, Shekhar, Shivani, Ge, Yirou, Li, Yixi, Kona, Uma, Pang, Linsey, Mehrotra, Prakhar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ada-RS: Adaptive Rejection Sampling for Selective Thinking
by: Ge, Yirou, et al.
Published: (2026)
by: Ge, Yirou, et al.
Published: (2026)
NEMO-4-PAYPAL: Leveraging NVIDIA's Nemo Framework for empowering PayPal's Commerce Agent
by: Garg, Sudhanshu, et al.
Published: (2025)
by: Garg, Sudhanshu, et al.
Published: (2025)
Backdoor Collapse: Eliminating Unknown Threats via Known Backdoor Aggregation in Language Models
by: Lin, Liang, et al.
Published: (2025)
by: Lin, Liang, et al.
Published: (2025)
MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in Model Context Protocol Agents
by: Zhou, Zhenhong, et al.
Published: (2026)
by: Zhou, Zhenhong, et al.
Published: (2026)
THE MAKING OF MODERNITY: VIOLENCE AND SOCIAL REVOLUTION IN THE SOUTH ASIAN CONTEXT
by: Prakash Kona
Published: (2018)
by: Prakash Kona
Published: (2018)
Calibrating Attribution Proxies for Reward Allocation in Participatory Weather Sensing
by: Ballandies, Mark C., et al.
Published: (2026)
by: Ballandies, Mark C., et al.
Published: (2026)
Philosophy of Health Sciences Library Management--A Panel Discussion
by: Kona, Martha, et al.
Published: (1976)
by: Kona, Martha, et al.
Published: (1976)
$C^\ast$-extreme points of unital completely positive maps invariant under group action
by: Kulkarni, Chaitanya J.
Published: (2026)
by: Kulkarni, Chaitanya J.
Published: (2026)
Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
PROXIMA: A Reliability Scoring Framework for Proxy Metrics in Online Controlled Experiments
by: Amudala, Avinash
Published: (2026)
by: Amudala, Avinash
Published: (2026)
Youth Smart-City Readiness — replication package
by: Kóňa, Andrej, et al.
Published: (2026)
by: Kóňa, Andrej, et al.
Published: (2026)
Bootstrapped Mixed Rewards for RL Post-Training: Injecting Canonical Action Order
by: Gupta, Prakhar, et al.
Published: (2025)
by: Gupta, Prakhar, et al.
Published: (2025)
Pivoting Retail Supply Chain with Deep Generative Techniques: Taxonomy, Survey and Insights
by: Wang, Yuan, et al.
Published: (2024)
by: Wang, Yuan, et al.
Published: (2024)
SuperdropNet: a Stable and Accurate Machine Learning Proxy for Droplet-based Cloud Microphysics
by: Sharma, Shivani, et al.
Published: (2024)
by: Sharma, Shivani, et al.
Published: (2024)
Structural Gender Bias in Credit Scoring: Proxy Leakage
by: SD, Navya, et al.
Published: (2026)
by: SD, Navya, et al.
Published: (2026)
On (2,2)-decomposable genus 4 Jacobians
by: Bruin, Nils, et al.
Published: (2023)
by: Bruin, Nils, et al.
Published: (2023)
Generalized Orthogonal Measures on the Space of Unital Completely Positive Maps
by: Bhattacharya, Angshuman, et al.
Published: (2022)
by: Bhattacharya, Angshuman, et al.
Published: (2022)
Barycentric decompositions in the space of weak expectations
by: Bhattacharya, Angshuman, et al.
Published: (2022)
by: Bhattacharya, Angshuman, et al.
Published: (2022)
Ergodic decomposition in the space of unital completely positive maps
by: Bhattacharya, Angshuman, et al.
Published: (2023)
by: Bhattacharya, Angshuman, et al.
Published: (2023)
vApps: Verifiable Applications at Internet Scale
by: Zhang, Isaac, et al.
Published: (2025)
by: Zhang, Isaac, et al.
Published: (2025)
The Delusional Hedge Algorithm as a Model of Human Learning from Diverse Opinions
by: Chuang, Yun-Shiuan, et al.
Published: (2024)
by: Chuang, Yun-Shiuan, et al.
Published: (2024)
From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agents
by: Gao, Jiaxuan, et al.
Published: (2026)
by: Gao, Jiaxuan, et al.
Published: (2026)
Truncated Step-Level Sampling with Process Rewards for Retrieval-Augmented Reasoning
by: Samarinas, Chris, et al.
Published: (2026)
by: Samarinas, Chris, et al.
Published: (2026)
Reimagining Self-Adaptation in the Age of Large Language Models
by: Donakanti, Raghav, et al.
Published: (2024)
by: Donakanti, Raghav, et al.
Published: (2024)
Extreme points of unital completely positive maps invariant under partial action
by: Kulkarni, Chaitanya J., et al.
Published: (2025)
by: Kulkarni, Chaitanya J., et al.
Published: (2025)
Direct integral of locally Hilbert spaces
by: Kulkarni, Chaitanya J., et al.
Published: (2025)
by: Kulkarni, Chaitanya J., et al.
Published: (2025)
Direct integral of locally Hilbert spaces over a locally measure space
by: Kulkarni, Chaitanya J., et al.
Published: (2025)
by: Kulkarni, Chaitanya J., et al.
Published: (2025)
On the Haagerup property for partial crossed products
by: Hossain, Md Amir, et al.
Published: (2026)
by: Hossain, Md Amir, et al.
Published: (2026)
Direct Integral and Decompoisitions of Locally Hilbert spaces
by: Kulkarni, Chaitanya J., et al.
Published: (2024)
by: Kulkarni, Chaitanya J., et al.
Published: (2024)
Sequential Recommendation via Adaptive Robust Attention with Multi-dimensional Embeddings
by: Pang, Linsey, et al.
Published: (2024)
by: Pang, Linsey, et al.
Published: (2024)
Logical Relations for Formally Verified Authenticated Data Structures
by: Gregersen, Simon Oddershede, et al.
Published: (2025)
by: Gregersen, Simon Oddershede, et al.
Published: (2025)
Reward Hacking Mitigation using Verifiable Composite Rewards
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
KnotResolver: Tracking self-intersecting filaments in microscopy using directed graphs
by: Khatri, Dhruv, et al.
Published: (2024)
by: Khatri, Dhruv, et al.
Published: (2024)
Physical Principles of Size and Frequency Scaling of Active Cytoskeletal Spirals
by: Soni, Aman, et al.
Published: (2025)
by: Soni, Aman, et al.
Published: (2025)
Robust Optimization for Mitigating Reward Hacking with Correlated Proxies
by: Liu, Zixuan, et al.
Published: (2026)
by: Liu, Zixuan, et al.
Published: (2026)
Rethinking the Role of Proxy Rewards in Language Model Alignment
by: Kim, Sungdong, et al.
Published: (2024)
by: Kim, Sungdong, et al.
Published: (2024)
Translation as a Scalable Proxy for Multilingual Evaluation
by: Issaka, Sheriff, et al.
Published: (2026)
by: Issaka, Sheriff, et al.
Published: (2026)
Call for a Pacific Centre for Disease Control!
by: George P. Drewett, et al.
Published: (2025)
by: George P. Drewett, et al.
Published: (2025)
Verifiable Process Rewards for Agentic Reasoning
by: Yuan, Huining, et al.
Published: (2026)
by: Yuan, Huining, et al.
Published: (2026)
Temporal Context Awareness: A Defense Framework Against Multi-turn Manipulation Attacks on Large Language Models
by: Kulkarni, Prashant, et al.
Published: (2025)
by: Kulkarni, Prashant, et al.
Published: (2025)
Similar Items
-
Ada-RS: Adaptive Rejection Sampling for Selective Thinking
by: Ge, Yirou, et al.
Published: (2026) -
NEMO-4-PAYPAL: Leveraging NVIDIA's Nemo Framework for empowering PayPal's Commerce Agent
by: Garg, Sudhanshu, et al.
Published: (2025) -
Backdoor Collapse: Eliminating Unknown Threats via Known Backdoor Aggregation in Language Models
by: Lin, Liang, et al.
Published: (2025) -
MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in Model Context Protocol Agents
by: Zhou, Zhenhong, et al.
Published: (2026) -
THE MAKING OF MODERNITY: VIOLENCE AND SOCIAL REVOLUTION IN THE SOUTH ASIAN CONTEXT
by: Prakash Kona
Published: (2018)