Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gu, Jiawei, Liang, Shangsong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
von: Huang, Jen-tse, et al.
Veröffentlicht: (2024)
von: Huang, Jen-tse, et al.
Veröffentlicht: (2024)
On the Decision-Making Abilities in Role-Playing using Large Language Models
von: Shen, Chenglei, et al.
Veröffentlicht: (2024)
von: Shen, Chenglei, et al.
Veröffentlicht: (2024)
D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation
von: Zhou, Weibo, et al.
Veröffentlicht: (2025)
von: Zhou, Weibo, et al.
Veröffentlicht: (2025)
S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
von: Sun, Shaoning, et al.
Veröffentlicht: (2025)
von: Sun, Shaoning, et al.
Veröffentlicht: (2025)
CLEX: Continuous Length Extrapolation for Large Language Models
von: Chen, Guanzheng, et al.
Veröffentlicht: (2023)
von: Chen, Guanzheng, et al.
Veröffentlicht: (2023)
Effective Distillation of Table-based Reasoning Ability from LLMs
von: Yang, Bohao, et al.
Veröffentlicht: (2023)
von: Yang, Bohao, et al.
Veröffentlicht: (2023)
Towards Cost-Effective Reward Guided Text Generation
von: Rashid, Ahmad, et al.
Veröffentlicht: (2025)
von: Rashid, Ahmad, et al.
Veröffentlicht: (2025)
Cascaded Language Models for Cost-effective Human-AI Decision-Making
von: Fanconi, Claudio, et al.
Veröffentlicht: (2025)
von: Fanconi, Claudio, et al.
Veröffentlicht: (2025)
Cognitive Bias in Decision-Making with LLMs
von: Echterhoff, Jessica, et al.
Veröffentlicht: (2024)
von: Echterhoff, Jessica, et al.
Veröffentlicht: (2024)
S3D: A Simple and Cost-Effective Self-Speculative Decoding Scheme for Low-Memory GPUs
von: Zhong, Wei, et al.
Veröffentlicht: (2024)
von: Zhong, Wei, et al.
Veröffentlicht: (2024)
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2024)
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2024)
Cost-Effective Proxy Reward Model Construction with On-Policy and Active Learning
von: Chen, Yifang, et al.
Veröffentlicht: (2024)
von: Chen, Yifang, et al.
Veröffentlicht: (2024)
Evaluating the Bias in LLMs for Surveying Opinion and Decision Making in Healthcare
von: Khaokaew, Yonchanok, et al.
Veröffentlicht: (2025)
von: Khaokaew, Yonchanok, et al.
Veröffentlicht: (2025)
Cost-Aware Diffusion Draft Trees for Speculative Decoding
von: Zhang, Shuai, et al.
Veröffentlicht: (2026)
von: Zhang, Shuai, et al.
Veröffentlicht: (2026)
Boosting Reward Model with Preference-Conditional Multi-Aspect Synthetic Data Generation
von: Shen, Jiaming, et al.
Veröffentlicht: (2024)
von: Shen, Jiaming, et al.
Veröffentlicht: (2024)
Out-of-Vocabulary Sampling Boosts Speculative Decoding
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases
von: Qiu, Pengcheng, et al.
Veröffentlicht: (2025)
von: Qiu, Pengcheng, et al.
Veröffentlicht: (2025)
Cost-Effective Hallucination Detection for LLMs
von: Valentin, Simon, et al.
Veröffentlicht: (2024)
von: Valentin, Simon, et al.
Veröffentlicht: (2024)
Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
von: Li, Bolian, et al.
Veröffentlicht: (2025)
von: Li, Bolian, et al.
Veröffentlicht: (2025)
Haste Makes Waste: Evaluating Planning Abilities of LLMs for Efficient and Feasible Multitasking with Time Constraints Between Actions
von: Wu, Zirui, et al.
Veröffentlicht: (2025)
von: Wu, Zirui, et al.
Veröffentlicht: (2025)
Harnessing LLMs Explanations to Boost Surrogate Models in Tabular Data Classification
von: Shi, Ruxue, et al.
Veröffentlicht: (2025)
von: Shi, Ruxue, et al.
Veröffentlicht: (2025)
LETToT: Label-Free Evaluation of Large Language Models On Tourism Using Expert Tree-of-Thought
von: Qi, Ruiyan, et al.
Veröffentlicht: (2025)
von: Qi, Ruiyan, et al.
Veröffentlicht: (2025)
Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
von: Zheng, Qinqing, et al.
Veröffentlicht: (2024)
von: Zheng, Qinqing, et al.
Veröffentlicht: (2024)
Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative Decoding
von: Wang, Pei-Shuo, et al.
Veröffentlicht: (2025)
von: Wang, Pei-Shuo, et al.
Veröffentlicht: (2025)
Discrimination by LLMs: Cross-lingual Bias Assessment and Mitigation in Decision-Making and Summarisation
von: Huijzer, Willem, et al.
Veröffentlicht: (2025)
von: Huijzer, Willem, et al.
Veröffentlicht: (2025)
Efficient Reasoning for LLMs through Speculative Chain-of-Thought
von: Wang, Jikai, et al.
Veröffentlicht: (2025)
von: Wang, Jikai, et al.
Veröffentlicht: (2025)
Accelerating Production LLMs with Combined Token/Embedding Speculators
von: Wertheimer, Davis, et al.
Veröffentlicht: (2024)
von: Wertheimer, Davis, et al.
Veröffentlicht: (2024)
Cost-Efficient Estimation of General Abilities Across Benchmarks
von: Krumdick, Michael, et al.
Veröffentlicht: (2026)
von: Krumdick, Michael, et al.
Veröffentlicht: (2026)
Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
von: Kuang, Jiayi, et al.
Veröffentlicht: (2025)
von: Kuang, Jiayi, et al.
Veröffentlicht: (2025)
Self-Generated Critiques Boost Reward Modeling for Language Models
von: Yu, Yue, et al.
Veröffentlicht: (2024)
von: Yu, Yue, et al.
Veröffentlicht: (2024)
Intrinsic Mutual Information as a Modulator for Preference Optimization
von: Liao, Peng, et al.
Veröffentlicht: (2026)
von: Liao, Peng, et al.
Veröffentlicht: (2026)
Don't Ignore Dual Logic Ability of LLMs while Privatizing: A Data-Intensive Analysis in Medical Domain
von: Du, Yanrui, et al.
Veröffentlicht: (2023)
von: Du, Yanrui, et al.
Veröffentlicht: (2023)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
von: Li, Manling, et al.
Veröffentlicht: (2024)
von: Li, Manling, et al.
Veröffentlicht: (2024)
Exploring the Sensitivity of LLMs' Decision-Making Capabilities: Insights from Prompt Variation and Hyperparameters
von: Loya, Manikanta, et al.
Veröffentlicht: (2023)
von: Loya, Manikanta, et al.
Veröffentlicht: (2023)
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making
von: Kim, Yubin, et al.
Veröffentlicht: (2024)
von: Kim, Yubin, et al.
Veröffentlicht: (2024)
Speculating LLMs' Chinese Training Data Pollution from Their Tokens
von: Zhang, Qingjie, et al.
Veröffentlicht: (2025)
von: Zhang, Qingjie, et al.
Veröffentlicht: (2025)
Balancing Cost and Effectiveness of Synthetic Data Generation Strategies for LLMs
von: Chan, Yung-Chieh, et al.
Veröffentlicht: (2024)
von: Chan, Yung-Chieh, et al.
Veröffentlicht: (2024)
Temperature-Centric Investigation of Speculative Decoding with Knowledge Distillation
von: Ouyang, Siru, et al.
Veröffentlicht: (2024)
von: Ouyang, Siru, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
von: Huang, Jen-tse, et al.
Veröffentlicht: (2024) -
On the Decision-Making Abilities in Role-Playing using Large Language Models
von: Shen, Chenglei, et al.
Veröffentlicht: (2024) -
D-SCoRE: Document-Centric Segmentation and CoT Reasoning with Structured Export for QA-CoT Data Generation
von: Zhou, Weibo, et al.
Veröffentlicht: (2025) -
S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
von: Sun, Shaoning, et al.
Veröffentlicht: (2025) -
CLEX: Continuous Length Extrapolation for Large Language Models
von: Chen, Guanzheng, et al.
Veröffentlicht: (2023)