Predicting Rewards Alongside Tokens: Non-disruptive Parameter Insertion for Efficient Inference Intervention in Large Language Model
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yuan, Chenhan, Huang, Fei, Peng, Ru, Lu, Keming, Yu, Bowen, Zhou, Chang, Zhou, Jingren |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment
par: Lu, Keming, et autres
Publié: (2024)
par: Lu, Keming, et autres
Publié: (2024)
Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
par: Dong, Guanting, et autres
Publié: (2024)
par: Dong, Guanting, et autres
Publié: (2024)
Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment
par: Lu, Keming, et autres
Publié: (2024)
par: Lu, Keming, et autres
Publié: (2024)
How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition
par: Dong, Guanting, et autres
Publié: (2023)
par: Dong, Guanting, et autres
Publié: (2023)
Speculative Contrastive Decoding
par: Yuan, Hongyi, et autres
Publié: (2023)
par: Yuan, Hongyi, et autres
Publié: (2023)
Self-Steering Optimization: Autonomous Preference Optimization for Large Language Models
par: Xiang, Hao, et autres
Publié: (2024)
par: Xiang, Hao, et autres
Publié: (2024)
Language Confusion Gate: Language-Aware Decoding Through Model Self-Distillation
par: Zhang, Collin, et autres
Publié: (2025)
par: Zhang, Collin, et autres
Publié: (2025)
AutoLogi: Automated Generation of Logic Puzzles for Evaluating Reasoning Abilities of Large Language Models
par: Zhu, Qin, et autres
Publié: (2025)
par: Zhu, Qin, et autres
Publié: (2025)
CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
par: Hu, Xiaomeng, et autres
Publié: (2025)
par: Hu, Xiaomeng, et autres
Publié: (2025)
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
par: Chen, Yanxi, et autres
Publié: (2023)
par: Chen, Yanxi, et autres
Publié: (2023)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
par: Huang, Xiaohu, et autres
Publié: (2024)
par: Huang, Xiaohu, et autres
Publié: (2024)
SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models
par: Lu, Xudong, et autres
Publié: (2024)
par: Lu, Xudong, et autres
Publié: (2024)
A Unified View of Delta Parameter Editing in Post-Trained Large-Scale Models
par: Tang, Qiaoyu, et autres
Publié: (2024)
par: Tang, Qiaoyu, et autres
Publié: (2024)
Inspo: Writing with Crowds Alongside AI
par: Huang, Chieh-Yang, et autres
Publié: (2023)
par: Huang, Chieh-Yang, et autres
Publié: (2023)
ProcessBench: Identifying Process Errors in Mathematical Reasoning
par: Zheng, Chujie, et autres
Publié: (2024)
par: Zheng, Chujie, et autres
Publié: (2024)
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
par: Qiu, Zihan, et autres
Publié: (2025)
par: Qiu, Zihan, et autres
Publié: (2025)
Efficient Pre-Training with Token Superposition
par: Peng, Bowen, et autres
Publié: (2026)
par: Peng, Bowen, et autres
Publié: (2026)
An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
par: Chen, Liang, et autres
Publié: (2024)
par: Chen, Liang, et autres
Publié: (2024)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
par: Zhang, Zhenru, et autres
Publié: (2025)
par: Zhang, Zhenru, et autres
Publié: (2025)
Spatio-Temporal Token Pruning for Efficient High-Resolution GUI Agents
par: Xu, Zhou, et autres
Publié: (2026)
par: Xu, Zhou, et autres
Publié: (2026)
Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning
par: Guan, Xin, et autres
Publié: (2026)
par: Guan, Xin, et autres
Publié: (2026)
Provably Efficient Online RLHF with One-Pass Reward Modeling
par: Li, Long-Fei, et autres
Publié: (2025)
par: Li, Long-Fei, et autres
Publié: (2025)
ReaLM: Reliable and Efficient Large Language Model Inference with Statistical Algorithm-Based Fault Tolerance
par: Xie, Tong, et autres
Publié: (2025)
par: Xie, Tong, et autres
Publié: (2025)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
par: Shao, Chenze, et autres
Publié: (2024)
par: Shao, Chenze, et autres
Publié: (2024)
AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator
par: Fan, Zhihao, et autres
Publié: (2024)
par: Fan, Zhihao, et autres
Publié: (2024)
Automated Profile Inference with Language Model Agents
par: Du, Yuntao, et autres
Publié: (2025)
par: Du, Yuntao, et autres
Publié: (2025)
Revealing Behavioral Plasticity in Large Language Models: A Token-Conditional Perspective
par: Mao, Liyuan, et autres
Publié: (2026)
par: Mao, Liyuan, et autres
Publié: (2026)
Theosis Within, Alongside, and Outside the Bible
par: John W. Martens
Publié: (2026)
par: John W. Martens
Publié: (2026)
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution
par: Li, Jiahui, et autres
Publié: (2024)
par: Li, Jiahui, et autres
Publié: (2024)
Efficient Temporal Tokenization for Mobility Prediction with Large Language Models
par: He, Haoyu, et autres
Publié: (2025)
par: He, Haoyu, et autres
Publié: (2025)
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
par: Yuan, Zenghui, et autres
Publié: (2025)
par: Yuan, Zenghui, et autres
Publié: (2025)
Are Large Language Models True Healthcare Jacks-of-All-Trades? Benchmarking Across Health Professions Beyond Physician Exams
par: Luo, Zheheng, et autres
Publié: (2024)
par: Luo, Zheheng, et autres
Publié: (2024)
Efficient Hybrid Inference for LLMs: Reward-Based Token Modelling with Selective Cloud Assistance
par: MS, Adarsh, et autres
Publié: (2024)
par: MS, Adarsh, et autres
Publié: (2024)
Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation
par: Zhu, Qin, et autres
Publié: (2024)
par: Zhu, Qin, et autres
Publié: (2024)
STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference
par: Guo, Yichen, et autres
Publié: (2025)
par: Guo, Yichen, et autres
Publié: (2025)
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
par: Yang, Qian, et autres
Publié: (2024)
par: Yang, Qian, et autres
Publié: (2024)
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
par: Zhang, Yuan, et autres
Publié: (2024)
par: Zhang, Yuan, et autres
Publié: (2024)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
par: Chen, Yanxi, et autres
Publié: (2024)
par: Chen, Yanxi, et autres
Publié: (2024)
Multi-Token Residual Prediction
par: Xu, Yufeng, et autres
Publié: (2026)
par: Xu, Yufeng, et autres
Publié: (2026)
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
par: Yang, Shu-wen, et autres
Publié: (2025)
par: Yang, Shu-wen, et autres
Publié: (2025)
Documents similaires
-
Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-Alignment
par: Lu, Keming, et autres
Publié: (2024) -
Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
par: Dong, Guanting, et autres
Publié: (2024) -
Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment
par: Lu, Keming, et autres
Publié: (2024) -
How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition
par: Dong, Guanting, et autres
Publié: (2023) -
Speculative Contrastive Decoding
par: Yuan, Hongyi, et autres
Publié: (2023)