BOW: Reinforcement Learning for Bottlenecked Next Word Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Ming, Xu, Zhikun, Dineen, Jacob, Ye, Xiao, Zhou, Ben |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ToW: Thoughts of Words Improve Reasoning in Large Language Models
by: Xu, Zhikun, et al.
Published: (2024)
by: Xu, Zhikun, et al.
Published: (2024)
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
by: Dineen, Jacob, et al.
Published: (2026)
by: Dineen, Jacob, et al.
Published: (2026)
CC-LEARN: Cohort-based Consistency Learning
by: Ye, Xiao, et al.
Published: (2025)
by: Ye, Xiao, et al.
Published: (2025)
QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
by: Dineen, Jacob, et al.
Published: (2025)
by: Dineen, Jacob, et al.
Published: (2025)
On Support Samples of Next Word Prediction
by: Li, Yuqian, et al.
Published: (2025)
by: Li, Yuqian, et al.
Published: (2025)
Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
by: Ye, Xiao, et al.
Published: (2025)
by: Ye, Xiao, et al.
Published: (2025)
Can You Learn Semantics Through Next-Word Prediction? The Case of Entailment
by: Merrill, William, et al.
Published: (2024)
by: Merrill, William, et al.
Published: (2024)
Reliable Use of Lemmas via Eligibility Reasoning and Section$-$Aware Reinforcement Learning
by: Xu, Zhikun, et al.
Published: (2026)
by: Xu, Zhikun, et al.
Published: (2026)
Unbiased Visual Reasoning with Controlled Visual Inputs
by: Li, Zhaonan, et al.
Published: (2025)
by: Li, Zhaonan, et al.
Published: (2025)
RECAP: Transparent Inference-Time Emotion Alignment for Medical Dialogue Systems
by: Srinivasan, Adarsh, et al.
Published: (2025)
by: Srinivasan, Adarsh, et al.
Published: (2025)
Expression Syntax Information Bottleneck for Math Word Problems
by: Xiong, Jing, et al.
Published: (2023)
by: Xiong, Jing, et al.
Published: (2023)
HieroLM: Egyptian Hieroglyph Recovery with Next Word Prediction Language Model
by: Cai, Xuheng, et al.
Published: (2025)
by: Cai, Xuheng, et al.
Published: (2025)
Reinforced Fast Weights with Next-Sequence Prediction
by: Hwang, Hee Seung, et al.
Published: (2026)
by: Hwang, Hee Seung, et al.
Published: (2026)
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
by: Liu, Qin, et al.
Published: (2025)
by: Liu, Qin, et al.
Published: (2025)
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
by: Ilia, Evgenia, et al.
Published: (2024)
by: Ilia, Evgenia, et al.
Published: (2024)
Fix the Structural Bottleneck: Context Compression via Explicit Information Transmission
by: Ye, Jiangnan, et al.
Published: (2026)
by: Ye, Jiangnan, et al.
Published: (2026)
ThinkTuning: Instilling Cognitive Reflections without Distillation
by: RRV, Aswin, et al.
Published: (2025)
by: RRV, Aswin, et al.
Published: (2025)
Policy Learning with a Language Bottleneck
by: Srivastava, Megha, et al.
Published: (2024)
by: Srivastava, Megha, et al.
Published: (2024)
From Words to Worth: Newborn Article Impact Prediction with LLM
by: Zhao, Penghai, et al.
Published: (2024)
by: Zhao, Penghai, et al.
Published: (2024)
Deepfake Word Detection by Next-token Prediction using Fine-tuned Whisper
by: Tran, Hoan My, et al.
Published: (2026)
by: Tran, Hoan My, et al.
Published: (2026)
Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models
by: Li, Zongqian, et al.
Published: (2026)
by: Li, Zongqian, et al.
Published: (2026)
On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms
by: Trauger, Jacob, et al.
Published: (2025)
by: Trauger, Jacob, et al.
Published: (2025)
QoNext: Towards Next-generation QoE for Foundation Models
by: Guo, Yijin, et al.
Published: (2025)
by: Guo, Yijin, et al.
Published: (2025)
Reinforcement World Model Learning for LLM-based Agents
by: Yu, Xiao, et al.
Published: (2026)
by: Yu, Xiao, et al.
Published: (2026)
Next Word Suggestion using Graph Neural Network
by: Magar, Abisha Thapa, et al.
Published: (2025)
by: Magar, Abisha Thapa, et al.
Published: (2025)
Label Words as Local Task Vectors in In-Context Learning
by: Zheng, Bowen, et al.
Published: (2024)
by: Zheng, Bowen, et al.
Published: (2024)
We Need Knowledge Distillation for Solving Math Word Problems
by: Shen, Zhenquan, et al.
Published: (2025)
by: Shen, Zhenquan, et al.
Published: (2025)
From Graph to Word Bag: Introducing Domain Knowledge to Confusing Charge Prediction
by: Li, Ang, et al.
Published: (2024)
by: Li, Ang, et al.
Published: (2024)
Visual Grounding Helps Learn Word Meanings in Low-Data Regimes
by: Zhuang, Chengxu, et al.
Published: (2023)
by: Zhuang, Chengxu, et al.
Published: (2023)
TAT-R1: Terminology-Aware Translation with Reinforcement Learning and Word Alignment
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
CLIF: Concept-Level Influence Functions for Transparent Bottleneck Models
by: Sun, Yike, et al.
Published: (2026)
by: Sun, Yike, et al.
Published: (2026)
Predicting the Target Word of Game-playing Conversations using a Low-Rank Dialect Adapter for Decoder Models
by: Srirag, Dipankar, et al.
Published: (2024)
by: Srirag, Dipankar, et al.
Published: (2024)
Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering
by: Fu, Xingyu, et al.
Published: (2023)
by: Fu, Xingyu, et al.
Published: (2023)
Cross-Lingual Word Alignment for ASEAN Languages with Contrastive Learning
by: Zhang, Jingshen, et al.
Published: (2024)
by: Zhang, Jingshen, et al.
Published: (2024)
Co-Evolving LLM Coder and Unit Tester via Reinforcement Learning
by: Wang, Yinjie, et al.
Published: (2025)
by: Wang, Yinjie, et al.
Published: (2025)
Skill Reuse as Compression in Agentic RL
by: Xu, Zhikun, et al.
Published: (2026)
by: Xu, Zhikun, et al.
Published: (2026)
Learning Intrinsic Dimension via Information Bottleneck for Explainable Aspect-based Sentiment Analysis
by: Cheng, Zhenxiao, et al.
Published: (2024)
by: Cheng, Zhenxiao, et al.
Published: (2024)
Reasoning Bias of Next Token Prediction Training
by: Lin, Pengxiao, et al.
Published: (2025)
by: Lin, Pengxiao, et al.
Published: (2025)
Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models
by: Wang, Yinjie, et al.
Published: (2025)
by: Wang, Yinjie, et al.
Published: (2025)
Beyond Next Word Prediction: Developing Comprehensive Evaluation Frameworks for measuring LLM performance on real world applications
by: Agrawal, Vishakha, et al.
Published: (2025)
by: Agrawal, Vishakha, et al.
Published: (2025)
Similar Items
-
ToW: Thoughts of Words Improve Reasoning in Large Language Models
by: Xu, Zhikun, et al.
Published: (2024) -
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
by: Dineen, Jacob, et al.
Published: (2026) -
CC-LEARN: Cohort-based Consistency Learning
by: Ye, Xiao, et al.
Published: (2025) -
QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
by: Dineen, Jacob, et al.
Published: (2025) -
On Support Samples of Next Word Prediction
by: Li, Yuqian, et al.
Published: (2025)