LLMs Are In-Context Bandit Reinforcement Learners
Fuente:
arXiv
Salvato in:
| Autori principali: | Monea, Giovanni, Bosselut, Antoine, Brantley, Kianté, Artzi, Yoav |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Surprising Failure? Multimodal LLMs and the NLVR Challenge
di: Wu, Anne, et al.
Pubblicazione: (2024)
di: Wu, Anne, et al.
Pubblicazione: (2024)
Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons
di: Monea, Giovanni, et al.
Pubblicazione: (2025)
di: Monea, Giovanni, et al.
Pubblicazione: (2025)
No Mean Feat: Simple, Strong Baselines for Context Compression
di: Feldman, Yair, et al.
Pubblicazione: (2025)
di: Feldman, Yair, et al.
Pubblicazione: (2025)
Talk Less, Interact Better: Evaluating In-context Conversational Adaptation in Multimodal LLMs
di: Hua, Yilun, et al.
Pubblicazione: (2024)
di: Hua, Yilun, et al.
Pubblicazione: (2024)
Post-training for Efficient Communication via Convention Formation
di: Hua, Yilun, et al.
Pubblicazione: (2025)
di: Hua, Yilun, et al.
Pubblicazione: (2025)
Reward Is Enough: LLMs Are In-Context Reinforcement Learners
di: Song, Kefan, et al.
Pubblicazione: (2025)
di: Song, Kefan, et al.
Pubblicazione: (2025)
CoGen: Learning from Feedback with Coupled Comprehension and Generation
di: Gul, Mustafa Omer, et al.
Pubblicazione: (2024)
di: Gul, Mustafa Omer, et al.
Pubblicazione: (2024)
Policy-Gradient Training of Language Models for Ranking
di: Gao, Ge, et al.
Pubblicazione: (2023)
di: Gao, Ge, et al.
Pubblicazione: (2023)
Value-Guided Search for Efficient Chain-of-Thought Reasoning
di: Wang, Kaiwen, et al.
Pubblicazione: (2025)
di: Wang, Kaiwen, et al.
Pubblicazione: (2025)
Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
di: Bayazit, Deniz, et al.
Pubblicazione: (2025)
di: Bayazit, Deniz, et al.
Pubblicazione: (2025)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
di: Gao, Zhaolin, et al.
Pubblicazione: (2024)
di: Gao, Zhaolin, et al.
Pubblicazione: (2024)
Dataset Reset Policy Optimization for RLHF
di: Chang, Jonathan D., et al.
Pubblicazione: (2024)
di: Chang, Jonathan D., et al.
Pubblicazione: (2024)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
di: Zhou, Jin Peng, et al.
Pubblicazione: (2025)
di: Zhou, Jin Peng, et al.
Pubblicazione: (2025)
Reliable Evaluation and Benchmarks for Statement Autoformalization
di: Poiroux, Auguste, et al.
Pubblicazione: (2024)
di: Poiroux, Auguste, et al.
Pubblicazione: (2024)
Expressive Value Learning for Scalable Offline Reinforcement Learning
di: Espinosa-Dice, Nicolas, et al.
Pubblicazione: (2025)
di: Espinosa-Dice, Nicolas, et al.
Pubblicazione: (2025)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
di: Foroutan, Negar, et al.
Pubblicazione: (2025)
di: Foroutan, Negar, et al.
Pubblicazione: (2025)
ConLID: Supervised Contrastive Learning for Low-Resource Language Identification
di: Foroutan, Negar, et al.
Pubblicazione: (2025)
di: Foroutan, Negar, et al.
Pubblicazione: (2025)
Mixtures of In-Context Learners
di: Hong, Giwon, et al.
Pubblicazione: (2024)
di: Hong, Giwon, et al.
Pubblicazione: (2024)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
di: Bayazit, Deniz, et al.
Pubblicazione: (2023)
di: Bayazit, Deniz, et al.
Pubblicazione: (2023)
Knot So Simple: A Minimalistic Environment for Spatial Reasoning
di: Chen, Zizhao, et al.
Pubblicazione: (2025)
di: Chen, Zizhao, et al.
Pubblicazione: (2025)
A Logical Fallacy-Informed Framework for Argument Generation
di: Mouchel, Luca, et al.
Pubblicazione: (2024)
di: Mouchel, Luca, et al.
Pubblicazione: (2024)
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
di: Monea, Giovanni, et al.
Pubblicazione: (2023)
di: Monea, Giovanni, et al.
Pubblicazione: (2023)
Large Language Models are Miscalibrated In-Context Learners
di: Li, Chengzu, et al.
Pubblicazione: (2023)
di: Li, Chengzu, et al.
Pubblicazione: (2023)
Large Language Models are Biased Reinforcement Learners
di: Hayes, William M., et al.
Pubblicazione: (2024)
di: Hayes, William M., et al.
Pubblicazione: (2024)
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
di: Foroutan, Negar, et al.
Pubblicazione: (2025)
di: Foroutan, Negar, et al.
Pubblicazione: (2025)
Retrospective Learning from Interactions
di: Chen, Zizhao, et al.
Pubblicazione: (2024)
di: Chen, Zizhao, et al.
Pubblicazione: (2024)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
di: Xia, Fanzeng, et al.
Pubblicazione: (2024)
di: Xia, Fanzeng, et al.
Pubblicazione: (2024)
Online Personalizing White-box LLMs Generation with Neural Bandits
di: Chen, Zekai, et al.
Pubblicazione: (2024)
di: Chen, Zekai, et al.
Pubblicazione: (2024)
Evaluating Language Model Agency through Negotiations
di: Davidson, Tim R., et al.
Pubblicazione: (2024)
di: Davidson, Tim R., et al.
Pubblicazione: (2024)
Rational Metareasoning for Large Language Models
di: De Sabbata, C. Nicolò, et al.
Pubblicazione: (2024)
di: De Sabbata, C. Nicolò, et al.
Pubblicazione: (2024)
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
di: Gao, Silin, et al.
Pubblicazione: (2025)
di: Gao, Silin, et al.
Pubblicazione: (2025)
Pre-training Limited Memory Language Models with Internal and External Knowledge
di: Zhao, Linxi, et al.
Pubblicazione: (2025)
di: Zhao, Linxi, et al.
Pubblicazione: (2025)
Iterative Deployment Improves Planning Skills in LLMs
di: Corrêa, Augusto B., et al.
Pubblicazione: (2025)
di: Corrêa, Augusto B., et al.
Pubblicazione: (2025)
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation
di: Oertell, Owen, et al.
Pubblicazione: (2024)
di: Oertell, Owen, et al.
Pubblicazione: (2024)
LLMs as Zero-shot Graph Learners: Alignment of GNN Representations with LLM Token Embeddings
di: Wang, Duo, et al.
Pubblicazione: (2024)
di: Wang, Duo, et al.
Pubblicazione: (2024)
Is In-Context Learning Sufficient for Instruction Following in LLMs?
di: Zhao, Hao, et al.
Pubblicazione: (2024)
di: Zhao, Hao, et al.
Pubblicazione: (2024)
Learning and Enforcing Context-Sensitive Control for LLMs
di: Albinhassan, Mohammad, et al.
Pubblicazione: (2026)
di: Albinhassan, Mohammad, et al.
Pubblicazione: (2026)
Understanding Understanding: A Pragmatic Framework Motivated by Large Language Models
di: Leyton-Brown, Kevin, et al.
Pubblicazione: (2024)
di: Leyton-Brown, Kevin, et al.
Pubblicazione: (2024)
Adversarial Imitation Learning via Boosting
di: Chang, Jonathan D., et al.
Pubblicazione: (2024)
di: Chang, Jonathan D., et al.
Pubblicazione: (2024)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
di: Xiao, Chaojun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A Surprising Failure? Multimodal LLMs and the NLVR Challenge
di: Wu, Anne, et al.
Pubblicazione: (2024) -
Breadcrumbs Reasoning: Memory-Efficient Reasoning with Compression Beacons
di: Monea, Giovanni, et al.
Pubblicazione: (2025) -
No Mean Feat: Simple, Strong Baselines for Context Compression
di: Feldman, Yair, et al.
Pubblicazione: (2025) -
Talk Less, Interact Better: Evaluating In-context Conversational Adaptation in Multimodal LLMs
di: Hua, Yilun, et al.
Pubblicazione: (2024) -
Post-training for Efficient Communication via Convention Formation
di: Hua, Yilun, et al.
Pubblicazione: (2025)