Online Personalizing White-box LLMs Generation with Neural Bandits
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Zekai, Daniel, Weeden, Chen, Po-yu, Buet-Golfouse, Francois |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Instruction Learning Paradigms: A Dual Perspective on White-box and Black-box LLMs
por: Ren, Yanwei, et al.
Publicado: (2025)
por: Ren, Yanwei, et al.
Publicado: (2025)
Bridging Online and Offline RL: Contextual Bandit Learning for Multi-Turn Code Generation
por: Chen, Ziru, et al.
Publicado: (2026)
por: Chen, Ziru, et al.
Publicado: (2026)
LLMs Are In-Context Bandit Reinforcement Learners
por: Monea, Giovanni, et al.
Publicado: (2024)
por: Monea, Giovanni, et al.
Publicado: (2024)
Learning to Correct for QA Reasoning with Black-box LLMs
por: Kim, Jaehyung, et al.
Publicado: (2024)
por: Kim, Jaehyung, et al.
Publicado: (2024)
Surrogate modeling for interpreting black-box LLMs in medical predictions
por: Han, Changho, et al.
Publicado: (2026)
por: Han, Changho, et al.
Publicado: (2026)
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements
por: Li, Ming, et al.
Publicado: (2024)
por: Li, Ming, et al.
Publicado: (2024)
Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
por: Lin, Xiaofeng, et al.
Publicado: (2026)
por: Lin, Xiaofeng, et al.
Publicado: (2026)
Automated Black-box Prompt Engineering for Personalized Text-to-Image Generation
por: He, Yutong, et al.
Publicado: (2024)
por: He, Yutong, et al.
Publicado: (2024)
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
por: Chen, Sanxing, et al.
Publicado: (2025)
por: Chen, Sanxing, et al.
Publicado: (2025)
Jump Starting Bandits with LLM-Generated Prior Knowledge
por: Alamdari, Parand A., et al.
Publicado: (2024)
por: Alamdari, Parand A., et al.
Publicado: (2024)
Differentially Private Fine-Tuning of Diffusion Models
por: Tsai, Yu-Lin, et al.
Publicado: (2024)
por: Tsai, Yu-Lin, et al.
Publicado: (2024)
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers
por: Huang, Zekai, et al.
Publicado: (2025)
por: Huang, Zekai, et al.
Publicado: (2025)
Selective Prompting Tuning for Personalized Conversations with LLMs
por: Huang, Qiushi, et al.
Publicado: (2024)
por: Huang, Qiushi, et al.
Publicado: (2024)
Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments
por: Zhang, Ziyuan, et al.
Publicado: (2025)
por: Zhang, Ziyuan, et al.
Publicado: (2025)
HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs
por: Li, Qing, et al.
Publicado: (2025)
por: Li, Qing, et al.
Publicado: (2025)
Few-shot Personalization of LLMs with Mis-aligned Responses
por: Kim, Jaehyung, et al.
Publicado: (2024)
por: Kim, Jaehyung, et al.
Publicado: (2024)
Rewarding Graph Reasoning Process makes LLMs more Generalized Reasoners
por: Peng, Miao, et al.
Publicado: (2025)
por: Peng, Miao, et al.
Publicado: (2025)
LoRe: Personalizing LLMs via Low-Rank Reward Modeling
por: Bose, Avinandan, et al.
Publicado: (2025)
por: Bose, Avinandan, et al.
Publicado: (2025)
Understanding Chain-of-Thought in LLMs through Information Theory
por: Ton, Jean-Francois, et al.
Publicado: (2024)
por: Ton, Jean-Francois, et al.
Publicado: (2024)
Can Post-Training Transform LLMs into Causal Reasoners?
por: Chen, Junqi, et al.
Publicado: (2026)
por: Chen, Junqi, et al.
Publicado: (2026)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
por: Liu, Yijun, et al.
Publicado: (2024)
por: Liu, Yijun, et al.
Publicado: (2024)
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
por: Jiang, Dongwei, et al.
Publicado: (2024)
por: Jiang, Dongwei, et al.
Publicado: (2024)
A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms
por: Li, Yingru, et al.
Publicado: (2025)
por: Li, Yingru, et al.
Publicado: (2025)
Knowledge Editing on Black-box Large Language Models
por: Song, Xiaoshuai, et al.
Publicado: (2024)
por: Song, Xiaoshuai, et al.
Publicado: (2024)
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
por: Xu, Wanghan, et al.
Publicado: (2025)
por: Xu, Wanghan, et al.
Publicado: (2025)
RadioRAG: Online Retrieval-augmented Generation for Radiology Question Answering
por: Arasteh, Soroosh Tayebi, et al.
Publicado: (2024)
por: Arasteh, Soroosh Tayebi, et al.
Publicado: (2024)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
por: Mahdavi, Sadegh, et al.
Publicado: (2025)
por: Mahdavi, Sadegh, et al.
Publicado: (2025)
Learning Retrieval Augmentation for Personalized Dialogue Generation
por: Huang, Qiushi, et al.
Publicado: (2024)
por: Huang, Qiushi, et al.
Publicado: (2024)
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
por: Zhou, Cai, et al.
Publicado: (2026)
por: Zhou, Cai, et al.
Publicado: (2026)
A Comprehensive Sustainable Framework for Machine Learning and Artificial Intelligence
por: Pagliari, Roberto, et al.
Publicado: (2024)
por: Pagliari, Roberto, et al.
Publicado: (2024)
PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching
por: Chen, Ruishuo, et al.
Publicado: (2026)
por: Chen, Ruishuo, et al.
Publicado: (2026)
Can LLMs Follow Simple Rules?
por: Mu, Norman, et al.
Publicado: (2023)
por: Mu, Norman, et al.
Publicado: (2023)
Memory-Efficient LLM Training with Online Subspace Descent
por: Liang, Kaizhao, et al.
Publicado: (2024)
por: Liang, Kaizhao, et al.
Publicado: (2024)
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning
por: Chen, Zui, et al.
Publicado: (2024)
por: Chen, Zui, et al.
Publicado: (2024)
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
por: Liu, Wanlong, et al.
Publicado: (2024)
por: Liu, Wanlong, et al.
Publicado: (2024)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
por: Xia, Fanzeng, et al.
Publicado: (2024)
por: Xia, Fanzeng, et al.
Publicado: (2024)
Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction
por: Tsaknakis, Ioannis, et al.
Publicado: (2025)
por: Tsaknakis, Ioannis, et al.
Publicado: (2025)
No One Size Fits All: QueryBandits for Hallucination Mitigation
por: Cho, Nicole, et al.
Publicado: (2026)
por: Cho, Nicole, et al.
Publicado: (2026)
Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization
por: Xiong, Boya, et al.
Publicado: (2025)
por: Xiong, Boya, et al.
Publicado: (2025)
Sample-Efficient Alignment for LLMs
por: Liu, Zichen, et al.
Publicado: (2024)
por: Liu, Zichen, et al.
Publicado: (2024)
Ejemplares similares
-
Instruction Learning Paradigms: A Dual Perspective on White-box and Black-box LLMs
por: Ren, Yanwei, et al.
Publicado: (2025) -
Bridging Online and Offline RL: Contextual Bandit Learning for Multi-Turn Code Generation
por: Chen, Ziru, et al.
Publicado: (2026) -
LLMs Are In-Context Bandit Reinforcement Learners
por: Monea, Giovanni, et al.
Publicado: (2024) -
Learning to Correct for QA Reasoning with Black-box LLMs
por: Kim, Jaehyung, et al.
Publicado: (2024) -
Surrogate modeling for interpreting black-box LLMs in medical predictions
por: Han, Changho, et al.
Publicado: (2026)