Learning to Reason in LLMs by Expectation Maximization
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Junghyun, Kveton, Branislav, Rao, Anup, Mukherjee, Subhojyoti, Rossi, Ryan A., Choudhary, Sunav, Siu, Alexa |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2025)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2025)
Experimental Design for Active Transductive Inference in Large Language Models
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
Efficient and Interpretable Bandit Algorithms
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2023)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2023)
Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
Partial Policy Gradients for RL in LLMs
di: Mathur, Puneet, et al.
Pubblicazione: (2026)
di: Mathur, Puneet, et al.
Pubblicazione: (2026)
Agentic Planning with Reasoning for Image Styling via Offline RL
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2026)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2026)
Quantitative LLM Judges
di: Sahoo, Aishwarya, et al.
Pubblicazione: (2025)
di: Sahoo, Aishwarya, et al.
Pubblicazione: (2025)
Off-Policy Evaluation from Logged Human Feedback
di: Bhargava, Aniruddha, et al.
Pubblicazione: (2024)
di: Bhargava, Aniruddha, et al.
Pubblicazione: (2024)
RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
di: Fernandez, Nigel, et al.
Pubblicazione: (2025)
di: Fernandez, Nigel, et al.
Pubblicazione: (2025)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
di: Lee, Daeun, et al.
Pubblicazione: (2025)
di: Lee, Daeun, et al.
Pubblicazione: (2025)
FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain
di: Deb, Rohan, et al.
Pubblicazione: (2025)
di: Deb, Rohan, et al.
Pubblicazione: (2025)
Logits are All We Need to Adapt Closed Models
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
di: Hiranandani, Gaurush, et al.
Pubblicazione: (2025)
Hallucination Diversity-Aware Active Learning for Text Summarization
di: Xia, Yu, et al.
Pubblicazione: (2024)
di: Xia, Yu, et al.
Pubblicazione: (2024)
ML-Tool-Bench: Tool-Augmented Planning for ML Tasks
di: Chittepu, Yaswanth, et al.
Pubblicazione: (2025)
di: Chittepu, Yaswanth, et al.
Pubblicazione: (2025)
Optimal Design for Human Preference Elicitation
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
From Selection to Generation: A Survey of LLM-based Active Learning
di: Xia, Yu, et al.
Pubblicazione: (2025)
di: Xia, Yu, et al.
Pubblicazione: (2025)
Learning from a single labeled face and a stream of unlabeled data
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
di: Kveton, Branislav, et al.
Pubblicazione: (2026)
Stepwise Credit Assignment for GRPO on Flow-Matching Models
di: Savani, Yash, et al.
Pubblicazione: (2026)
di: Savani, Yash, et al.
Pubblicazione: (2026)
MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization
di: Tanjim, Md Mehrab, et al.
Pubblicazione: (2026)
di: Tanjim, Md Mehrab, et al.
Pubblicazione: (2026)
Few-Shot Graph Out-of-Distribution Detection with LLMs
di: Xu, Haoyan, et al.
Pubblicazione: (2025)
di: Xu, Haoyan, et al.
Pubblicazione: (2025)
To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization
di: Wang, Haozhe, et al.
Pubblicazione: (2025)
di: Wang, Haozhe, et al.
Pubblicazione: (2025)
Learning When to Attend: Conditional Memory Access for Long-Context LLMs
di: Choudhary, Sakshi, et al.
Pubblicazione: (2026)
di: Choudhary, Sakshi, et al.
Pubblicazione: (2026)
Fake or Compromised? Making Sense of Malicious Clients in Federated Learning
di: Mozaffari, Hamid, et al.
Pubblicazione: (2024)
di: Mozaffari, Hamid, et al.
Pubblicazione: (2024)
OCEAN: Offline Chain-of-thought Evaluation and Alignment in Large Language Models
di: Wu, Junda, et al.
Pubblicazione: (2024)
di: Wu, Junda, et al.
Pubblicazione: (2024)
Detecting Training Data of Large Language Models via Expectation Maximization
di: Kim, Gyuwan, et al.
Pubblicazione: (2024)
di: Kim, Gyuwan, et al.
Pubblicazione: (2024)
Steering MoE LLMs via Expert (De)Activation
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2025)
di: Fayyaz, Mohsen, et al.
Pubblicazione: (2025)
Next Word Suggestion using Graph Neural Network
di: Magar, Abisha Thapa, et al.
Pubblicazione: (2025)
di: Magar, Abisha Thapa, et al.
Pubblicazione: (2025)
REA-RL: Reflection-Aware Online Reinforcement Learning for Efficient Reasoning
di: Deng, Hexuan, et al.
Pubblicazione: (2025)
di: Deng, Hexuan, et al.
Pubblicazione: (2025)
A Survey on Stability of Learning with Limited Labelled Data and its Sensitivity to the Effects of Randomness
di: Pecher, Branislav, et al.
Pubblicazione: (2023)
di: Pecher, Branislav, et al.
Pubblicazione: (2023)
On Sensitivity of Learning with Limited Labelled Data to the Effects of Randomness: Impact of Interactions and Systematic Choices
di: Pecher, Branislav, et al.
Pubblicazione: (2024)
di: Pecher, Branislav, et al.
Pubblicazione: (2024)
LLMs Meet Finance: Fine-Tuning Foundation Models for the Open FinLLM Leaderboard
di: Rao, Varun, et al.
Pubblicazione: (2025)
di: Rao, Varun, et al.
Pubblicazione: (2025)
Reinforcement Learning with Conditional Expectation Reward
di: Xiao, Changyi, et al.
Pubblicazione: (2026)
di: Xiao, Changyi, et al.
Pubblicazione: (2026)
Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't
di: Dang, Quy-Anh, et al.
Pubblicazione: (2025)
di: Dang, Quy-Anh, et al.
Pubblicazione: (2025)
MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
di: Taghanaki, Saeid Asgari, et al.
Pubblicazione: (2024)
di: Taghanaki, Saeid Asgari, et al.
Pubblicazione: (2024)
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
di: Kim, Jeonghye, et al.
Pubblicazione: (2026)
di: Kim, Jeonghye, et al.
Pubblicazione: (2026)
Active Learning for Direct Preference Optimization
di: Kveton, Branislav, et al.
Pubblicazione: (2025)
di: Kveton, Branislav, et al.
Pubblicazione: (2025)
Efficient Continual Pre-training of LLMs for Low-resource Languages
di: Nag, Arijit, et al.
Pubblicazione: (2024)
di: Nag, Arijit, et al.
Pubblicazione: (2024)
Representation Learning with Conditional Information Flow Maximization
di: Hu, Dou, et al.
Pubblicazione: (2024)
di: Hu, Dou, et al.
Pubblicazione: (2024)
Reasoning Boosts Opinion Alignment in LLMs
di: Berdoz, Frédéric, et al.
Pubblicazione: (2026)
di: Berdoz, Frédéric, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2025) -
Experimental Design for Active Transductive Inference in Large Language Models
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024) -
AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models
di: Kveton, Branislav, et al.
Pubblicazione: (2026) -
Efficient and Interpretable Bandit Algorithms
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2023) -
Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)