PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play
Fuente:
arXiv
Saved in:
| Main Authors: | Castanyer, Roger Creus, Bradway, Geoffrey, Wolf, Lorenz, Lin, Maxwill, Mavor-Parker, Augustine N., Sargent, Matthew James |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
unix-ctf: Procedural Environments for Unix-Competence Reinforcement Learning
by: Bradway, Geoffrey, et al.
Published: (2026)
by: Bradway, Geoffrey, et al.
Published: (2026)
xgenius: LLM-Oriented Autonomous Research Platform for SLURM Clusters
by: Creus Castanyer, Roger
Published: (2026)
by: Creus Castanyer, Roger
Published: (2026)
Improving Intrinsic Exploration by Creating Stationary Objectives
by: Castanyer, Roger Creus, et al.
Published: (2023)
by: Castanyer, Roger Creus, et al.
Published: (2023)
Agentick: A Unified Benchmark for General Sequential Decision-Making Agents
by: Castanyer, Roger Creus, et al.
Published: (2026)
by: Castanyer, Roger Creus, et al.
Published: (2026)
Frequency and Generalisation of Periodic Activation Functions in Reinforcement Learning
by: Mavor-Parker, Augustine N., et al.
Published: (2024)
by: Mavor-Parker, Augustine N., et al.
Published: (2024)
Surprise-Adaptive Intrinsic Motivation for Unsupervised Reinforcement Learning
by: Hugessen, Adriana, et al.
Published: (2024)
by: Hugessen, Adriana, et al.
Published: (2024)
Using Forwards-Backwards Models to Approximate MDP Homomorphisms
by: Mavor-Parker, Augustine N., et al.
Published: (2022)
by: Mavor-Parker, Augustine N., et al.
Published: (2022)
Seirênes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement Learning
by: Castanyer, Roger Creus, et al.
Published: (2025)
by: Castanyer, Roger Creus, et al.
Published: (2025)
RLeXplore: Accelerating Research in Intrinsically-Motivated Reinforcement Learning
by: Yuan, Mingqi, et al.
Published: (2024)
by: Yuan, Mingqi, et al.
Published: (2024)
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
by: Chen, Jiaqi, et al.
Published: (2025)
by: Chen, Jiaqi, et al.
Published: (2025)
Align and Filter: Improving Performance in Asynchronous On-Policy RL
by: Honari, Homayoun, et al.
Published: (2026)
by: Honari, Homayoun, et al.
Published: (2026)
BuilderBench: The Building Blocks of Intelligent Agents
by: Ghugare, Raj, et al.
Published: (2025)
by: Ghugare, Raj, et al.
Published: (2025)
ID-LoRA: Identity-Driven Audio-Video Personalization with In-Context LoRA
by: Dahan, Aviad, et al.
Published: (2026)
by: Dahan, Aviad, et al.
Published: (2026)
How to Stay Curious while Avoiding Noisy TVs using Aleatoric Uncertainty Estimation
by: Mavor-Parker, Augustine N., et al.
Published: (2021)
by: Mavor-Parker, Augustine N., et al.
Published: (2021)
Self-Evolving Curriculum for LLM Reasoning
by: Chen, Xiaoyin, et al.
Published: (2025)
by: Chen, Xiaoyin, et al.
Published: (2025)
LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystem
by: Liu, Hongyi, et al.
Published: (2024)
by: Liu, Hongyi, et al.
Published: (2024)
The Role of the Public Libraries in Adult Independent Learning. Part II. Final Report.
by: Mavor, Anne S., et al.
Published: (1976)
by: Mavor, Anne S., et al.
Published: (1976)
The Role of the Public Libraries in Adult Independent Learning. Final Report.
by: Mavor, Anne S., et al.
Published: (1976)
by: Mavor, Anne S., et al.
Published: (1976)
An Overview of the National Adult Independent Learning Project
by: Mavor, Anne S., et al.
Published: (1976)
by: Mavor, Anne S., et al.
Published: (1976)
Non-Traditional Study, The Public Library Approach: Research Study Group for the Development of an Integrated Common Data System. Final Report.
by: Mavor, Anne S., et al.
Published: (1976)
by: Mavor, Anne S., et al.
Published: (1976)
LoRA-Drop: Temporal LoRA Decoding for Efficient LLM Inference
by: Rajabzadeh, Hossein, et al.
Published: (2026)
by: Rajabzadeh, Hossein, et al.
Published: (2026)
SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks
by: Kwan, Wai-Chung, et al.
Published: (2026)
by: Kwan, Wai-Chung, et al.
Published: (2026)
Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning
by: Castanyer, Roger Creus, et al.
Published: (2025)
by: Castanyer, Roger Creus, et al.
Published: (2025)
Mask the Target: A Plug-and-Play Regularizer Against LoRA Forgetting
by: Xu, Runze, et al.
Published: (2026)
by: Xu, Runze, et al.
Published: (2026)
Quantitat d'aigua als embassaments de les Conques Internes de Catalunya
by: Creus, Esteve
Published: (2026)
by: Creus, Esteve
Published: (2026)
WHEN HARRY MET ZUCKERMAN: SELF-REFLEXIVITY AND METAFICTION IN PHILIP ROTH AND WOODY ALLEN
by: Tomás Creus
Published: (2006)
by: Tomás Creus
Published: (2006)
eXistenZ, de David Cronenberg: ciberficciones para la posthumanidad
by: Laura Borràs Castanyer
Published: (2003)
by: Laura Borràs Castanyer
Published: (2003)
El mercado laboral indio
by: Prisca Castanyer Bonnin
Published: (2006)
by: Prisca Castanyer Bonnin
Published: (2006)
MaLoRA: Gated Modality LoRA for Key-Space Alignment in Multimodal LLM Fine-Tuning
by: Zheng, Xinhan, et al.
Published: (2025)
by: Zheng, Xinhan, et al.
Published: (2025)
Tina: Tiny Reasoning Models via LoRA
by: Wang, Shangshang, et al.
Published: (2025)
by: Wang, Shangshang, et al.
Published: (2025)
Dual-Phase LLM Reasoning: Self-Evolved Mathematical Frameworks
by: Liu, ShaoZhen, et al.
Published: (2026)
by: Liu, ShaoZhen, et al.
Published: (2026)
Population-Evolve: a Parallel Sampling and Evolutionary Method for LLM Math Reasoning
by: Zhang, Yanzhi, et al.
Published: (2025)
by: Zhang, Yanzhi, et al.
Published: (2025)
Neeko: Leveraging Dynamic LoRA for Efficient Multi-Character Role-Playing Agent
by: Yu, Xiaoyan, et al.
Published: (2024)
by: Yu, Xiaoyan, et al.
Published: (2024)
PALoRA: Projection-Adaptive LoRA for Preserving Reasoning in Large Language Models
by: Bilgin, Mustafa Hayri, et al.
Published: (2026)
by: Bilgin, Mustafa Hayri, et al.
Published: (2026)
Universe Routing: Why Self-Evolving Agents Need Epistemic Control
by: Wang, Zhaohui Geoffrey
Published: (2026)
by: Wang, Zhaohui Geoffrey
Published: (2026)
LoRA-PAR: A Flexible Dual-System LoRA Partitioning Approach to Efficient LLM Fine-Tuning
by: Huang, Yining, et al.
Published: (2025)
by: Huang, Yining, et al.
Published: (2025)
Merging LoRAs like Playing LEGO: Pushing the Modularity of LoRA to Extremes Through Rank-Wise Clustering
by: Zhao, Ziyu, et al.
Published: (2024)
by: Zhao, Ziyu, et al.
Published: (2024)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
by: Sheng, Ying, et al.
Published: (2023)
by: Sheng, Ying, et al.
Published: (2023)
HiCoLoRA: Addressing Context-Prompt Misalignment via Hierarchical Collaborative LoRA for Zero-Shot DST
by: Zhang, Shuyu, et al.
Published: (2025)
by: Zhang, Shuyu, et al.
Published: (2025)
Similar Items
-
unix-ctf: Procedural Environments for Unix-Competence Reinforcement Learning
by: Bradway, Geoffrey, et al.
Published: (2026) -
xgenius: LLM-Oriented Autonomous Research Platform for SLURM Clusters
by: Creus Castanyer, Roger
Published: (2026) -
Improving Intrinsic Exploration by Creating Stationary Objectives
by: Castanyer, Roger Creus, et al.
Published: (2023) -
Agentick: A Unified Benchmark for General Sequential Decision-Making Agents
by: Castanyer, Roger Creus, et al.
Published: (2026) -
Frequency and Generalisation of Periodic Activation Functions in Reinforcement Learning
by: Mavor-Parker, Augustine N., et al.
Published: (2024)