Investigating Regularization of Self-Play Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Alami, Reda, Abubaker, Abdalgader, Achab, Mastane, Seddik, Mohamed El Amine, Lahlou, Salem |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets
di: Younsi, Adam, et al.
Pubblicazione: (2025)
di: Younsi, Adam, et al.
Pubblicazione: (2025)
A Bregman firmly nonexpansive proximal operator for baryconvex optimization
di: Achab, Mastane
Pubblicazione: (2024)
di: Achab, Mastane
Pubblicazione: (2024)
PORT: Preference Optimization on Reasoning Traces
di: Lahlou, Salem, et al.
Pubblicazione: (2024)
di: Lahlou, Salem, et al.
Pubblicazione: (2024)
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
di: Mansouri, Omar El, et al.
Pubblicazione: (2025)
di: Mansouri, Omar El, et al.
Pubblicazione: (2025)
Alignment with Preference Optimization Is All You Need for LLM Safety
di: Alami, Reda, et al.
Pubblicazione: (2024)
di: Alami, Reda, et al.
Pubblicazione: (2024)
How Does Attention Help? Insights from Random Matrices on Signal Recovery from Sequence Models
di: Seddik, Mohamed El Amine
Pubblicazione: (2026)
di: Seddik, Mohamed El Amine
Pubblicazione: (2026)
Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory
di: Firdoussi, Aymane El, et al.
Pubblicazione: (2024)
di: Firdoussi, Aymane El, et al.
Pubblicazione: (2024)
High-dimensional Learning with Noisy Labels
di: Firdoussi, Aymane El, et al.
Pubblicazione: (2024)
di: Firdoussi, Aymane El, et al.
Pubblicazione: (2024)
Semantic Communication meets System 2 ML: How Abstraction, Compositionality and Emergent Languages Shape Intelligence
di: Bennis, Mehdi, et al.
Pubblicazione: (2025)
di: Bennis, Mehdi, et al.
Pubblicazione: (2025)
$α$-LoRA: Effective Fine-Tuning via Base Model Rescaling
di: Firdoussi, Aymane El, et al.
Pubblicazione: (2025)
di: Firdoussi, Aymane El, et al.
Pubblicazione: (2025)
Communication Enables Cooperation in LLM Agents: A Comparison with Curriculum-Based Approaches
di: Madmoun, Hachem, et al.
Pubblicazione: (2025)
di: Madmoun, Hachem, et al.
Pubblicazione: (2025)
Falcon-H1R: Pushing the Reasoning Frontiers with a Hybrid Model for Efficient Test-Time Scaling
di: Falcon LLM Team, et al.
Pubblicazione: (2026)
di: Falcon LLM Team, et al.
Pubblicazione: (2026)
Performance Gaps in Multi-view Clustering under the Nested Matrix-Tensor Model
di: Lebeau, Hugo, et al.
Pubblicazione: (2024)
di: Lebeau, Hugo, et al.
Pubblicazione: (2024)
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
di: Seddik, Mohamed El Amine, et al.
Pubblicazione: (2024)
di: Seddik, Mohamed El Amine, et al.
Pubblicazione: (2024)
Likelihood-guided Regularization in Attention Based Models
di: Salem, Mohamed, et al.
Pubblicazione: (2025)
di: Salem, Mohamed, et al.
Pubblicazione: (2025)
High-Dimensional Analysis of Bootstrap Ensemble Classifiers
di: Tiomoko, Malik, et al.
Pubblicazione: (2025)
di: Tiomoko, Malik, et al.
Pubblicazione: (2025)
Improved Exploration in GFlownets via Enhanced Epistemic Neural Networks
di: Muhammad, Sajan, et al.
Pubblicazione: (2025)
di: Muhammad, Sajan, et al.
Pubblicazione: (2025)
RSPO: Regularized Self-Play Alignment of Large Language Models
di: Tang, Xiaohang, et al.
Pubblicazione: (2025)
di: Tang, Xiaohang, et al.
Pubblicazione: (2025)
On the Stability of the Jacobian Matrix in Deep Neural Networks
di: Dadoun, Benjamin, et al.
Pubblicazione: (2025)
di: Dadoun, Benjamin, et al.
Pubblicazione: (2025)
On Generalization for Generative Flow Networks
di: Krichel, Anas, et al.
Pubblicazione: (2024)
di: Krichel, Anas, et al.
Pubblicazione: (2024)
On the Privacy Risks of Spiking Neural Networks: A Membership Inference Analysis
di: Guan, Junyi, et al.
Pubblicazione: (2025)
di: Guan, Junyi, et al.
Pubblicazione: (2025)
HyperKKL: Enabling Non-Autonomous State Estimation through Dynamic Weight Conditioning
di: Shaaban, Yahia Salaheldin, et al.
Pubblicazione: (2026)
di: Shaaban, Yahia Salaheldin, et al.
Pubblicazione: (2026)
Valid Feature-Level Inference for Tabular Foundation Models via the Conditional Randomization Test
di: Salem, Mohamed
Pubblicazione: (2026)
di: Salem, Mohamed
Pubblicazione: (2026)
Loss-Guided Auxiliary Agents for Overcoming Mode Collapse in GFlowNets
di: Malek, Idriss, et al.
Pubblicazione: (2025)
di: Malek, Idriss, et al.
Pubblicazione: (2025)
Comparative Performance Analysis of Quantum Machine Learning Architectures for Credit Card Fraud Detection
di: Alami, Mansour El, et al.
Pubblicazione: (2024)
di: Alami, Mansour El, et al.
Pubblicazione: (2024)
Curriculum-Augmented GFlowNets For mRNA Sequence Generation
di: Laajil, Aya, et al.
Pubblicazione: (2025)
di: Laajil, Aya, et al.
Pubblicazione: (2025)
Avoid What You Know: Divergent Trajectory Balance for GFlowNets
di: Dall'Antonia, Pedro, et al.
Pubblicazione: (2026)
di: Dall'Antonia, Pedro, et al.
Pubblicazione: (2026)
Do Vision and Language Encoders Represent the World Similarly?
di: Maniparambil, Mayug, et al.
Pubblicazione: (2024)
di: Maniparambil, Mayug, et al.
Pubblicazione: (2024)
PoliTune: Analyzing the Impact of Data Selection and Fine-Tuning on Economic and Political Biases in Large Language Models
di: Agiza, Ahmed, et al.
Pubblicazione: (2024)
di: Agiza, Ahmed, et al.
Pubblicazione: (2024)
PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models
di: Seddik, Issam, et al.
Pubblicazione: (2025)
di: Seddik, Issam, et al.
Pubblicazione: (2025)
Zero-Shot Off-Policy Learning
di: Asadulaev, Arip, et al.
Pubblicazione: (2026)
di: Asadulaev, Arip, et al.
Pubblicazione: (2026)
torchgfn: A PyTorch GFlowNet library
di: Viviano, Joseph D., et al.
Pubblicazione: (2023)
di: Viviano, Joseph D., et al.
Pubblicazione: (2023)
Reinforcement Learning-enabled Satellite Constellation Reconfiguration and Retasking for Mission-Critical Applications
di: Alami, Hassan El, et al.
Pubblicazione: (2024)
di: Alami, Hassan El, et al.
Pubblicazione: (2024)
GFlowNet Foundations
di: Bengio, Yoshua, et al.
Pubblicazione: (2021)
di: Bengio, Yoshua, et al.
Pubblicazione: (2021)
Diffusion-based Adversarial Purification for Intrusion Detection
di: Merzouk, Mohamed Amine, et al.
Pubblicazione: (2024)
di: Merzouk, Mohamed Amine, et al.
Pubblicazione: (2024)
SVRPBench: A Realistic Benchmark for Stochastic Vehicle Routing Problem
di: Heakl, Ahmed, et al.
Pubblicazione: (2025)
di: Heakl, Ahmed, et al.
Pubblicazione: (2025)
SCAFFLSA: Taming Heterogeneity in Federated Linear Stochastic Approximation and TD Learning
di: Mangold, Paul, et al.
Pubblicazione: (2024)
di: Mangold, Paul, et al.
Pubblicazione: (2024)
Self-Play Preference Optimization for Language Model Alignment
di: Wu, Yue, et al.
Pubblicazione: (2024)
di: Wu, Yue, et al.
Pubblicazione: (2024)
VARS-FL: Validation-Aligned Client Selection for Non-IID Federated Learning in IoT Systems
di: Lakas, Mohamed, et al.
Pubblicazione: (2026)
di: Lakas, Mohamed, et al.
Pubblicazione: (2026)
BackPlay: Head-Only Look-Back Self-Correction for Diffusion Language Models
di: Liu, Liming, et al.
Pubblicazione: (2026)
di: Liu, Liming, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets
di: Younsi, Adam, et al.
Pubblicazione: (2025) -
A Bregman firmly nonexpansive proximal operator for baryconvex optimization
di: Achab, Mastane
Pubblicazione: (2024) -
PORT: Preference Optimization on Reasoning Traces
di: Lahlou, Salem, et al.
Pubblicazione: (2024) -
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
di: Mansouri, Omar El, et al.
Pubblicazione: (2025) -
Alignment with Preference Optimization Is All You Need for LLM Safety
di: Alami, Reda, et al.
Pubblicazione: (2024)