R-Diverse: Mitigating Diversity Illusion in Self-Play LLM Training
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Gengsheng, He, Jinghan, Wang, Shijie, Zhang, Dan, Liu, Ruiqi, Zhang, Renrui, Yao, Zijun, Fang, Junfeng, Guo, Haiyun, Wang, Jinqiao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration
by: He, Jinghan, et al.
Published: (2026)
by: He, Jinghan, et al.
Published: (2026)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
by: Hua, Zhenglin, et al.
Published: (2025)
by: Hua, Zhenglin, et al.
Published: (2025)
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
by: Li, Gengsheng, et al.
Published: (2026)
by: Li, Gengsheng, et al.
Published: (2026)
SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language Models
by: He, Jinghan, et al.
Published: (2024)
by: He, Jinghan, et al.
Published: (2024)
Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
by: He, Jinghan, et al.
Published: (2024)
by: He, Jinghan, et al.
Published: (2024)
TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval
by: Hao, Xiangzhao, et al.
Published: (2026)
by: Hao, Xiangzhao, et al.
Published: (2026)
Rubric-based On-policy Distillation
by: Fang, Junfeng, et al.
Published: (2026)
by: Fang, Junfeng, et al.
Published: (2026)
Diversity-oriented Data Augmentation with Large Language Models
by: Wang, Zaitian, et al.
Published: (2025)
by: Wang, Zaitian, et al.
Published: (2025)
ST-Prune: Training-Free Spatio-Temporal Token Pruning for Vision-Language Models in Autonomous Driving
by: Sha, Lin, et al.
Published: (2026)
by: Sha, Lin, et al.
Published: (2026)
PASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual Learning
by: Hou, Zhiyan, et al.
Published: (2026)
by: Hou, Zhiyan, et al.
Published: (2026)
UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
by: Guo, Hongyu, et al.
Published: (2025)
by: Guo, Hongyu, et al.
Published: (2025)
UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs
by: Ni, Shuo, et al.
Published: (2026)
by: Ni, Shuo, et al.
Published: (2026)
We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
by: Fang, Junfeng, et al.
Published: (2025)
by: Fang, Junfeng, et al.
Published: (2025)
Rethinking Representativeness and Diversity in Dynamic Data Selection
by: Zhou, Yuzhe, et al.
Published: (2026)
by: Zhou, Yuzhe, et al.
Published: (2026)
Monocular Lane Detection Based on Deep Learning: A Survey
by: He, Xin, et al.
Published: (2024)
by: He, Xin, et al.
Published: (2024)
Diversity of Play
by: Fuchs, Mathias
Published: (2020)
by: Fuchs, Mathias
Published: (2020)
WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrieval
by: Wang, Tianyue, et al.
Published: (2026)
by: Wang, Tianyue, et al.
Published: (2026)
MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis
by: Guo, Haiyun, et al.
Published: (2025)
by: Guo, Haiyun, et al.
Published: (2025)
AAformer: Auto-Aligned Transformer for Person Re-Identification
by: Zhu, Kuan, et al.
Published: (2021)
by: Zhu, Kuan, et al.
Published: (2021)
The Species Fixed Law III: The Diversity Illusion
by: ANG, FOO SENG
Published: (2026)
by: ANG, FOO SENG
Published: (2026)
Revealing and Mitigating the Challenge of Detecting Character Knowledge Errors in LLM Role-Playing
by: Zhang, Wenyuan, et al.
Published: (2024)
by: Zhang, Wenyuan, et al.
Published: (2024)
FOCUS: Fine-grained Optimization with Semantic Guided Understanding for Pedestrian Attributes Recognition
by: An, Hongyan, et al.
Published: (2025)
by: An, Hongyan, et al.
Published: (2025)
Mitigating the Negative Impact of Over-association for Conversational Query Production
by: Wang, Ante, et al.
Published: (2024)
by: Wang, Ante, et al.
Published: (2024)
Enhancing Multi-Modal LLMs Reasoning via Difficulty-Aware Group Normalization
by: Li, Jinghan, et al.
Published: (2026)
by: Li, Jinghan, et al.
Published: (2026)
DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO
by: Liu, Henglin, et al.
Published: (2025)
by: Liu, Henglin, et al.
Published: (2025)
DCAST: Diverse Class-Aware Self-Training Mitigates Selection Bias for Fairer Learning
by: Tepeli, Yasin I., et al.
Published: (2024)
by: Tepeli, Yasin I., et al.
Published: (2024)
Internalizing World Models via Self-Play Finetuning for Agentic RL
by: Chen, Shiqi, et al.
Published: (2025)
by: Chen, Shiqi, et al.
Published: (2025)
Flow decomposition for heat equations with memory
by: Wang, Gengsheng, et al.
Published: (2021)
by: Wang, Gengsheng, et al.
Published: (2021)
Sampling Observability for Heat Equations with Memory
by: Ma, Lingying, et al.
Published: (2024)
by: Ma, Lingying, et al.
Published: (2024)
Periodic propagation of singularities for heat equations with time delay
by: Wang, Gengsheng, et al.
Published: (2025)
by: Wang, Gengsheng, et al.
Published: (2025)
Observability for heat equations with time-dependent analytic memory
by: Wang, Gengsheng, et al.
Published: (2021)
by: Wang, Gengsheng, et al.
Published: (2021)
Learning Diverse Policies with Soft Self-Generated Guidance
by: Wang, Guojian, et al.
Published: (2024)
by: Wang, Guojian, et al.
Published: (2024)
Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
by: Zhang, Jiayi, et al.
Published: (2025)
by: Zhang, Jiayi, et al.
Published: (2025)
On the Effect of Sampling Diversity in Scaling LLM Inference
by: Wang, Tianchun, et al.
Published: (2025)
by: Wang, Tianchun, et al.
Published: (2025)
Can Large Language Models Play Games? A Case Study of A Self-Play Approach
by: Guo, Hongyi, et al.
Published: (2024)
by: Guo, Hongyi, et al.
Published: (2024)
Breaking the Illusion: Consensus-Based Generative Mitigation of Adversarial Illusions in Multi-Modal Embeddings
by: Akbarian, Fatemeh, et al.
Published: (2025)
by: Akbarian, Fatemeh, et al.
Published: (2025)
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
by: Fan, Chongyu, et al.
Published: (2024)
by: Fan, Chongyu, et al.
Published: (2024)
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
by: Fang, Junfeng, et al.
Published: (2025)
by: Fang, Junfeng, et al.
Published: (2025)
Efficient, Adaptive Near-Field Beam Training based on Linear Bandit
by: Liu, Junchi, et al.
Published: (2026)
by: Liu, Junchi, et al.
Published: (2026)
Molecular Diversity of Gustatory Receptors in Insects: From Structure to Ecological Adaptation
by: Rongrong Xu, et al.
Published: (2026)
by: Rongrong Xu, et al.
Published: (2026)
Similar Items
-
Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration
by: He, Jinghan, et al.
Published: (2026) -
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
by: Hua, Zhenglin, et al.
Published: (2025) -
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
by: Li, Gengsheng, et al.
Published: (2026) -
SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language Models
by: He, Jinghan, et al.
Published: (2024) -
Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
by: He, Jinghan, et al.
Published: (2024)