Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Wu, Zhou, Yigeng, Shi, Zesheng, Wang, Yequan, Zhang, Min, Li, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mitigating Context-Memory Conflicts in LLMs through Dynamic Cognitive Reconciliation Decoding
von: Zhou, Yigeng, et al.
Veröffentlicht: (2026)
von: Zhou, Yigeng, et al.
Veröffentlicht: (2026)
Multi-objective Large Language Model Alignment with Hierarchical Experts
von: Li, Zhuo, et al.
Veröffentlicht: (2025)
von: Li, Zhuo, et al.
Veröffentlicht: (2025)
Safety Alignment via Constrained Knowledge Unlearning
von: Shi, Zesheng, et al.
Veröffentlicht: (2025)
von: Shi, Zesheng, et al.
Veröffentlicht: (2025)
Multimodal Reasoning with Multimodal Knowledge Graph
von: Lee, Junlin, et al.
Veröffentlicht: (2024)
von: Lee, Junlin, et al.
Veröffentlicht: (2024)
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
von: Wu, Xinwei, et al.
Veröffentlicht: (2025)
von: Wu, Xinwei, et al.
Veröffentlicht: (2025)
Commonsense Knowledge Editing Based on Free-Text in LLMs
von: Huang, Xiusheng, et al.
Veröffentlicht: (2024)
von: Huang, Xiusheng, et al.
Veröffentlicht: (2024)
SPFT-SQL: Enhancing Large Language Model for Text-to-SQL Parsing by Self-Play Fine-Tuning
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuning
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
60 Data Points are Sufficient to Fine-Tune LLMs for Question-Answering
von: Ye, Junjie, et al.
Veröffentlicht: (2024)
von: Ye, Junjie, et al.
Veröffentlicht: (2024)
PsyMem: Fine-grained psychological alignment and Explicit Memory Control for Advanced Role-Playing LLMs
von: Cheng, Xilong, et al.
Veröffentlicht: (2025)
von: Cheng, Xilong, et al.
Veröffentlicht: (2025)
AnyTaskTune: Advanced Domain-Specific Solutions through Task-Fine-Tuning
von: Cui, Jiaxi, et al.
Veröffentlicht: (2024)
von: Cui, Jiaxi, et al.
Veröffentlicht: (2024)
RankAdaptor: Hierarchical Rank Allocation for Efficient Fine-Tuning Pruned LLMs via Performance Model
von: Zhou, Changhai, et al.
Veröffentlicht: (2024)
von: Zhou, Changhai, et al.
Veröffentlicht: (2024)
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
von: Guo, Song, et al.
Veröffentlicht: (2024)
von: Guo, Song, et al.
Veröffentlicht: (2024)
Step-On-Feet Tuning: Scaling Self-Alignment of LLMs via Bootstrapping
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
von: Chen, Zixiang, et al.
Veröffentlicht: (2024)
von: Chen, Zixiang, et al.
Veröffentlicht: (2024)
Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy
von: Shi, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Shi, Xiaofeng, et al.
Veröffentlicht: (2025)
ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning
von: Li, Xianming, et al.
Veröffentlicht: (2026)
von: Li, Xianming, et al.
Veröffentlicht: (2026)
Playing Language Game with LLMs Leads to Jailbreaking
von: Peng, Yu, et al.
Veröffentlicht: (2024)
von: Peng, Yu, et al.
Veröffentlicht: (2024)
SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-Reflection
von: Liu, Liangxin, et al.
Veröffentlicht: (2024)
von: Liu, Liangxin, et al.
Veröffentlicht: (2024)
Natural Language Fine-Tuning
von: Liu, Jia, et al.
Veröffentlicht: (2024)
von: Liu, Jia, et al.
Veröffentlicht: (2024)
Mind the Quote: Enabling Quotation-Aware Dialogue in LLMs via Plug-and-Play Modules
von: Zhang, Yueqi, et al.
Veröffentlicht: (2025)
von: Zhang, Yueqi, et al.
Veröffentlicht: (2025)
Self-Play Fine-Tuning of Diffusion Models for Text-to-Image Generation
von: Yuan, Huizhuo, et al.
Veröffentlicht: (2024)
von: Yuan, Huizhuo, et al.
Veröffentlicht: (2024)
Data Analysis and Performance Evaluation of Simulation Deduction Based on LLMs
von: Zhang, Shansi, et al.
Veröffentlicht: (2025)
von: Zhang, Shansi, et al.
Veröffentlicht: (2025)
Using LLMs for Automated Privacy Policy Analysis: Prompt Engineering, Fine-Tuning and Explainability
von: Chen, Yuxin, et al.
Veröffentlicht: (2025)
von: Chen, Yuxin, et al.
Veröffentlicht: (2025)
Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches
von: Yousefiramandi, Amirhossein, et al.
Veröffentlicht: (2025)
von: Yousefiramandi, Amirhossein, et al.
Veröffentlicht: (2025)
SELT: Self-Evaluation Tree Search for LLMs with Task Decomposition
von: Wu, Mengsong, et al.
Veröffentlicht: (2025)
von: Wu, Mengsong, et al.
Veröffentlicht: (2025)
HydraLoRA: An Asymmetric LoRA Architecture for Efficient Fine-Tuning
von: Tian, Chunlin, et al.
Veröffentlicht: (2024)
von: Tian, Chunlin, et al.
Veröffentlicht: (2024)
Parameter-Efficient Fine-Tuning for Medical Text Summarization: A Comparative Study of Lora, Prompt Tuning, and Full Fine-Tuning
von: Shernazarov, Ulugbek, et al.
Veröffentlicht: (2026)
von: Shernazarov, Ulugbek, et al.
Veröffentlicht: (2026)
LIFT: Improving Long Context Understanding Through Long Input Fine-Tuning
von: Mao, Yansheng, et al.
Veröffentlicht: (2024)
von: Mao, Yansheng, et al.
Veröffentlicht: (2024)
Deconfounded Causality-aware Parameter-Efficient Fine-Tuning for Problem-Solving Improvement of LLMs
von: Wang, Ruoyu, et al.
Veröffentlicht: (2024)
von: Wang, Ruoyu, et al.
Veröffentlicht: (2024)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
von: Zhou, Sifan, et al.
Veröffentlicht: (2025)
von: Zhou, Sifan, et al.
Veröffentlicht: (2025)
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds
von: Wang, Lei, et al.
Veröffentlicht: (2024)
von: Wang, Lei, et al.
Veröffentlicht: (2024)
Phased Instruction Fine-Tuning for Large Language Models
von: Pang, Wei, et al.
Veröffentlicht: (2024)
von: Pang, Wei, et al.
Veröffentlicht: (2024)
Fine-Tuning LLMs for Reliable Medical Question-Answering Services
von: Anaissi, Ali, et al.
Veröffentlicht: (2024)
von: Anaissi, Ali, et al.
Veröffentlicht: (2024)
Positive and Risky Message Assessment for Music Products
von: Zhang, Yigeng, et al.
Veröffentlicht: (2023)
von: Zhang, Yigeng, et al.
Veröffentlicht: (2023)
Fine-Grained Behavior Simulation with Role-Playing Large Language Model on Social Media
von: Li, Kun, et al.
Veröffentlicht: (2024)
von: Li, Kun, et al.
Veröffentlicht: (2024)
GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation
von: Chen, Zihong, et al.
Veröffentlicht: (2025)
von: Chen, Zihong, et al.
Veröffentlicht: (2025)
LLMs for Explainable Business Decision-Making: A Reinforcement Learning Fine-Tuning Approach
von: Cheng, Xiang, et al.
Veröffentlicht: (2025)
von: Cheng, Xiang, et al.
Veröffentlicht: (2025)
Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
What Do LLMs Know About Alzheimer's Disease? Multi-loss Fine-Tuning and Probing for AD Detection
von: Jiang, Lei, et al.
Veröffentlicht: (2026)
von: Jiang, Lei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Mitigating Context-Memory Conflicts in LLMs through Dynamic Cognitive Reconciliation Decoding
von: Zhou, Yigeng, et al.
Veröffentlicht: (2026) -
Multi-objective Large Language Model Alignment with Hierarchical Experts
von: Li, Zhuo, et al.
Veröffentlicht: (2025) -
Safety Alignment via Constrained Knowledge Unlearning
von: Shi, Zesheng, et al.
Veröffentlicht: (2025) -
Multimodal Reasoning with Multimodal Knowledge Graph
von: Lee, Junlin, et al.
Veröffentlicht: (2024) -
Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity
von: Wu, Xinwei, et al.
Veröffentlicht: (2025)