Mixing Expert Knowledge: Bring Human Thoughts Back To the Game of Go
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Yichuan, Li, Linyang, Chen, Yongkang, Li, Peiji, Ye, Jiasheng, Guo, Qipeng, Lin, Dahua, Chen, Kai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic
von: Ma, Yichuan, et al.
Veröffentlicht: (2026)
von: Ma, Yichuan, et al.
Veröffentlicht: (2026)
InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
von: Li, Peiji, et al.
Veröffentlicht: (2025)
von: Li, Peiji, et al.
Veröffentlicht: (2025)
UnitCoder: Scalable Iterative Code Synthesis with Unit Test Guidance
von: Ma, Yichuan, et al.
Veröffentlicht: (2025)
von: Ma, Yichuan, et al.
Veröffentlicht: (2025)
FastMCTS: A Simple Sampling Strategy for Data Synthesis
von: Li, Peiji, et al.
Veröffentlicht: (2025)
von: Li, Peiji, et al.
Veröffentlicht: (2025)
TL-GRPO: Turn-Level RL for Reasoning-Guided Iterative Optimization
von: Li, Peiji, et al.
Veröffentlicht: (2026)
von: Li, Peiji, et al.
Veröffentlicht: (2026)
Case2Code: Scalable Synthetic Data for Code Generation
von: Shao, Yunfan, et al.
Veröffentlicht: (2024)
von: Shao, Yunfan, et al.
Veröffentlicht: (2024)
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents
von: Li, Xiaozhe, et al.
Veröffentlicht: (2026)
von: Li, Xiaozhe, et al.
Veröffentlicht: (2026)
Unearthing Large Scale Domain-Specific Knowledge from Public Corpora
von: Fei, Zhaoye, et al.
Veröffentlicht: (2024)
von: Fei, Zhaoye, et al.
Veröffentlicht: (2024)
F-Eval: Assessing Fundamental Abilities with Refined Evaluation Methods
von: Sun, Yu, et al.
Veröffentlicht: (2024)
von: Sun, Yu, et al.
Veröffentlicht: (2024)
Beyond Mode Collapse: Distribution Matching for Diverse Reasoning
von: Li, Xiaozhe, et al.
Veröffentlicht: (2026)
von: Li, Xiaozhe, et al.
Veröffentlicht: (2026)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
von: Wang, Bo, et al.
Veröffentlicht: (2025)
von: Wang, Bo, et al.
Veröffentlicht: (2025)
LongWanjuan: Towards Systematic Measurement for Long Text Quality
von: Lv, Kai, et al.
Veröffentlicht: (2024)
von: Lv, Kai, et al.
Veröffentlicht: (2024)
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2024)
Balanced Data Sampling for Language Model Training with Clustering
von: Shao, Yunfan, et al.
Veröffentlicht: (2024)
von: Shao, Yunfan, et al.
Veröffentlicht: (2024)
Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back Home
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025)
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025)
What are the Essential Factors in Crafting Effective Long Context Multi-Hop Instruction Datasets? Insights and Best Practices
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
Mix of Experts Language Model for Named Entity Recognition
von: Chen, Xinwei, et al.
Veröffentlicht: (2024)
von: Chen, Xinwei, et al.
Veröffentlicht: (2024)
HERGC: Heterogeneous Experts Representation and Generative Completion for Multimodal Knowledge Graphs
von: Xiao, Yongkang, et al.
Veröffentlicht: (2025)
von: Xiao, Yongkang, et al.
Veröffentlicht: (2025)
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
von: Hua, Zhouqi, et al.
Veröffentlicht: (2025)
von: Hua, Zhouqi, et al.
Veröffentlicht: (2025)
Data-free Weight Compress and Denoise for Large Language Models
von: Peng, Runyu, et al.
Veröffentlicht: (2024)
von: Peng, Runyu, et al.
Veröffentlicht: (2024)
Scaling Behavior for Large Language Models regarding Numeral Systems: An Example using Pythia
von: Zhou, Zhejian, et al.
Veröffentlicht: (2024)
von: Zhou, Zhejian, et al.
Veröffentlicht: (2024)
AlchemistCoder: Harmonizing and Eliciting Code Capability by Hindsight Tuning on Multi-source Data
von: Song, Zifan, et al.
Veröffentlicht: (2024)
von: Song, Zifan, et al.
Veröffentlicht: (2024)
Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment
von: Liu, Zhili, et al.
Veröffentlicht: (2024)
von: Liu, Zhili, et al.
Veröffentlicht: (2024)
DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting
von: Lv, Kai, et al.
Veröffentlicht: (2025)
von: Lv, Kai, et al.
Veröffentlicht: (2025)
Identifying Semantic Induction Heads to Understand In-Context Learning
von: Ren, Jie, et al.
Veröffentlicht: (2024)
von: Ren, Jie, et al.
Veröffentlicht: (2024)
OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification
von: Wu, Zijian, et al.
Veröffentlicht: (2025)
von: Wu, Zijian, et al.
Veröffentlicht: (2025)
LEAN-GitHub: Compiling GitHub LEAN repositories for a versatile LEAN prover
von: Wu, Zijian, et al.
Veröffentlicht: (2024)
von: Wu, Zijian, et al.
Veröffentlicht: (2024)
CritiQ: Mining Data Quality Criteria from Human Preferences
von: Guo, Honglin, et al.
Veröffentlicht: (2025)
von: Guo, Honglin, et al.
Veröffentlicht: (2025)
FBQuant: FeedBack Quantization for Large Language Models
von: Liu, Yijiang, et al.
Veröffentlicht: (2025)
von: Liu, Yijiang, et al.
Veröffentlicht: (2025)
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
Prune, Interpret, Evaluate: A Cross-Layer Transcoder-Native Framework for Efficient Circuit Discovery via Feature Attribution
von: Chen, Qinhao, et al.
Veröffentlicht: (2026)
von: Chen, Qinhao, et al.
Veröffentlicht: (2026)
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
von: Ying, Huaiyuan, et al.
Veröffentlicht: (2024)
von: Ying, Huaiyuan, et al.
Veröffentlicht: (2024)
Cultivating Game Sense for Yourself: Making VLMs Gaming Experts
von: Lu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Lu, Wenxuan, et al.
Veröffentlicht: (2025)
DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels
von: Xu, Zhe, et al.
Veröffentlicht: (2024)
von: Xu, Zhe, et al.
Veröffentlicht: (2024)
MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems
von: Zhang, Xuanming, et al.
Veröffentlicht: (2025)
von: Zhang, Xuanming, et al.
Veröffentlicht: (2025)
RouteGoT: Node-Adaptive Routing for Cost-Efficient Graph of Thoughts Reasoning
von: Liu, Yuhang, et al.
Veröffentlicht: (2026)
von: Liu, Yuhang, et al.
Veröffentlicht: (2026)
Wrong-of-Thought: An Integrated Reasoning Framework with Multi-Perspective Verification and Wrong Information
von: Zhang, Yongheng, et al.
Veröffentlicht: (2024)
von: Zhang, Yongheng, et al.
Veröffentlicht: (2024)
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning
von: Shan, Zikang, et al.
Veröffentlicht: (2026)
von: Shan, Zikang, et al.
Veröffentlicht: (2026)
From Thinking to Output: Chain-of-Thought and Text Generation Characteristics in Reasoning Language Models
von: Liu, Junhao, et al.
Veröffentlicht: (2025)
von: Liu, Junhao, et al.
Veröffentlicht: (2025)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic
von: Ma, Yichuan, et al.
Veröffentlicht: (2026) -
InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
von: Li, Peiji, et al.
Veröffentlicht: (2025) -
UnitCoder: Scalable Iterative Code Synthesis with Unit Test Guidance
von: Ma, Yichuan, et al.
Veröffentlicht: (2025) -
FastMCTS: A Simple Sampling Strategy for Data Synthesis
von: Li, Peiji, et al.
Veröffentlicht: (2025) -
TL-GRPO: Turn-Level RL for Reasoning-Guided Iterative Optimization
von: Li, Peiji, et al.
Veröffentlicht: (2026)