Enregistré dans:
| Auteurs principaux: | Qi, Xianbiao, Chen, Marco, Ye, Jiaquan, He, Yelin, Xiao, Rong |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2602.04669 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SimpleGPT: Improving GPT via A Simple Normalization Strategy
par: Chen, Marco, et autres
Publié: (2026)
par: Chen, Marco, et autres
Publié: (2026)
DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD
par: Qi, Xianbiao, et autres
Publié: (2025)
par: Qi, Xianbiao, et autres
Publié: (2025)
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
par: Yang, Dongquan, et autres
Publié: (2025)
par: Yang, Dongquan, et autres
Publié: (2025)
Muon is Scalable for LLM Training
par: Liu, Jingyuan, et autres
Publié: (2025)
par: Liu, Jingyuan, et autres
Publié: (2025)
Taming Transformer Without Using Learning Rate Warmup
par: Qi, Xianbiao, et autres
Publié: (2025)
par: Qi, Xianbiao, et autres
Publié: (2025)
Mousse: Rectifying the Geometry of Muon with Curvature-Aware Preconditioning
par: Zhang, Yechen, et autres
Publié: (2026)
par: Zhang, Yechen, et autres
Publié: (2026)
Why Does ChatGPT "Delve" So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models
par: Juzek, Tom S., et autres
Publié: (2024)
par: Juzek, Tom S., et autres
Publié: (2024)
DeepCritic: Deliberate Critique with Large Language Models
par: Yang, Wenkai, et autres
Publié: (2025)
par: Yang, Wenkai, et autres
Publié: (2025)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
par: Tong, Chengzhuo, et autres
Publié: (2025)
par: Tong, Chengzhuo, et autres
Publié: (2025)
Real Deep Research for AI, Robotics and Beyond
par: Zou, Xueyan, et autres
Publié: (2025)
par: Zou, Xueyan, et autres
Publié: (2025)
An Analysis and Mitigation of the Reversal Curse
par: Lv, Ang, et autres
Publié: (2023)
par: Lv, Ang, et autres
Publié: (2023)
Frac-Connections: Fractional Extension of Hyper-Connections
par: Zhu, Defa, et autres
Publié: (2025)
par: Zhu, Defa, et autres
Publié: (2025)
Neutral Residues: Revisiting Adapters for Model Extension
par: Talla, Franck Signe, et autres
Publié: (2024)
par: Talla, Franck Signe, et autres
Publié: (2024)
Beyond Isolation: Multi-Agent Synergy for Improving Knowledge Graph Construction
par: Ye, Hongbin, et autres
Publié: (2023)
par: Ye, Hongbin, et autres
Publié: (2023)
Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond
par: Chen, Rubing, et autres
Publié: (2025)
par: Chen, Rubing, et autres
Publié: (2025)
DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
par: Zheng, Yuxiang, et autres
Publié: (2025)
par: Zheng, Yuxiang, et autres
Publié: (2025)
RHYTHM: Reasoning with Hierarchical Temporal Tokenization for Human Mobility
par: He, Haoyu, et autres
Publié: (2025)
par: He, Haoyu, et autres
Publié: (2025)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
par: Gunjal, Anisha, et autres
Publié: (2025)
par: Gunjal, Anisha, et autres
Publié: (2025)
Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
par: Ye, Ziyi, et autres
Publié: (2024)
par: Ye, Ziyi, et autres
Publié: (2024)
Training-Free Exponential Context Extension via Cascading KV Cache
par: Willette, Jeffrey, et autres
Publié: (2024)
par: Willette, Jeffrey, et autres
Publié: (2024)
YaRN: Efficient Context Window Extension of Large Language Models
par: Peng, Bowen, et autres
Publié: (2023)
par: Peng, Bowen, et autres
Publié: (2023)
CogAtom: From Cognitive Atoms to Olympiad-level Mathematical Reasoning in Large Language Models
par: Chen, Zhuofan, et autres
Publié: (2025)
par: Chen, Zhuofan, et autres
Publié: (2025)
A Theoretical Understanding of Chain-of-Thought: Coherent Reasoning and Error-Aware Demonstration
par: Cui, Yingqian, et autres
Publié: (2024)
par: Cui, Yingqian, et autres
Publié: (2024)
CORI: CJKV Benchmark with Romanization Integration -- A step towards Cross-lingual Transfer Beyond Textual Scripts
par: Nguyen, Hoang H., et autres
Publié: (2024)
par: Nguyen, Hoang H., et autres
Publié: (2024)
Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation
par: Zhang, Wenbo, et autres
Publié: (2025)
par: Zhang, Wenbo, et autres
Publié: (2025)
Beyond Correlation: Refutation-Validated Aspect-Based Sentiment Analysis for Explainable Energy Market Returns
par: van der Heever, Wihan, et autres
Publié: (2026)
par: van der Heever, Wihan, et autres
Publié: (2026)
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
par: Shin, Haebin, et autres
Publié: (2025)
par: Shin, Haebin, et autres
Publié: (2025)
SpanNorm: Reconciling Training Stability and Performance in Deep Transformers
par: Wang, Chao, et autres
Publié: (2026)
par: Wang, Chao, et autres
Publié: (2026)
Beyond Introspection: Reinforcing Thinking via Externalist Behavioral Feedback
par: Yang, Diji, et autres
Publié: (2024)
par: Yang, Diji, et autres
Publié: (2024)
Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed
par: Fu, Yonggan, et autres
Publié: (2025)
par: Fu, Yonggan, et autres
Publié: (2025)
Generation Meets Verification: Accelerating Large Language Model Inference with Smart Parallel Auto-Correct Decoding
par: Yi, Hanling, et autres
Publié: (2024)
par: Yi, Hanling, et autres
Publié: (2024)
NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning
par: Liu, Wei, et autres
Publié: (2025)
par: Liu, Wei, et autres
Publié: (2025)
TTSR: Test-Time Self-Reflection for Continual Reasoning Improvement
par: He, Haoyang, et autres
Publié: (2026)
par: He, Haoyang, et autres
Publié: (2026)
Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs
par: Bhattacharyya, Sree, et autres
Publié: (2026)
par: Bhattacharyya, Sree, et autres
Publié: (2026)
Enhancing Quantitative Reasoning Skills of Large Language Models through Dimension Perception
par: Huang, Yuncheng, et autres
Publié: (2023)
par: Huang, Yuncheng, et autres
Publié: (2023)
Deep Dense Exploration for LLM Reinforcement Learning via Pivot-Driven Resampling
par: Guo, Yiran, et autres
Publié: (2026)
par: Guo, Yiran, et autres
Publié: (2026)
ChemAmp: Amplified Chemistry Tools via Composable Agents
par: Li, Zhucong, et autres
Publié: (2025)
par: Li, Zhucong, et autres
Publié: (2025)
SLOT: Sample-specific Language Model Optimization at Test-time
par: Hu, Yang, et autres
Publié: (2025)
par: Hu, Yang, et autres
Publié: (2025)
GoRA: Gradient-driven Adaptive Low Rank Adaptation
par: He, Haonan, et autres
Publié: (2025)
par: He, Haonan, et autres
Publié: (2025)
DeepOnto: A Python Package for Ontology Engineering with Deep Learning
par: He, Yuan, et autres
Publié: (2023)
par: He, Yuan, et autres
Publié: (2023)
Documents similaires
-
SimpleGPT: Improving GPT via A Simple Normalization Strategy
par: Chen, Marco, et autres
Publié: (2026) -
DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD
par: Qi, Xianbiao, et autres
Publié: (2025) -
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs
par: Yang, Dongquan, et autres
Publié: (2025) -
Muon is Scalable for LLM Training
par: Liu, Jingyuan, et autres
Publié: (2025) -
Taming Transformer Without Using Learning Rate Warmup
par: Qi, Xianbiao, et autres
Publié: (2025)