KunlunBaize: LLM with Multi-Scale Convolution and Multi-Token Prediction Under TransformerX Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Cheng, Liu, Jiexiong, Chen, Yixuan, Jia, Yanqin, Li, Zhepeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models
von: Li, Cheng, et al.
Veröffentlicht: (2025)
von: Li, Cheng, et al.
Veröffentlicht: (2025)
Video-VoT-R1: An efficient video inference model integrating image packing and AoE architecture
von: Li, Cheng, et al.
Veröffentlicht: (2025)
von: Li, Cheng, et al.
Veröffentlicht: (2025)
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
von: Li, Cheng, et al.
Veröffentlicht: (2025)
von: Li, Cheng, et al.
Veröffentlicht: (2025)
Self-Distillation for Multi-Token Prediction
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026)
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
Pair-In, Pair-Out: Latent Multi-Token Prediction for Efficient LLMs
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
Navigating Large-Scale Document Collections: MuDABench for Multi-Document Analytical QA
von: Li, Zhanli, et al.
Veröffentlicht: (2026)
von: Li, Zhanli, et al.
Veröffentlicht: (2026)
VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
A Scaling Law for Token Efficiency in LLM Fine-Tuning Under Fixed Compute Budgets
von: Lagasse, Ryan, et al.
Veröffentlicht: (2025)
von: Lagasse, Ryan, et al.
Veröffentlicht: (2025)
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
von: Qiu, Zeju, et al.
Veröffentlicht: (2026)
von: Qiu, Zeju, et al.
Veröffentlicht: (2026)
Is Your LLM Really Mastering the Concept? A Multi-Agent Benchmark
von: Xu, Shuhang, et al.
Veröffentlicht: (2025)
von: Xu, Shuhang, et al.
Veröffentlicht: (2025)
Pre-Training Curriculum for Multi-Token Prediction in Language Models
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2025)
von: Aynetdinov, Ansar, et al.
Veröffentlicht: (2025)
Multi-LLM Collaborative Search for Complex Problem Solving
von: Yang, Sen, et al.
Veröffentlicht: (2025)
von: Yang, Sen, et al.
Veröffentlicht: (2025)
Token Statistics Reveal Conversational Drift in Multi-turn LLM Interaction
von: Hafez, Wael, et al.
Veröffentlicht: (2026)
von: Hafez, Wael, et al.
Veröffentlicht: (2026)
JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency
von: Cai, Aichen, et al.
Veröffentlicht: (2026)
von: Cai, Aichen, et al.
Veröffentlicht: (2026)
Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models
von: Liu, Peijie, et al.
Veröffentlicht: (2025)
von: Liu, Peijie, et al.
Veröffentlicht: (2025)
S$^2$-MAD: Breaking the Token Barrier to Enhance Multi-Agent Debate Efficiency
von: Zeng, Yuting, et al.
Veröffentlicht: (2025)
von: Zeng, Yuting, et al.
Veröffentlicht: (2025)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
von: Li, Zheng, et al.
Veröffentlicht: (2025)
von: Li, Zheng, et al.
Veröffentlicht: (2025)
Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs
von: Mei, Yixuan, et al.
Veröffentlicht: (2026)
von: Mei, Yixuan, et al.
Veröffentlicht: (2026)
Zodiac: A Cardiologist-Level LLM Framework for Multi-Agent Diagnostics
von: Zhou, Yuan, et al.
Veröffentlicht: (2024)
von: Zhou, Yuan, et al.
Veröffentlicht: (2024)
Optima: Optimizing Effectiveness and Efficiency for LLM-Based Multi-Agent System
von: Chen, Weize, et al.
Veröffentlicht: (2024)
von: Chen, Weize, et al.
Veröffentlicht: (2024)
AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent Collaboration
von: Wang, Zhexuan, et al.
Veröffentlicht: (2025)
von: Wang, Zhexuan, et al.
Veröffentlicht: (2025)
Direct Multi-Token Decoding
von: Luo, Xuan, et al.
Veröffentlicht: (2025)
von: Luo, Xuan, et al.
Veröffentlicht: (2025)
Adaptive Stopping for Multi-Turn LLM Reasoning
von: Zhou, Xiaofan, et al.
Veröffentlicht: (2026)
von: Zhou, Xiaofan, et al.
Veröffentlicht: (2026)
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Multi-Token Prediction Needs Registers
von: Gerontopoulos, Anastasios, et al.
Veröffentlicht: (2025)
von: Gerontopoulos, Anastasios, et al.
Veröffentlicht: (2025)
CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question Answering
von: Yang, Hao, et al.
Veröffentlicht: (2026)
von: Yang, Hao, et al.
Veröffentlicht: (2026)
Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement
von: Zhong, Qimin, et al.
Veröffentlicht: (2026)
von: Zhong, Qimin, et al.
Veröffentlicht: (2026)
Learning to Align Multi-Faceted Evaluation: A Unified and Robust Framework
von: Xu, Kaishuai, et al.
Veröffentlicht: (2025)
von: Xu, Kaishuai, et al.
Veröffentlicht: (2025)
Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference
von: Li, Zeping, et al.
Veröffentlicht: (2024)
von: Li, Zeping, et al.
Veröffentlicht: (2024)
LLM-Oriented Token-Adaptive Knowledge Distillation
von: Xie, Xurong, et al.
Veröffentlicht: (2025)
von: Xie, Xurong, et al.
Veröffentlicht: (2025)
Consistency of Large Reasoning Models Under Multi-Turn Attacks
von: Li, Yubo, et al.
Veröffentlicht: (2026)
von: Li, Yubo, et al.
Veröffentlicht: (2026)
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
von: Pan, Wenbo, et al.
Veröffentlicht: (2025)
Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning
von: Pang, Jinlong, et al.
Veröffentlicht: (2025)
von: Pang, Jinlong, et al.
Veröffentlicht: (2025)
Hybrid OCR-LLM Framework for Enterprise-Scale Document Information Extraction Under Copy-heavy Task
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
von: Wang, Zilong, et al.
Veröffentlicht: (2025)
Token-level Data Selection for Safe LLM Fine-tuning
von: Li, Yanping, et al.
Veröffentlicht: (2026)
von: Li, Yanping, et al.
Veröffentlicht: (2026)
A Multi-LLM Debiasing Framework
von: Owens, Deonna M., et al.
Veröffentlicht: (2024)
von: Owens, Deonna M., et al.
Veröffentlicht: (2024)
Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling
von: Yu, Yao-Ching, et al.
Veröffentlicht: (2024)
von: Yu, Yao-Ching, et al.
Veröffentlicht: (2024)
LAWCAT: Efficient Distillation from Quadratic to Linear Attention with Convolution across Tokens for Long Context Modeling
von: Liu, Zeyu, et al.
Veröffentlicht: (2025)
von: Liu, Zeyu, et al.
Veröffentlicht: (2025)
Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization
von: Xing, Tiancheng, et al.
Veröffentlicht: (2025)
von: Xing, Tiancheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models
von: Li, Cheng, et al.
Veröffentlicht: (2025) -
Video-VoT-R1: An efficient video inference model integrating image packing and AoE architecture
von: Li, Cheng, et al.
Veröffentlicht: (2025) -
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
von: Li, Cheng, et al.
Veröffentlicht: (2025) -
Self-Distillation for Multi-Token Prediction
von: Zhao, Guoliang, et al.
Veröffentlicht: (2026) -
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)