Gespeichert in:
| Hauptverfasser: | Shi, Luohe, Yao, Yao, Li, Zuchao, Zhang, Lefei, Zhao, Hai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2409.20181 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
von: Shi, Luohe, et al.
Veröffentlicht: (2025)
von: Shi, Luohe, et al.
Veröffentlicht: (2025)
Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
von: Shi, Luohe, et al.
Veröffentlicht: (2024)
von: Shi, Luohe, et al.
Veröffentlicht: (2024)
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
von: Shi, Luohe, et al.
Veröffentlicht: (2025)
von: Shi, Luohe, et al.
Veröffentlicht: (2025)
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models
von: Yao, Yao, et al.
Veröffentlicht: (2023)
von: Yao, Yao, et al.
Veröffentlicht: (2023)
Sparse is Enough in Fine-tuning Pre-trained Large Language Models
von: Song, Weixi, et al.
Veröffentlicht: (2023)
von: Song, Weixi, et al.
Veröffentlicht: (2023)
Segment First or Comprehend First? Explore the Limit of Unsupervised Word Segmentation with Large Language Models
von: Zhang, Zihong, et al.
Veröffentlicht: (2025)
von: Zhang, Zihong, et al.
Veröffentlicht: (2025)
Faster MoE LLM Inference for Extremely Large Models
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
GKT: A Novel Guidance-Based Knowledge Transfer Framework For Efficient Cloud-edge Collaboration LLM Deployment
von: Yao, Yao, et al.
Veröffentlicht: (2024)
von: Yao, Yao, et al.
Veröffentlicht: (2024)
From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic Horizons
von: Ma, Xiangyu, et al.
Veröffentlicht: (2026)
von: Ma, Xiangyu, et al.
Veröffentlicht: (2026)
SirLLM: Streaming Infinite Retentive LLM
von: Yao, Yao, et al.
Veröffentlicht: (2024)
von: Yao, Yao, et al.
Veröffentlicht: (2024)
Venturing into Uncharted Waters: The Navigation Compass from Transformer to Mamba
von: Zou, Yuchen, et al.
Veröffentlicht: (2024)
von: Zou, Yuchen, et al.
Veröffentlicht: (2024)
ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language Models
von: Guo, Jiani, et al.
Veröffentlicht: (2025)
von: Guo, Jiani, et al.
Veröffentlicht: (2025)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
A Coin Has Two Sides: A Novel Detector-Corrector Framework for Chinese Spelling Correction
von: Zeng, Xiangke, et al.
Veröffentlicht: (2024)
von: Zeng, Xiangke, et al.
Veröffentlicht: (2024)
Model Hemorrhage and the Robustness Limits of Large Language Models
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
VHASR: A Multimodal Speech Recognition System With Vision Hotwords
von: Hu, Jiliang, et al.
Veröffentlicht: (2024)
von: Hu, Jiliang, et al.
Veröffentlicht: (2024)
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
How Deep is Love in LLMs' Hearts? Exploring Semantic Size in Human-like Cognition
von: Yao, Yao, et al.
Veröffentlicht: (2025)
von: Yao, Yao, et al.
Veröffentlicht: (2025)
Label Drop for Multi-Aspect Relation Modeling in Universal Information Extraction
von: Yang, Lu, et al.
Veröffentlicht: (2025)
von: Yang, Lu, et al.
Veröffentlicht: (2025)
Towards Trustable Language Models: Investigating Information Quality of Large Language Models
von: Rejeleene, Rick, et al.
Veröffentlicht: (2024)
von: Rejeleene, Rick, et al.
Veröffentlicht: (2024)
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
SongSong: A Time Phonograph for Chinese SongCi Music from Thousand of Years Away
von: Li, Jiajia, et al.
Veröffentlicht: (2026)
von: Li, Jiajia, et al.
Veröffentlicht: (2026)
A Learning Rate Path Switching Training Paradigm for Version Updates of Large Language Models
von: Wang, Zhihao, et al.
Veröffentlicht: (2024)
von: Wang, Zhihao, et al.
Veröffentlicht: (2024)
Centroid-centered Modeling for Efficient Vision Transformer Pre-training
von: Yan, Xin, et al.
Veröffentlicht: (2023)
von: Yan, Xin, et al.
Veröffentlicht: (2023)
MGIMM: Multi-Granularity Instruction Multimodal Model for Attribute-Guided Remote Sensing Image Detailed Description
von: Yang, Cong, et al.
Veröffentlicht: (2024)
von: Yang, Cong, et al.
Veröffentlicht: (2024)
DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
DHI: Leveraging Diverse Hallucination Induction for Enhanced Contrastive Factuality Control in Large Language Models
von: Guo, Jiani, et al.
Veröffentlicht: (2026)
von: Guo, Jiani, et al.
Veröffentlicht: (2026)
SongSage: A Large Musical Language Model with Lyric Generative Pre-training
von: Guo, Jiani, et al.
Veröffentlicht: (2026)
von: Guo, Jiani, et al.
Veröffentlicht: (2026)
Multi-modal Auto-regressive Modeling via Visual Words
von: Peng, Tianshuo, et al.
Veröffentlicht: (2024)
von: Peng, Tianshuo, et al.
Veröffentlicht: (2024)
Entropy-Based Decoding for Retrieval-Augmented Large Language Models
von: Qiu, Zexuan, et al.
Veröffentlicht: (2024)
von: Qiu, Zexuan, et al.
Veröffentlicht: (2024)
Thinking with Nothinking Calibration: A New In-Context Learning Paradigm in Reasoning Large Language Models
von: Wu, Haotian, et al.
Veröffentlicht: (2025)
von: Wu, Haotian, et al.
Veröffentlicht: (2025)
Aggressive Post-Training Compression on Extremely Large Language Models
von: Zhang, Zining, et al.
Veröffentlicht: (2024)
von: Zhang, Zining, et al.
Veröffentlicht: (2024)
VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models
von: Tran, Hieu, et al.
Veröffentlicht: (2024)
von: Tran, Hieu, et al.
Veröffentlicht: (2024)
Decoding in Order-Agnostic Language Models: Chain-Rule Deviation and Uniform Spreading
von: Yao, Lin
Veröffentlicht: (2026)
von: Yao, Lin
Veröffentlicht: (2026)
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding
von: Yao, Feiyu, et al.
Veröffentlicht: (2025)
von: Yao, Feiyu, et al.
Veröffentlicht: (2025)
Data Augmentation using Large Language Models: Data Perspectives, Learning Paradigms and Challenges
von: Ding, Bosheng, et al.
Veröffentlicht: (2024)
von: Ding, Bosheng, et al.
Veröffentlicht: (2024)
How Do Large Language Models Learn Concepts During Continual Pre-Training?
von: Yao, Barry Menglong, et al.
Veröffentlicht: (2026)
von: Yao, Barry Menglong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
von: Shi, Luohe, et al.
Veröffentlicht: (2025) -
Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
von: Shi, Luohe, et al.
Veröffentlicht: (2024) -
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
von: Shi, Luohe, et al.
Veröffentlicht: (2025) -
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
von: Zhang, Zihong, et al.
Veröffentlicht: (2026) -
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers
von: Tang, Zicong, et al.
Veröffentlicht: (2025)