Reference Trustable Decoding: A Training-Free Augmentation Paradigm for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Luohe, Yao, Yao, Li, Zuchao, Zhang, Lefei, Zhao, Hai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
by: Shi, Luohe, et al.
Published: (2025)
by: Shi, Luohe, et al.
Published: (2025)
Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
by: Shi, Luohe, et al.
Published: (2024)
by: Shi, Luohe, et al.
Published: (2024)
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
by: Shi, Luohe, et al.
Published: (2025)
by: Shi, Luohe, et al.
Published: (2025)
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
by: Zhang, Zihong, et al.
Published: (2026)
by: Zhang, Zihong, et al.
Published: (2026)
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models
by: Yao, Yao, et al.
Published: (2023)
by: Yao, Yao, et al.
Published: (2023)
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers
by: Tang, Zicong, et al.
Published: (2025)
by: Tang, Zicong, et al.
Published: (2025)
Sparse is Enough in Fine-tuning Pre-trained Large Language Models
by: Song, Weixi, et al.
Published: (2023)
by: Song, Weixi, et al.
Published: (2023)
Segment First or Comprehend First? Explore the Limit of Unsupervised Word Segmentation with Large Language Models
by: Zhang, Zihong, et al.
Published: (2025)
by: Zhang, Zihong, et al.
Published: (2025)
GKT: A Novel Guidance-Based Knowledge Transfer Framework For Efficient Cloud-edge Collaboration LLM Deployment
by: Yao, Yao, et al.
Published: (2024)
by: Yao, Yao, et al.
Published: (2024)
SirLLM: Streaming Infinite Retentive LLM
by: Yao, Yao, et al.
Published: (2024)
by: Yao, Yao, et al.
Published: (2024)
Faster MoE LLM Inference for Extremely Large Models
by: Yang, Haoqi, et al.
Published: (2025)
by: Yang, Haoqi, et al.
Published: (2025)
From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic Horizons
by: Ma, Xiangyu, et al.
Published: (2026)
by: Ma, Xiangyu, et al.
Published: (2026)
ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language Models
by: Guo, Jiani, et al.
Published: (2025)
by: Guo, Jiani, et al.
Published: (2025)
Venturing into Uncharted Waters: The Navigation Compass from Transformer to Mamba
by: Zou, Yuchen, et al.
Published: (2024)
by: Zou, Yuchen, et al.
Published: (2024)
A Coin Has Two Sides: A Novel Detector-Corrector Framework for Chinese Spelling Correction
by: Zeng, Xiangke, et al.
Published: (2024)
by: Zeng, Xiangke, et al.
Published: (2024)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
by: Tang, Zicong, et al.
Published: (2025)
by: Tang, Zicong, et al.
Published: (2025)
Model Hemorrhage and the Robustness Limits of Large Language Models
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
VHASR: A Multimodal Speech Recognition System With Vision Hotwords
by: Hu, Jiliang, et al.
Published: (2024)
by: Hu, Jiliang, et al.
Published: (2024)
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
by: Yang, Haoqi, et al.
Published: (2025)
by: Yang, Haoqi, et al.
Published: (2025)
Towards Trustable Language Models: Investigating Information Quality of Large Language Models
by: Rejeleene, Rick, et al.
Published: (2024)
by: Rejeleene, Rick, et al.
Published: (2024)
How Deep is Love in LLMs' Hearts? Exploring Semantic Size in Human-like Cognition
by: Yao, Yao, et al.
Published: (2025)
by: Yao, Yao, et al.
Published: (2025)
Label Drop for Multi-Aspect Relation Modeling in Universal Information Extraction
by: Yang, Lu, et al.
Published: (2025)
by: Yang, Lu, et al.
Published: (2025)
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
A Learning Rate Path Switching Training Paradigm for Version Updates of Large Language Models
by: Wang, Zhihao, et al.
Published: (2024)
by: Wang, Zhihao, et al.
Published: (2024)
SongSong: A Time Phonograph for Chinese SongCi Music from Thousand of Years Away
by: Li, Jiajia, et al.
Published: (2026)
by: Li, Jiajia, et al.
Published: (2026)
DHI: Leveraging Diverse Hallucination Induction for Enhanced Contrastive Factuality Control in Large Language Models
by: Guo, Jiani, et al.
Published: (2026)
by: Guo, Jiani, et al.
Published: (2026)
DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
Entropy-Based Decoding for Retrieval-Augmented Large Language Models
by: Qiu, Zexuan, et al.
Published: (2024)
by: Qiu, Zexuan, et al.
Published: (2024)
Centroid-centered Modeling for Efficient Vision Transformer Pre-training
by: Yan, Xin, et al.
Published: (2023)
by: Yan, Xin, et al.
Published: (2023)
MGIMM: Multi-Granularity Instruction Multimodal Model for Attribute-Guided Remote Sensing Image Detailed Description
by: Yang, Cong, et al.
Published: (2024)
by: Yang, Cong, et al.
Published: (2024)
SongSage: A Large Musical Language Model with Lyric Generative Pre-training
by: Guo, Jiani, et al.
Published: (2026)
by: Guo, Jiani, et al.
Published: (2026)
Thinking with Nothinking Calibration: A New In-Context Learning Paradigm in Reasoning Large Language Models
by: Wu, Haotian, et al.
Published: (2025)
by: Wu, Haotian, et al.
Published: (2025)
Decoding in Order-Agnostic Language Models: Chain-Rule Deviation and Uniform Spreading
by: Yao, Lin
Published: (2026)
by: Yao, Lin
Published: (2026)
RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models
by: Tran, Hieu, et al.
Published: (2024)
by: Tran, Hieu, et al.
Published: (2024)
Aggressive Post-Training Compression on Extremely Large Language Models
by: Zhang, Zining, et al.
Published: (2024)
by: Zhang, Zining, et al.
Published: (2024)
Data Augmentation using Large Language Models: Data Perspectives, Learning Paradigms and Challenges
by: Ding, Bosheng, et al.
Published: (2024)
by: Ding, Bosheng, et al.
Published: (2024)
How Do Large Language Models Learn Concepts During Continual Pre-Training?
by: Yao, Barry Menglong, et al.
Published: (2026)
by: Yao, Barry Menglong, et al.
Published: (2026)
A Systematic Study of Training-Free Methods for Trustworthy Large Language Models
by: Si, Wai Man, et al.
Published: (2026)
by: Si, Wai Man, et al.
Published: (2026)
SEPTQ: A Simple and Effective Post-Training Quantization Paradigm for Large Language Models
by: Liu, Han, et al.
Published: (2026)
by: Liu, Han, et al.
Published: (2026)
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding
by: Yao, Feiyu, et al.
Published: (2025)
by: Yao, Feiyu, et al.
Published: (2025)
Similar Items
-
Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
by: Shi, Luohe, et al.
Published: (2025) -
Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
by: Shi, Luohe, et al.
Published: (2024) -
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
by: Shi, Luohe, et al.
Published: (2025) -
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
by: Zhang, Zihong, et al.
Published: (2026) -
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models
by: Yao, Yao, et al.
Published: (2023)