SirLLM: Streaming Infinite Retentive LLM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yao, Yao, Li, Zuchao, Zhao, Hai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GKT: A Novel Guidance-Based Knowledge Transfer Framework For Efficient Cloud-edge Collaboration LLM Deployment
von: Yao, Yao, et al.
Veröffentlicht: (2024)
von: Yao, Yao, et al.
Veröffentlicht: (2024)
Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
von: Shi, Luohe, et al.
Veröffentlicht: (2024)
von: Shi, Luohe, et al.
Veröffentlicht: (2024)
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models
von: Yao, Yao, et al.
Veröffentlicht: (2023)
von: Yao, Yao, et al.
Veröffentlicht: (2023)
Reference Trustable Decoding: A Training-Free Augmentation Paradigm for Large Language Models
von: Shi, Luohe, et al.
Veröffentlicht: (2024)
von: Shi, Luohe, et al.
Veröffentlicht: (2024)
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
von: Shi, Luohe, et al.
Veröffentlicht: (2025)
von: Shi, Luohe, et al.
Veröffentlicht: (2025)
How Deep is Love in LLMs' Hearts? Exploring Semantic Size in Human-like Cognition
von: Yao, Yao, et al.
Veröffentlicht: (2025)
von: Yao, Yao, et al.
Veröffentlicht: (2025)
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
Faster MoE LLM Inference for Extremely Large Models
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)
Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions
von: Ma, Xinbei, et al.
Veröffentlicht: (2024)
von: Ma, Xinbei, et al.
Veröffentlicht: (2024)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
von: Xu, Ruyi, et al.
Veröffentlicht: (2025)
Uni-ASR: Unified LLM-Based Architecture for Non-Streaming and Streaming Automatic Speech Recognition
von: Xia, Yinfeng, et al.
Veröffentlicht: (2026)
von: Xia, Yinfeng, et al.
Veröffentlicht: (2026)
DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
LESA: Learnable LLM Layer Scaling-Up
von: Yang, Yifei, et al.
Veröffentlicht: (2025)
von: Yang, Yifei, et al.
Veröffentlicht: (2025)
Venturing into Uncharted Waters: The Navigation Compass from Transformer to Mamba
von: Zou, Yuchen, et al.
Veröffentlicht: (2024)
von: Zou, Yuchen, et al.
Veröffentlicht: (2024)
ChainStream: An LLM-based Framework for Unified Synthetic Sensing
von: Liu, Jiacheng, et al.
Veröffentlicht: (2024)
von: Liu, Jiacheng, et al.
Veröffentlicht: (2024)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
von: Yang, Dongjie, et al.
Veröffentlicht: (2024)
von: Yang, Dongjie, et al.
Veröffentlicht: (2024)
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
A Coin Has Two Sides: A Novel Detector-Corrector Framework for Chinese Spelling Correction
von: Zeng, Xiangke, et al.
Veröffentlicht: (2024)
von: Zeng, Xiangke, et al.
Veröffentlicht: (2024)
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
von: Shi, Luohe, et al.
Veröffentlicht: (2025)
von: Shi, Luohe, et al.
Veröffentlicht: (2025)
LLM as Effective Streaming Processor: Bridging Streaming-Batch Mismatches with Group Position Encoding
von: Tong, Junlong, et al.
Veröffentlicht: (2025)
von: Tong, Junlong, et al.
Veröffentlicht: (2025)
DIY-MKG: An LLM-Based Polyglot Language Learning System
von: Tang, Kenan, et al.
Veröffentlicht: (2025)
von: Tang, Kenan, et al.
Veröffentlicht: (2025)
Sparse is Enough in Fine-tuning Pre-trained Large Language Models
von: Song, Weixi, et al.
Veröffentlicht: (2023)
von: Song, Weixi, et al.
Veröffentlicht: (2023)
Semantics-Preserved Distortion for Personal Privacy Protection in Information Management
von: Li, Jiajia, et al.
Veröffentlicht: (2022)
von: Li, Jiajia, et al.
Veröffentlicht: (2022)
Segment First or Comprehend First? Explore the Limit of Unsupervised Word Segmentation with Large Language Models
von: Zhang, Zihong, et al.
Veröffentlicht: (2025)
von: Zhang, Zihong, et al.
Veröffentlicht: (2025)
Merging Beyond: Streaming LLM Updates via Activation-Guided Rotations
von: Yao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Yao, Yuxuan, et al.
Veröffentlicht: (2026)
Wide-Horizon Thinking and Simulation-Based Evaluation for Real-World LLM Planning with Multifaceted Constraints
von: Yang, Dongjie, et al.
Veröffentlicht: (2025)
von: Yang, Dongjie, et al.
Veröffentlicht: (2025)
Efficient and Effective Internal Memory Retrieval for LLM-Based Healthcare Prediction
von: Li, Mingchen, et al.
Veröffentlicht: (2026)
von: Li, Mingchen, et al.
Veröffentlicht: (2026)
Game Development as Human-LLM Interaction
von: Hong, Jiale, et al.
Veröffentlicht: (2024)
von: Hong, Jiale, et al.
Veröffentlicht: (2024)
Dissecting Human and LLM Preferences
von: Li, Junlong, et al.
Veröffentlicht: (2024)
von: Li, Junlong, et al.
Veröffentlicht: (2024)
Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph
von: Zhao, Yibo, et al.
Veröffentlicht: (2024)
von: Zhao, Yibo, et al.
Veröffentlicht: (2024)
ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language Models
von: Guo, Jiani, et al.
Veröffentlicht: (2025)
von: Guo, Jiani, et al.
Veröffentlicht: (2025)
Smart Audit System Empowered by LLM
von: Yao, Xu, et al.
Veröffentlicht: (2024)
von: Yao, Xu, et al.
Veröffentlicht: (2024)
Semantic Graphs for Syntactic Simplification: A Revisit from the Age of LLM
von: Yao, Peiran, et al.
Veröffentlicht: (2024)
von: Yao, Peiran, et al.
Veröffentlicht: (2024)
Dual Reasoning: A GNN-LLM Collaborative Framework for Knowledge Graph Question Answering
von: Liu, Guangyi, et al.
Veröffentlicht: (2024)
von: Liu, Guangyi, et al.
Veröffentlicht: (2024)
Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
von: Yao, Xintong
Veröffentlicht: (2026)
von: Yao, Xintong
Veröffentlicht: (2026)
VHASR: A Multimodal Speech Recognition System With Vision Hotwords
von: Hu, Jiliang, et al.
Veröffentlicht: (2024)
von: Hu, Jiliang, et al.
Veröffentlicht: (2024)
OPEN-THEATRE: An Open-Source Toolkit for LLM-based Interactive Drama
von: Xu, Tianyang, et al.
Veröffentlicht: (2025)
von: Xu, Tianyang, et al.
Veröffentlicht: (2025)
Rethinking LLM Language Adaptation: A Case Study on Chinese Mixtral
von: Cui, Yiming, et al.
Veröffentlicht: (2024)
von: Cui, Yiming, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
GKT: A Novel Guidance-Based Knowledge Transfer Framework For Efficient Cloud-edge Collaboration LLM Deployment
von: Yao, Yao, et al.
Veröffentlicht: (2024) -
Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
von: Shi, Luohe, et al.
Veröffentlicht: (2024) -
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models
von: Yao, Yao, et al.
Veröffentlicht: (2023) -
Reference Trustable Decoding: A Training-Free Augmentation Paradigm for Large Language Models
von: Shi, Luohe, et al.
Veröffentlicht: (2024) -
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
von: Yang, Haoqi, et al.
Veröffentlicht: (2025)