Gespeichert in:
| Hauptverfasser: | Xiao, Qingfa, Wang, Jiachuan, Li, Haoyang, Deng, Cheng, Tang, Jiaqi, Li, Shuangyin, Zhang, Yongqi, Wang, Jun, Chen, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.13542 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adapting Pretrained Language Models for Citation Classification via Self-Supervised Contrastive Learning
von: Li, Tong, et al.
Veröffentlicht: (2025)
von: Li, Tong, et al.
Veröffentlicht: (2025)
R^2AG: Incorporating Retrieval Information into Retrieval Augmented Generation
von: Ye, Fuda, et al.
Veröffentlicht: (2024)
von: Ye, Fuda, et al.
Veröffentlicht: (2024)
PLM: Efficient Peripheral Language Models Hardware-Co-Designed for Ubiquitous Computing
von: Deng, Cheng, et al.
Veröffentlicht: (2025)
von: Deng, Cheng, et al.
Veröffentlicht: (2025)
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
AGRAG: Advanced Graph-based Retrieval-Augmented Generation for LLMs
von: Wang, Yubo, et al.
Veröffentlicht: (2025)
von: Wang, Yubo, et al.
Veröffentlicht: (2025)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
von: Tang, Jiaming, et al.
Veröffentlicht: (2024)
Query-focused and Memory-aware Reranker for Long Context Processing
von: Li, Yuqing, et al.
Veröffentlicht: (2026)
von: Li, Yuqing, et al.
Veröffentlicht: (2026)
Route Before Retrieve: Activating Latent Routing Abilities of LLMs for RAG vs. Long-Context Selection
von: Chen, Yiwen, et al.
Veröffentlicht: (2026)
von: Chen, Yiwen, et al.
Veröffentlicht: (2026)
Long-Short Alignment for Effective Long-Context Modeling in LLMs
von: Du, Tianqi, et al.
Veröffentlicht: (2025)
von: Du, Tianqi, et al.
Veröffentlicht: (2025)
Unshackling Context Length: An Efficient Selective Attention Approach through Query-Key Compression
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
Hamming Attention Distillation: Binarizing Keys and Queries for Efficient Long-Context Transformers
von: Horton, Mark, et al.
Veröffentlicht: (2025)
von: Horton, Mark, et al.
Veröffentlicht: (2025)
Learning Towards Emergence: Paving the Way to Induce Emergence by Inhibiting Monosemantic Neurons on Pre-trained Models
von: Wang, Jiachuan, et al.
Veröffentlicht: (2025)
von: Wang, Jiachuan, et al.
Veröffentlicht: (2025)
Cross-domain-aware Worker Selection with Training for Crowdsourced Annotation
von: Sun, Yushi, et al.
Veröffentlicht: (2024)
von: Sun, Yushi, et al.
Veröffentlicht: (2024)
S$^{2}$-DMs:Skip-Step Diffusion Models
von: Wang, Yixuan, et al.
Veröffentlicht: (2024)
von: Wang, Yixuan, et al.
Veröffentlicht: (2024)
Context Matters: Query-aware Dynamic Long Sequence Modeling of Gigapixel Images
von: Guo, Zhengrui, et al.
Veröffentlicht: (2025)
von: Guo, Zhengrui, et al.
Veröffentlicht: (2025)
TongSearch-QR: Reinforced Query Reasoning for Retrieval
von: Qin, Xubo, et al.
Veröffentlicht: (2025)
von: Qin, Xubo, et al.
Veröffentlicht: (2025)
Inference Scaling for Long-Context Retrieval Augmented Generation
von: Yue, Zhenrui, et al.
Veröffentlicht: (2024)
von: Yue, Zhenrui, et al.
Veröffentlicht: (2024)
Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing
von: Ye, Xiaoju, et al.
Veröffentlicht: (2025)
von: Ye, Xiaoju, et al.
Veröffentlicht: (2025)
Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models
von: Lin, Zhenghao, et al.
Veröffentlicht: (2025)
von: Lin, Zhenghao, et al.
Veröffentlicht: (2025)
QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
von: Yan, Jianxin, et al.
Veröffentlicht: (2026)
von: Yan, Jianxin, et al.
Veröffentlicht: (2026)
CHESS: Context-aware Hierarchical Efficient Semantic Selection for Long-Context LLM Inference
von: Fei, Chao, et al.
Veröffentlicht: (2026)
von: Fei, Chao, et al.
Veröffentlicht: (2026)
LooGLE: Can Long-Context Language Models Understand Long Contexts?
von: Li, Jiaqi, et al.
Veröffentlicht: (2023)
von: Li, Jiaqi, et al.
Veröffentlicht: (2023)
Enhancing Long Context Performance in LLMs Through Inner Loop Query Mechanism
von: Tang, Yimin, et al.
Veröffentlicht: (2024)
von: Tang, Yimin, et al.
Veröffentlicht: (2024)
Exploring Training and Inference Scaling Laws in Generative Retrieval
von: Cai, Hongru, et al.
Veröffentlicht: (2025)
von: Cai, Hongru, et al.
Veröffentlicht: (2025)
Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
RAT: Retrieval Augmented Thoughts Elicit Context-Aware Reasoning in Long-Horizon Generation
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding
von: Chen, Guanzheng, et al.
Veröffentlicht: (2025)
von: Chen, Guanzheng, et al.
Veröffentlicht: (2025)
Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMs
von: Cui, Wanyun, et al.
Veröffentlicht: (2025)
von: Cui, Wanyun, et al.
Veröffentlicht: (2025)
InstInfer: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
von: Pan, Xiurui, et al.
Veröffentlicht: (2024)
von: Pan, Xiurui, et al.
Veröffentlicht: (2024)
SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval
von: Zou, Zichen, et al.
Veröffentlicht: (2026)
von: Zou, Zichen, et al.
Veröffentlicht: (2026)
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
von: Tang, Hongyin, et al.
Veröffentlicht: (2024)
von: Tang, Hongyin, et al.
Veröffentlicht: (2024)
Training-Inference Consistent Segmented Execution for Long-Context LLMs
von: Shang, Xianpeng, et al.
Veröffentlicht: (2026)
von: Shang, Xianpeng, et al.
Veröffentlicht: (2026)
Beyond Model Base Retrieval: Weaving Knowledge to Master Fine-grained Neural Network Design
von: Wang, Jialiang, et al.
Veröffentlicht: (2025)
von: Wang, Jialiang, et al.
Veröffentlicht: (2025)
RxnNano:Training Compact LLMs for Chemical Reaction and Retrosynthesis Prediction via Hierarchical Curriculum Learning
von: Li, Ran, et al.
Veröffentlicht: (2026)
von: Li, Ran, et al.
Veröffentlicht: (2026)
Probe and Skip: Self-Predictive Token Skipping for Efficient Long-Context LLM Inference
von: Wu, Zimeng, et al.
Veröffentlicht: (2026)
von: Wu, Zimeng, et al.
Veröffentlicht: (2026)
Why Does the Effective Context Length of LLMs Fall Short?
von: An, Chenxin, et al.
Veröffentlicht: (2024)
von: An, Chenxin, et al.
Veröffentlicht: (2024)
A Posteriori Error Estimation Improved by a Reconstruction Operator for the Stokes Optimal Control Problem
von: Li, Jingshi, et al.
Veröffentlicht: (2025)
von: Li, Jingshi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Adapting Pretrained Language Models for Citation Classification via Self-Supervised Contrastive Learning
von: Li, Tong, et al.
Veröffentlicht: (2025) -
R^2AG: Incorporating Retrieval Information into Retrieval Augmented Generation
von: Ye, Fuda, et al.
Veröffentlicht: (2024) -
PLM: Efficient Peripheral Language Models Hardware-Co-Designed for Ubiquitous Computing
von: Deng, Cheng, et al.
Veröffentlicht: (2025) -
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
von: Li, Haoyang, et al.
Veröffentlicht: (2025) -
AGRAG: Advanced Graph-based Retrieval-Augmented Generation for LLMs
von: Wang, Yubo, et al.
Veröffentlicht: (2025)