CQIL: Inference Latency Optimization with Concurrent Computation of Quasi-Independent Layers
Fuente:
arXiv
Guardado en:
| Autores principales: | Zou, Longwei, Wang, Qingyang, Zhao, Han, Kong, Jiangang, Yang, Yi, Deng, Yangdong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Multi-Level Framework for Accelerating Training Transformer Models
por: Zou, Longwei, et al.
Publicado: (2024)
por: Zou, Longwei, et al.
Publicado: (2024)
InstCache: A Predictive Cache for LLM Serving
por: Zou, Longwei, et al.
Publicado: (2024)
por: Zou, Longwei, et al.
Publicado: (2024)
Scaling LLM Inference with Optimized Sample Compute Allocation
por: Zhang, Kexun, et al.
Publicado: (2024)
por: Zhang, Kexun, et al.
Publicado: (2024)
Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs
por: Hao, Jiangang
Publicado: (2026)
por: Hao, Jiangang
Publicado: (2026)
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
por: Ai, Xuan, et al.
Publicado: (2026)
por: Ai, Xuan, et al.
Publicado: (2026)
Probabilistic Concurrent Reasoning in Outcome Logic: Independence, Conditioning, and Invariants
por: Zilberstein, Noam, et al.
Publicado: (2024)
por: Zilberstein, Noam, et al.
Publicado: (2024)
Quasi-stratified Order Semantics of Concurrency
por: Koutny, Maciej, et al.
Publicado: (2024)
por: Koutny, Maciej, et al.
Publicado: (2024)
Slang Context-based Inference Enhancement via Greedy Search-Guided Chain-of-Thought Prompting
por: Cao, Jinghan, et al.
Publicado: (2026)
por: Cao, Jinghan, et al.
Publicado: (2026)
Dynamic Adaptive Attention and Supervised Contrastive Learning: A Novel Hybrid Framework for Text Sentiment Classification
por: Li, Qingyang
Publicado: (2026)
por: Li, Qingyang
Publicado: (2026)
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
por: Rodionov, Gleb, et al.
Publicado: (2025)
por: Rodionov, Gleb, et al.
Publicado: (2025)
LIONs: An Empirically Optimized Approach to Align Language Models
por: Yu, Xiao, et al.
Publicado: (2024)
por: Yu, Xiao, et al.
Publicado: (2024)
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
por: Gim, In, et al.
Publicado: (2023)
por: Gim, In, et al.
Publicado: (2023)
Latent Thought Models with Variational Bayes Inference-Time Computation
por: Kong, Deqian, et al.
Publicado: (2025)
por: Kong, Deqian, et al.
Publicado: (2025)
LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning
por: Kang, Beomseok, et al.
Publicado: (2025)
por: Kang, Beomseok, et al.
Publicado: (2025)
Expanding Computation Spaces of LLMs at Inference Time
por: Jang, Yoonna, et al.
Publicado: (2025)
por: Jang, Yoonna, et al.
Publicado: (2025)
Not All Layers of LLMs Are Necessary During Inference
por: Fan, Siqi, et al.
Publicado: (2024)
por: Fan, Siqi, et al.
Publicado: (2024)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
por: Yang, Seongjun, et al.
Publicado: (2023)
por: Yang, Seongjun, et al.
Publicado: (2023)
EmoVerse: Exploring Multimodal Large Language Models for Sentiment and Emotion Understanding
por: Li, Ao, et al.
Publicado: (2024)
por: Li, Ao, et al.
Publicado: (2024)
Acting Flatterers via LLMs Sycophancy: Combating Clickbait with LLMs Opposing-Stance Reasoning
por: Zhang, Chaowei, et al.
Publicado: (2026)
por: Zhang, Chaowei, et al.
Publicado: (2026)
Latency and Token-Aware Test-Time Compute
por: Huang, Jenny Y., et al.
Publicado: (2025)
por: Huang, Jenny Y., et al.
Publicado: (2025)
MemoryFormer: Minimize Transformer Computation by Removing Fully-Connected Layers
por: Ding, Ning, et al.
Publicado: (2024)
por: Ding, Ning, et al.
Publicado: (2024)
Inference Computation Scaling for Feature Augmentation in Recommendation Systems
por: Liu, Weihao, et al.
Publicado: (2025)
por: Liu, Weihao, et al.
Publicado: (2025)
Consecutive Batch Model Editing with HooK Layers
por: Li, Shuaiyi, et al.
Publicado: (2024)
por: Li, Shuaiyi, et al.
Publicado: (2024)
Inference Time Optimization with Confidence Dynamics
por: Wang, Yu, et al.
Publicado: (2026)
por: Wang, Yu, et al.
Publicado: (2026)
Test Security in Remote Testing Age: Perspectives from Process Data Analytics and AI
por: Hao, Jiangang, et al.
Publicado: (2024)
por: Hao, Jiangang, et al.
Publicado: (2024)
Inference Compute-Optimal Video Vision Language Models
por: Wang, Peiqi, et al.
Publicado: (2025)
por: Wang, Peiqi, et al.
Publicado: (2025)
MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts
por: Zhao, Yushu, et al.
Publicado: (2025)
por: Zhao, Yushu, et al.
Publicado: (2025)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
por: Husom, Erik Johannes, et al.
Publicado: (2025)
por: Husom, Erik Johannes, et al.
Publicado: (2025)
KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
por: Yang, Yifei, et al.
Publicado: (2024)
por: Yang, Yifei, et al.
Publicado: (2024)
CA*: Addressing Evaluation Pitfalls in Computation-Aware Latency for Simultaneous Speech Translation
por: Xu, Xi, et al.
Publicado: (2024)
por: Xu, Xi, et al.
Publicado: (2024)
GOSU: Retrieval-Augmented Generation with Global-Level Optimized Semantic Unit-Centric Framework
por: Zou, Xuecheng, et al.
Publicado: (2025)
por: Zou, Xuecheng, et al.
Publicado: (2025)
Iterative Layer Pruning for Efficient Translation Inference
por: Moslem, Yasmin, et al.
Publicado: (2025)
por: Moslem, Yasmin, et al.
Publicado: (2025)
DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
por: Yang, Ning, et al.
Publicado: (2025)
por: Yang, Ning, et al.
Publicado: (2025)
The Solution for the AIGC Inference Performance Optimization Competition
por: Pan, Sishun, et al.
Publicado: (2024)
por: Pan, Sishun, et al.
Publicado: (2024)
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
por: Li, Ruixiao, et al.
Publicado: (2025)
por: Li, Ruixiao, et al.
Publicado: (2025)
FlattenQuant: Breaking Through the Inference Compute-bound for Large Language Models with Per-tensor Quantization
por: Zhang, Yi, et al.
Publicado: (2024)
por: Zhang, Yi, et al.
Publicado: (2024)
AI-generated Essays: Characteristics and Implications on Automated Scoring and Academic Integrity
por: Zhong, Yang, et al.
Publicado: (2024)
por: Zhong, Yang, et al.
Publicado: (2024)
Optimizing Speech-Input Length for Speaker-Independent Depression Classification
por: Rutowski, Tomasz, et al.
Publicado: (2024)
por: Rutowski, Tomasz, et al.
Publicado: (2024)
StackingNet: Collective Inference Across Independent AI Foundation Models
por: Li, Siyang, et al.
Publicado: (2026)
por: Li, Siyang, et al.
Publicado: (2026)
TSO: Self-Training with Scaled Preference Optimization
por: Chen, Kaihui, et al.
Publicado: (2024)
por: Chen, Kaihui, et al.
Publicado: (2024)
Ejemplares similares
-
A Multi-Level Framework for Accelerating Training Transformer Models
por: Zou, Longwei, et al.
Publicado: (2024) -
InstCache: A Predictive Cache for LLM Serving
por: Zou, Longwei, et al.
Publicado: (2024) -
Scaling LLM Inference with Optimized Sample Compute Allocation
por: Zhang, Kexun, et al.
Publicado: (2024) -
Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs
por: Hao, Jiangang
Publicado: (2026) -
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
por: Ai, Xuan, et al.
Publicado: (2026)