vLLM Hook v0: A Plug-in for Programming Model Internals on vLLM
Fuente:
arXiv
Guardado en:
| Autores principales: | Ko, Ching-Yun, Chen, Pin-Yu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
When to Reason: Semantic Router for vLLM
por: Wang, Chen, et al.
Publicado: (2025)
por: Wang, Chen, et al.
Publicado: (2025)
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
por: Martinez, Matias
Publicado: (2024)
por: Martinez, Matias
Publicado: (2024)
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM
por: Kiruluta, Andrew, et al.
Publicado: (2025)
por: Kiruluta, Andrew, et al.
Publicado: (2025)
Embedding the Teacher: Distilling vLLM Preferences for Scalable Image Retrieval
por: He, Eric, et al.
Publicado: (2025)
por: He, Eric, et al.
Publicado: (2025)
98$\times$ Faster LLM Routing Without a Dedicated GPU: Flash Attention, Prompt Compression, and Near-Streaming for the vLLM Semantic Router
por: Liu, Xunzhuo, et al.
Publicado: (2026)
por: Liu, Xunzhuo, et al.
Publicado: (2026)
The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project
por: Chen, Huamin, et al.
Publicado: (2026)
por: Chen, Huamin, et al.
Publicado: (2026)
Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-tuning
por: Lee, Yu-Ang, et al.
Publicado: (2026)
por: Lee, Yu-Ang, et al.
Publicado: (2026)
vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models
por: Yin, Peiqi, et al.
Publicado: (2026)
por: Yin, Peiqi, et al.
Publicado: (2026)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
por: Kolluru, Saicharan
Publicado: (2025)
por: Kolluru, Saicharan
Publicado: (2025)
hdl2v: A Code Translation Dataset for Enhanced LLM Verilog Generation
por: Hong, Charles, et al.
Publicado: (2025)
por: Hong, Charles, et al.
Publicado: (2025)
AIOS Compiler: LLM as Interpreter for Natural Language Programming and Flow Programming of AI Agents
por: Xu, Shuyuan, et al.
Publicado: (2024)
por: Xu, Shuyuan, et al.
Publicado: (2024)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
por: Wei, Anjiang, et al.
Publicado: (2025)
por: Wei, Anjiang, et al.
Publicado: (2025)
Benchmarking Energy Efficiency of Large Language Models Using vLLM
por: Pronk, K., et al.
Publicado: (2025)
por: Pronk, K., et al.
Publicado: (2025)
vCache: Verified Semantic Prompt Caching
por: Schroeder, Luis Gaspar, et al.
Publicado: (2025)
por: Schroeder, Luis Gaspar, et al.
Publicado: (2025)
More Than a Score: Probing the Impact of Prompt Specificity on LLM Code Generation
por: Zi, Yangtian, et al.
Publicado: (2025)
por: Zi, Yangtian, et al.
Publicado: (2025)
Why Reasoning Matters? A Survey of Advancements in Multimodal Reasoning (v1)
por: Bi, Jing, et al.
Publicado: (2025)
por: Bi, Jing, et al.
Publicado: (2025)
GLOCON Database: Design Decisions and User Manual (v1.0)
por: Hürriyetoğlu, Ali, et al.
Publicado: (2024)
por: Hürriyetoğlu, Ali, et al.
Publicado: (2024)
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
por: Kandpal, Nikhil, et al.
Publicado: (2025)
por: Kandpal, Nikhil, et al.
Publicado: (2025)
Essential-Web v1.0: 24T tokens of organized web data
por: AI, Essential, et al.
Publicado: (2025)
por: AI, Essential, et al.
Publicado: (2025)
DistiLLM: Towards Streamlined Distillation for Large Language Models
por: Ko, Jongwoo, et al.
Publicado: (2024)
por: Ko, Jongwoo, et al.
Publicado: (2024)
From Belief Entrenchment to Robust Reasoning in LLM Agents
por: Oh, Jihwan, et al.
Publicado: (2025)
por: Oh, Jihwan, et al.
Publicado: (2025)
STAR: Spectral Truncation and Rescale for Model Merging
por: Lee, Yu-Ang, et al.
Publicado: (2025)
por: Lee, Yu-Ang, et al.
Publicado: (2025)
Decoding the Hook: A Multimodal LLM Framework for Analyzing the Hooking Period of Video Ads
por: Zhang, Kunpeng, et al.
Publicado: (2026)
por: Zhang, Kunpeng, et al.
Publicado: (2026)
QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic Approach
por: Dong, Shouyang, et al.
Publicado: (2025)
por: Dong, Shouyang, et al.
Publicado: (2025)
CodeCloak: A Method for Evaluating and Mitigating Code Leakage by LLM Code Assistants
por: Noah, Amit Finkman, et al.
Publicado: (2024)
por: Noah, Amit Finkman, et al.
Publicado: (2024)
Emergent Representations of Program Semantics in Language Models Trained on Programs
por: Jin, Charles, et al.
Publicado: (2023)
por: Jin, Charles, et al.
Publicado: (2023)
APPL: A Prompt Programming Language for Harmonious Integration of Programs and Large Language Model Prompts
por: Dong, Honghua, et al.
Publicado: (2024)
por: Dong, Honghua, et al.
Publicado: (2024)
Automatic Generation of Python Programs Using Context-Free Grammars
por: Yamani, Kamel, et al.
Publicado: (2024)
por: Yamani, Kamel, et al.
Publicado: (2024)
A Comparative Analysis of LLM Memorization at Statistical and Internal Levels: Cross-Model Commonalities and Model-Specific Signatures
por: Chen, Bowen, et al.
Publicado: (2026)
por: Chen, Bowen, et al.
Publicado: (2026)
vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models
por: Liu, Xunzhuo, et al.
Publicado: (2026)
por: Liu, Xunzhuo, et al.
Publicado: (2026)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
por: Barone, Antonio Valerio Miceli, et al.
Publicado: (2026)
por: Barone, Antonio Valerio Miceli, et al.
Publicado: (2026)
SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation
por: Petrukha, Ivan, et al.
Publicado: (2025)
por: Petrukha, Ivan, et al.
Publicado: (2025)
Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning Models
por: Yin, Huifeng, et al.
Publicado: (2025)
por: Yin, Huifeng, et al.
Publicado: (2025)
Automated Dynamic AI Inference Scaling on HPC-Infrastructure: Integrating Kubernetes, Slurm and vLLM
por: Trappen, Tim, et al.
Publicado: (2025)
por: Trappen, Tim, et al.
Publicado: (2025)
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
por: Wang, Xinyi, et al.
Publicado: (2024)
por: Wang, Xinyi, et al.
Publicado: (2024)
A Domain-Specific Language for LLM-Driven Trigger Generation in Multimodal Data Collection
por: Reis, Philipp, et al.
Publicado: (2026)
por: Reis, Philipp, et al.
Publicado: (2026)
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
por: Xia, Chunqiu Steven, et al.
Publicado: (2024)
por: Xia, Chunqiu Steven, et al.
Publicado: (2024)
BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs
por: Wang, Hongyu, et al.
Publicado: (2025)
por: Wang, Hongyu, et al.
Publicado: (2025)
Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications
por: Kudlur, Manjunath, et al.
Publicado: (2026)
por: Kudlur, Manjunath, et al.
Publicado: (2026)
Large Language Models can be Strong Self-Detoxifiers
por: Ko, Ching-Yun, et al.
Publicado: (2024)
por: Ko, Ching-Yun, et al.
Publicado: (2024)
Ejemplares similares
-
When to Reason: Semantic Router for vLLM
por: Wang, Chen, et al.
Publicado: (2025) -
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
por: Martinez, Matias
Publicado: (2024) -
History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM
por: Kiruluta, Andrew, et al.
Publicado: (2025) -
Embedding the Teacher: Distilling vLLM Preferences for Scalable Image Retrieval
por: He, Eric, et al.
Publicado: (2025) -
98$\times$ Faster LLM Routing Without a Dedicated GPU: Flash Attention, Prompt Compression, and Near-Streaming for the vLLM Semantic Router
por: Liu, Xunzhuo, et al.
Publicado: (2026)