Saved in:
| Main Authors: | Zhang, Weifang, Nie, Yuzhou, Pang, Bowen, Ma, Guangrui, Wu, Shining |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2606.00516 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ORPilot: A Production-Oriented Agentic LLM-for-OR Tool for Optimization Modeling
by: Xie, Guangrui
Published: (2026)
by: Xie, Guangrui
Published: (2026)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
by: Lyu, Hongtao, et al.
Published: (2025)
by: Lyu, Hongtao, et al.
Published: (2025)
AI Can Be Cognitively Biased: An Exploratory Study on Threshold Priming in LLM-Based Batch Relevance Assessment
by: Chen, Nuo, et al.
Published: (2024)
by: Chen, Nuo, et al.
Published: (2024)
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
by: Zheng, Zhen, et al.
Published: (2024)
by: Zheng, Zhen, et al.
Published: (2024)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
by: Zheng, Wanyi, et al.
Published: (2025)
by: Zheng, Wanyi, et al.
Published: (2025)
Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
by: Shrestha, Susav, et al.
Published: (2025)
by: Shrestha, Susav, et al.
Published: (2025)
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
by: He, Junhui, et al.
Published: (2024)
by: He, Junhui, et al.
Published: (2024)
From Static to Dynamic: A Streaming RAG Approach to Real-time Knowledge Base
by: Zhu, Yuzhou
Published: (2025)
by: Zhu, Yuzhou
Published: (2025)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
Membership Inference Attack with Partial Features
by: Wang, Xurun, et al.
Published: (2025)
by: Wang, Xurun, et al.
Published: (2025)
VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
by: Nie, Yuzhou, et al.
Published: (2025)
by: Nie, Yuzhou, et al.
Published: (2025)
HALO: Memory-Centric Heterogeneous Accelerator with 2.5D Integration for Low-Batch LLM Inference
by: Negi, Shubham, et al.
Published: (2025)
by: Negi, Shubham, et al.
Published: (2025)
No Free Lunch Theorem for Privacy-Preserving LLM Inference
by: Zhang, Xiaojin, et al.
Published: (2024)
by: Zhang, Xiaojin, et al.
Published: (2024)
Decentralized AI: Permissionless LLM Inference on POKT Network
by: Olshansky, Daniel, et al.
Published: (2024)
by: Olshansky, Daniel, et al.
Published: (2024)
ChameleonLLM: Batch-Aware Dynamic Low-Rank Adaptation via Inference-Time Clusters
by: Yuksel, Kamer Ali, et al.
Published: (2025)
by: Yuksel, Kamer Ali, et al.
Published: (2025)
Combining LLM Semantic Reasoning with GNN Structural Modeling for Multi-View Multi-Label Feature Selection
by: Chen, Zhiqi, et al.
Published: (2025)
by: Chen, Zhiqi, et al.
Published: (2025)
BlockBatch: Multi-Scale Consensus Decoding for Efficient Diffusion Language Model Inference
by: Wu, Xiaoyou, et al.
Published: (2026)
by: Wu, Xiaoyou, et al.
Published: (2026)
Unified Sparse-Matrix Representations for Diverse Neural Architectures
by: Zhu, Yuzhou
Published: (2025)
by: Zhu, Yuzhou
Published: (2025)
Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
by: Pang, Yuqi, et al.
Published: (2025)
by: Pang, Yuqi, et al.
Published: (2025)
Batch-of-Thought: Cross-Instance Learning for Enhanced LLM Reasoning
by: Yang, Xuan, et al.
Published: (2026)
by: Yang, Xuan, et al.
Published: (2026)
SynerDiff: Synergetic Continuous Batching for Fast and Parallel Diffusion Model Inference
by: Zhou, Ziqi, et al.
Published: (2026)
by: Zhou, Ziqi, et al.
Published: (2026)
Assessing the Human-Likeness of LLM-Driven Digital Twins in Simulating Health Care System Trust
by: Wu, Yuzhou, et al.
Published: (2025)
by: Wu, Yuzhou, et al.
Published: (2025)
Edge Intelligence Optimization for Large Language Model Inference with Batching and Quantization
by: Zhang, Xinyuan, et al.
Published: (2024)
by: Zhang, Xinyuan, et al.
Published: (2024)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
The Feasibility of Topic-Based Watermarking on Academic Peer Reviews
by: Nemecek, Alexander, et al.
Published: (2025)
by: Nemecek, Alexander, et al.
Published: (2025)
Batch Speculative Decoding Done Right
by: Zhang, Ranran Haoran, et al.
Published: (2025)
by: Zhang, Ranran Haoran, et al.
Published: (2025)
Deeper Insights into Deep Graph Convolutional Networks: Stability and Generalization
by: Yang, Guangrui, et al.
Published: (2024)
by: Yang, Guangrui, et al.
Published: (2024)
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization
by: Pang, Bowen, et al.
Published: (2025)
by: Pang, Bowen, et al.
Published: (2025)
AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
by: Wu, Yixin, et al.
Published: (2025)
by: Wu, Yixin, et al.
Published: (2025)
Batch-in-Batch: a new adversarial training framework for initial perturbation and sample selection
by: Wu, Yinting, et al.
Published: (2024)
by: Wu, Yinting, et al.
Published: (2024)
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
by: Zhang, Tianyi, et al.
Published: (2024)
by: Zhang, Tianyi, et al.
Published: (2024)
A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints
by: Nie, Chengyi, et al.
Published: (2026)
by: Nie, Chengyi, et al.
Published: (2026)
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
by: Ma, Songchen, et al.
Published: (2026)
by: Ma, Songchen, et al.
Published: (2026)
MELA: Multilingual Evaluation of Linguistic Acceptability
by: Zhang, Ziyin, et al.
Published: (2023)
by: Zhang, Ziyin, et al.
Published: (2023)
PQCache: Product Quantization-based KVCache for Long Context LLM Inference
by: Zhang, Hailin, et al.
Published: (2024)
by: Zhang, Hailin, et al.
Published: (2024)
Architectural Backdoors for Within-Batch Data Stealing and Model Inference Manipulation
by: Küchler, Nicolas, et al.
Published: (2025)
by: Küchler, Nicolas, et al.
Published: (2025)
Design and Evaluation of Cost-Aware PoQ for Decentralized LLM Inference
by: Tian, Arther, et al.
Published: (2025)
by: Tian, Arther, et al.
Published: (2025)
Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline
by: Fang, Zhiyuan, et al.
Published: (2025)
by: Fang, Zhiyuan, et al.
Published: (2025)
ODAR: Principled Adaptive Routing for LLM Reasoning via Active Inference
by: Ma, Siyuan, et al.
Published: (2026)
by: Ma, Siyuan, et al.
Published: (2026)
An Aligned Constraint Programming Model For Serial Batch Scheduling With Minimum Batch Size
by: Huertas, Jorge A., et al.
Published: (2025)
by: Huertas, Jorge A., et al.
Published: (2025)
Similar Items
-
ORPilot: A Production-Oriented Agentic LLM-for-OR Tool for Optimization Modeling
by: Xie, Guangrui
Published: (2026) -
FairBatching: Fairness-Aware Batch Formation for LLM Inference
by: Lyu, Hongtao, et al.
Published: (2025) -
AI Can Be Cognitively Biased: An Exploratory Study on Threshold Priming in LLM-Based Batch Relevance Assessment
by: Chen, Nuo, et al.
Published: (2024) -
BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
by: Zheng, Zhen, et al.
Published: (2024) -
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
by: Zheng, Wanyi, et al.
Published: (2025)