When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Haorui, He, Zhenghui, Liu, Xuanzi, Xu, Yang, Liu, Dongsheng, Ma, Jiakang, Wu, Lupan, Wu, Yangjie, Tang, Xiongchao, Shi, Tianhui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
von: Ma, Xinyue, et al.
Veröffentlicht: (2026)
von: Ma, Xinyue, et al.
Veröffentlicht: (2026)
Same Voice, Different Lab: On the Homogenization of Frontier LLM Personalities
von: Krishna, Avinash, et al.
Veröffentlicht: (2026)
von: Krishna, Avinash, et al.
Veröffentlicht: (2026)
Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders
von: Patel, Het, et al.
Veröffentlicht: (2026)
von: Patel, Het, et al.
Veröffentlicht: (2026)
A Performance Evaluation of a Quantized Large Language Model on Various Smartphones
von: Çöplü, Tolga, et al.
Veröffentlicht: (2023)
von: Çöplü, Tolga, et al.
Veröffentlicht: (2023)
When Are Two RLHF Objectives the Same?
von: Gaikwad, Madhava
Veröffentlicht: (2025)
von: Gaikwad, Madhava
Veröffentlicht: (2025)
R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k
von: Jaber, Jaber, et al.
Veröffentlicht: (2026)
von: Jaber, Jaber, et al.
Veröffentlicht: (2026)
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2025)
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2025)
Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction
von: Garcia, Gabriel
Veröffentlicht: (2026)
von: Garcia, Gabriel
Veröffentlicht: (2026)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
von: Kolluru, Saicharan
Veröffentlicht: (2025)
von: Kolluru, Saicharan
Veröffentlicht: (2025)
A Framework for Testing and Adapting REST APIs as LLM Tools
von: Bandlamudi, Jayachandu, et al.
Veröffentlicht: (2025)
von: Bandlamudi, Jayachandu, et al.
Veröffentlicht: (2025)
Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models
von: Syed, Mohammed Sameer, et al.
Veröffentlicht: (2026)
von: Syed, Mohammed Sameer, et al.
Veröffentlicht: (2026)
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2025)
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2025)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
von: Wu, Dekun, et al.
Veröffentlicht: (2023)
von: Wu, Dekun, et al.
Veröffentlicht: (2023)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
von: Topcu, Burak, et al.
Veröffentlicht: (2026)
von: Topcu, Burak, et al.
Veröffentlicht: (2026)
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
von: Shakerdargah, Mohammadali, et al.
Veröffentlicht: (2024)
von: Shakerdargah, Mohammadali, et al.
Veröffentlicht: (2024)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026)
von: Jo, Myeong Jun
Veröffentlicht: (2026)
Enhancing Wide-Angle Image Using Narrow-Angle View of the Same Scene
von: Safwan, Hussain Md., et al.
Veröffentlicht: (2025)
von: Safwan, Hussain Md., et al.
Veröffentlicht: (2025)
When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)
von: Agarwal, Mahak, et al.
Veröffentlicht: (2025)
von: Agarwal, Mahak, et al.
Veröffentlicht: (2025)
Statistical Scouting Finds Debate-Safe but Not Debate-Useful Cases: A Matched-Ceiling Study of Open-Weight LLM Reasoning Protocols
von: Hu, Julia, et al.
Veröffentlicht: (2026)
von: Hu, Julia, et al.
Veröffentlicht: (2026)
Schema as Parameterized Tools for Universal Information Extraction
von: Liang, Sheng, et al.
Veröffentlicht: (2025)
von: Liang, Sheng, et al.
Veröffentlicht: (2025)
Overhead Measurement Noise in Different Runtime Environments
von: Reichelt, David Georg, et al.
Veröffentlicht: (2024)
von: Reichelt, David Georg, et al.
Veröffentlicht: (2024)
ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
MarkLLM: An Open-Source Toolkit for LLM Watermarking
von: Pan, Leyi, et al.
Veröffentlicht: (2024)
von: Pan, Leyi, et al.
Veröffentlicht: (2024)
Exploiting Novel GPT-4 APIs
von: Pelrine, Kellin, et al.
Veröffentlicht: (2023)
von: Pelrine, Kellin, et al.
Veröffentlicht: (2023)
Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models
von: Dahir, Khalid Yusuf
Veröffentlicht: (2026)
von: Dahir, Khalid Yusuf
Veröffentlicht: (2026)
Evaluating LLM Metrics Through Real-World Capabilities
von: Miller, Justin K, et al.
Veröffentlicht: (2025)
von: Miller, Justin K, et al.
Veröffentlicht: (2025)
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
von: Sun, Kangkang, et al.
Veröffentlicht: (2026)
von: Sun, Kangkang, et al.
Veröffentlicht: (2026)
CWM: An Open-Weights LLM for Research on Code Generation with World Models
von: FAIR CodeGen team, et al.
Veröffentlicht: (2025)
von: FAIR CodeGen team, et al.
Veröffentlicht: (2025)
LLM-Based SQL Generation: Prompting, Self-Refinement, and Adaptive Weighted Majority Voting
von: Yang, Yu-Jie, et al.
Veröffentlicht: (2026)
von: Yang, Yu-Jie, et al.
Veröffentlicht: (2026)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
von: Souza, Débora, et al.
Veröffentlicht: (2026)
von: Souza, Débora, et al.
Veröffentlicht: (2026)
NameBERT: Scaling Name-Based Nationality Classification with LLM-Augmented Open Academic Data
von: Ming, Cong, et al.
Veröffentlicht: (2026)
von: Ming, Cong, et al.
Veröffentlicht: (2026)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
von: Liu, Fangxin, et al.
Veröffentlicht: (2025)
von: Liu, Fangxin, et al.
Veröffentlicht: (2025)
Internal APIs Are All You Need: Shadow APIs, Shared Discovery, and the Case Against Browser-First Agent Architectures
von: Tham, Lewis, et al.
Veröffentlicht: (2026)
von: Tham, Lewis, et al.
Veröffentlicht: (2026)
JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR
von: Chen, Xinjie, et al.
Veröffentlicht: (2026)
von: Chen, Xinjie, et al.
Veröffentlicht: (2026)
Unstructured Text Enhanced Open-domain Dialogue System: A Systematic Survey
von: Ma, Longxuan, et al.
Veröffentlicht: (2024)
von: Ma, Longxuan, et al.
Veröffentlicht: (2024)
In-Context Fixation: When Demonstrated Labels Override Semantics in Few-Shot Classification
von: Liu, Ming
Veröffentlicht: (2026)
von: Liu, Ming
Veröffentlicht: (2026)
The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
von: Ma, Xinyue, et al.
Veröffentlicht: (2026) -
Same Voice, Different Lab: On the Homogenization of Frontier LLM Personalities
von: Krishna, Avinash, et al.
Veröffentlicht: (2026) -
Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders
von: Patel, Het, et al.
Veröffentlicht: (2026) -
A Performance Evaluation of a Quantized Large Language Model on Various Smartphones
von: Çöplü, Tolga, et al.
Veröffentlicht: (2023) -
When Are Two RLHF Objectives the Same?
von: Gaikwad, Madhava
Veröffentlicht: (2025)