Efficiently Deploying LLMs with Controlled Risk
Fuente:
arXiv
Saved in:
| Main Authors: | Zellinger, Michael J., Thomson, Matt |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fail Fast, or Ask: Mitigating the Deficiencies of Reasoning LLMs with Human-in-the-Loop Systems Engineering
by: Zellinger, Michael J., et al.
Published: (2025)
by: Zellinger, Michael J., et al.
Published: (2025)
Rational Tuning of LLM Cascades via Probabilistic Modeling
by: Zellinger, Michael J., et al.
Published: (2025)
by: Zellinger, Michael J., et al.
Published: (2025)
Economic Evaluation of LLMs
by: Zellinger, Michael J., et al.
Published: (2025)
by: Zellinger, Michael J., et al.
Published: (2025)
Cost-Saving LLM Cascades with Early Abstention
by: Zellinger, Michael J., et al.
Published: (2025)
by: Zellinger, Michael J., et al.
Published: (2025)
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
by: Griffin, Charlie, et al.
Published: (2024)
by: Griffin, Charlie, et al.
Published: (2024)
Learning with Noisy Labels by Adaptive Gradient-Based Outlier Removal
by: Sedova, Anastasiia, et al.
Published: (2023)
by: Sedova, Anastasiia, et al.
Published: (2023)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
by: Zhang, Tuo, et al.
Published: (2025)
by: Zhang, Tuo, et al.
Published: (2025)
Counterfactual Reasoning with Knowledge Graph Embeddings
by: Zellinger, Lena, et al.
Published: (2024)
by: Zellinger, Lena, et al.
Published: (2024)
Near-Optimal Online Deployment and Routing for Streaming LLMs
by: Li, Shaoang, et al.
Published: (2025)
by: Li, Shaoang, et al.
Published: (2025)
Empirical Guidelines for Deploying LLMs onto Resource-constrained Edge Devices
by: Qin, Ruiyang, et al.
Published: (2024)
by: Qin, Ruiyang, et al.
Published: (2024)
Prompt Risk Control: A Rigorous Framework for Responsible Deployment of Large Language Models
by: Zollo, Thomas P., et al.
Published: (2023)
by: Zollo, Thomas P., et al.
Published: (2023)
Automatically Adaptive Conformal Risk Control
by: Blot, Vincent, et al.
Published: (2024)
by: Blot, Vincent, et al.
Published: (2024)
Conformal Selective Acting: Anytime-Valid Risk Control for RLVR-Trained LLMs
by: Khosravi, Hamed, et al.
Published: (2026)
by: Khosravi, Hamed, et al.
Published: (2026)
Iterative Deployment Improves Planning Skills in LLMs
by: Corrêa, Augusto B., et al.
Published: (2025)
by: Corrêa, Augusto B., et al.
Published: (2025)
What's the Magic Word? A Control Theory of LLM Prompting
by: Bhargava, Aman, et al.
Published: (2023)
by: Bhargava, Aman, et al.
Published: (2023)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
by: Turk, Matt
Published: (2026)
by: Turk, Matt
Published: (2026)
Risk Profiling and Modulation for LLMs
by: Wang, Yikai, et al.
Published: (2025)
by: Wang, Yikai, et al.
Published: (2025)
Resource-Efficient Generative AI Model Deployment in Mobile Edge Networks
by: Liang, Yuxin, et al.
Published: (2024)
by: Liang, Yuxin, et al.
Published: (2024)
Q-realign: Piggybacking Realignment on Quantization for Safe and Efficient LLM Deployment
by: Tan, Qitao, et al.
Published: (2026)
by: Tan, Qitao, et al.
Published: (2026)
Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
by: Taraghi, Mina, et al.
Published: (2025)
by: Taraghi, Mina, et al.
Published: (2025)
SymLight: Exploring Interpretable and Deployable Symbolic Policies for Traffic Signal Control
by: Liao, Xiao-Cheng, et al.
Published: (2025)
by: Liao, Xiao-Cheng, et al.
Published: (2025)
Reinforcing privacy reasoning in LLMs via normative simulacra from fiction
by: Franchi, Matt, et al.
Published: (2026)
by: Franchi, Matt, et al.
Published: (2026)
Co-Designing Binarized Transformer and Hardware Accelerator for Efficient End-to-End Edge Deployment
by: Ji, Yuhao, et al.
Published: (2024)
by: Ji, Yuhao, et al.
Published: (2024)
From Algorithm to Hardware: A Survey on Efficient and Safe Deployment of Deep Neural Networks
by: Geng, Xue, et al.
Published: (2024)
by: Geng, Xue, et al.
Published: (2024)
AmoebaLLM: Constructing Any-Shape Large Language Models for Efficient and Instant Deployment
by: Fu, Yonggan, et al.
Published: (2024)
by: Fu, Yonggan, et al.
Published: (2024)
Post-Training Quantization of OpenPangu Models for Efficient Deployment on Atlas A2
by: Luo, Yilun, et al.
Published: (2025)
by: Luo, Yilun, et al.
Published: (2025)
Performance Control in Early Exiting to Deploy Large Models at the Same Cost of Smaller Ones
by: Mofakhami, Mehrnaz, et al.
Published: (2024)
by: Mofakhami, Mehrnaz, et al.
Published: (2024)
DPF-CM: A Data Processing Framework with Privacy-Preserving Vector Databases for Chinese Medical LLMs Training and Deployment
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
by: Vulićević, Jelena Ilić
Published: (2026)
by: Vulićević, Jelena Ilić
Published: (2026)
SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment
by: Song, Yixin, et al.
Published: (2025)
by: Song, Yixin, et al.
Published: (2025)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
by: Lee, Deokjae, et al.
Published: (2025)
by: Lee, Deokjae, et al.
Published: (2025)
Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts?
by: Karim, Aabid, et al.
Published: (2025)
by: Karim, Aabid, et al.
Published: (2025)
Unified Parameter-Efficient Unlearning for LLMs
by: Ding, Chenlu, et al.
Published: (2024)
by: Ding, Chenlu, et al.
Published: (2024)
Selective Conformal Risk Control
by: Xu, Yunpeng, et al.
Published: (2025)
by: Xu, Yunpeng, et al.
Published: (2025)
Estimating Worst-Case Frontier Risks of Open-Weight LLMs
by: Wallace, Eric, et al.
Published: (2025)
by: Wallace, Eric, et al.
Published: (2025)
Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs
by: Huang, Luke J., et al.
Published: (2026)
by: Huang, Luke J., et al.
Published: (2026)
Rapid Deployment of DNNs for Edge Computing via Structured Pruning at Initialization
by: Eccles, Bailey J., et al.
Published: (2024)
by: Eccles, Bailey J., et al.
Published: (2024)
Designing and Deploying AI Models for Sustainable Logistics Optimization: A Case Study on Eco-Efficient Supply Chains in the USA
by: Shawon, Reza E Rabbi, et al.
Published: (2025)
by: Shawon, Reza E Rabbi, et al.
Published: (2025)
EL-MIA: Quantifying Membership Inference Risks of Sensitive Entities in LLMs
by: Satvaty, Ali, et al.
Published: (2025)
by: Satvaty, Ali, et al.
Published: (2025)
Frictive Policy Optimization for LLMs: Epistemic Intervention, Risk-Sensitive Control, and Reflective Alignment
by: Pustejovsky, James, et al.
Published: (2026)
by: Pustejovsky, James, et al.
Published: (2026)
Similar Items
-
Fail Fast, or Ask: Mitigating the Deficiencies of Reasoning LLMs with Human-in-the-Loop Systems Engineering
by: Zellinger, Michael J., et al.
Published: (2025) -
Rational Tuning of LLM Cascades via Probabilistic Modeling
by: Zellinger, Michael J., et al.
Published: (2025) -
Economic Evaluation of LLMs
by: Zellinger, Michael J., et al.
Published: (2025) -
Cost-Saving LLM Cascades with Early Abstention
by: Zellinger, Michael J., et al.
Published: (2025) -
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
by: Griffin, Charlie, et al.
Published: (2024)