Trustworthy and Controllable Professional Knowledge Utilization in Large Language Models with TEE-GPU Execution
Fuente:
arXiv
Saved in:
| Main Authors: | Cai, Yifeng, An, Zhida, Meng, Yuhan, Liu, Houqian, Wang, Pengli, Lei, Hanwen, Guo, Yao, Li, Ding |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgenTEE: Confidential LLM Agent Execution on Edge Devices
by: Abdollahi, Sina, et al.
Published: (2026)
by: Abdollahi, Sina, et al.
Published: (2026)
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
by: Song, Yixin, et al.
Published: (2023)
by: Song, Yixin, et al.
Published: (2023)
Efficient Function-as-a-Service for Large Language Models with TIDAL
by: Cui, Weihao, et al.
Published: (2025)
by: Cui, Weihao, et al.
Published: (2025)
NCCLbpf: Verified, Composable Policy Execution for GPU Collective Communication
by: Zheng, Yusheng
Published: (2026)
by: Zheng, Yusheng
Published: (2026)
BYOS: Knowledge-driven Large Language Models Bring Your Own Operating System More Excellent
by: Lin, Hongyu, et al.
Published: (2025)
by: Lin, Hongyu, et al.
Published: (2025)
Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale
by: Cui, Weihao, et al.
Published: (2025)
by: Cui, Weihao, et al.
Published: (2025)
Exploiting Dependency and Parallelism: Real-Time Scheduling and Analysis for GPU Tasks
by: Zhang, Yuanhai, et al.
Published: (2026)
by: Zhang, Yuanhai, et al.
Published: (2026)
Towards Fully-fledged GPU Multitasking via Proactive Memory Scheduling
by: Shen, Weihang, et al.
Published: (2025)
by: Shen, Weihang, et al.
Published: (2025)
Optimizing over FP/EDF Execution Times: Known Results and Open Problems
by: Bini, Enrico
Published: (2024)
by: Bini, Enrico
Published: (2024)
Interference-free Operating System: A 6 Years' Experience in Mitigating Cross-Core Interference in Linux
by: Deng, Zhaomeng, et al.
Published: (2024)
by: Deng, Zhaomeng, et al.
Published: (2024)
Optimizing Logical Execution Time Model for Both Determinism and Low Latency
by: Wang, Sen, et al.
Published: (2023)
by: Wang, Sen, et al.
Published: (2023)
UrgenGo: Urgency-Aware Transparent GPU Kernel Launching for Autonomous Driving
by: Zhu, Hanqi, et al.
Published: (2025)
by: Zhu, Hanqi, et al.
Published: (2025)
Securing Cloud File Systems with Trusted Execution
by: Burke, Quinn, et al.
Published: (2023)
by: Burke, Quinn, et al.
Published: (2023)
MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU
by: Yuan, Zhengqing, et al.
Published: (2026)
by: Yuan, Zhengqing, et al.
Published: (2026)
Skim: Speculative Execution for Fast and Efficient Web Agents
by: Wong, Mike, et al.
Published: (2026)
by: Wong, Mike, et al.
Published: (2026)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
by: Tofigh, Mani, et al.
Published: (2025)
by: Tofigh, Mani, et al.
Published: (2025)
Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
by: Xu, Yuhang, et al.
Published: (2025)
by: Xu, Yuhang, et al.
Published: (2025)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
Configuration Validation with Large Language Models
by: Lian, Xinyu, et al.
Published: (2023)
by: Lian, Xinyu, et al.
Published: (2023)
Characterizing Network Requirements for GPU API Remoting in AI Applications
by: Wang, Tianxia, et al.
Published: (2024)
by: Wang, Tianxia, et al.
Published: (2024)
Guidelines for Building Indexes on Partially Cache-Coherent CXL Shared Memory
by: Wu, Fangnuo, et al.
Published: (2025)
by: Wu, Fangnuo, et al.
Published: (2025)
SemaTune: Semantic-Aware Online OS Tuning with Large Language Models
by: Liargkovas, Georgios, et al.
Published: (2026)
by: Liargkovas, Georgios, et al.
Published: (2026)
An Early Exploration of Deep-Learning-Driven Prefetching for Far Memory
by: Huang, Yutong, et al.
Published: (2025)
by: Huang, Yutong, et al.
Published: (2025)
TierBPF: Page Migration Admission Control for Tiered Memory via eBPF
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
SARA: A Stall-Aware Memory Allocation Strategy for Mixed-Criticality Systems
by: Lee, Meng-Chia, et al.
Published: (2025)
by: Lee, Meng-Chia, et al.
Published: (2025)
RUISA Operational Ecosystem Architecture
by: AL Mohtar, Mouayad
Published: (2026)
by: AL Mohtar, Mouayad
Published: (2026)
Don't Let AI Agents YOLO Your Files: Shifting Information and Control to Filesystems for Agent Safety and Autonomy
by: Zhong, Shawn Wanxiang, et al.
Published: (2026)
by: Zhong, Shawn Wanxiang, et al.
Published: (2026)
MigGPT: Harnessing Large Language Models for Automated Migration of Out-of-Tree Linux Kernel Patches Across Versions
by: Dang, Pucheng, et al.
Published: (2025)
by: Dang, Pucheng, et al.
Published: (2025)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
by: Xing, Jiarong, et al.
Published: (2025)
by: Xing, Jiarong, et al.
Published: (2025)
Peformance Isolation for Inference Processes in Edge GPU Systems
by: Martín, Juan José, et al.
Published: (2026)
by: Martín, Juan José, et al.
Published: (2026)
A Task Equalization Allocation Algorithm Incorporating Blocking Estimation and Resource Similarity Analysis for Vehicle Control Real-Time Systems
by: Duan, Qianlong, et al.
Published: (2025)
by: Duan, Qianlong, et al.
Published: (2025)
Exploiting Application-to-Architecture Dependencies for Designing Scalable OS
by: Xiao, Yao, et al.
Published: (2025)
by: Xiao, Yao, et al.
Published: (2025)
TempoNet: Slack-Quantized Transformer-Guided Reinforcement Scheduler for Adaptive Deadline-Centric Real-Time Dispatchs
by: Fu, Rong, et al.
Published: (2026)
by: Fu, Rong, et al.
Published: (2026)
Performance Isolation and Semantic Determinism in Efficient GPU Spatial Sharing
by: Yang, Zhenyuan, et al.
Published: (2026)
by: Yang, Zhenyuan, et al.
Published: (2026)
An AI Agent Execution Environment to Safeguard User Data
by: Stanley, Robert, et al.
Published: (2026)
by: Stanley, Robert, et al.
Published: (2026)
Dissecting CXL Memory Performance at Scale: Analysis, Modeling, and Optimization
by: Liu, Jinshu, et al.
Published: (2024)
by: Liu, Jinshu, et al.
Published: (2024)
FusionANNS: An Efficient CPU/GPU Cooperative Processing Architecture for Billion-scale Approximate Nearest Neighbor Search
by: Tian, Bing, et al.
Published: (2024)
by: Tian, Bing, et al.
Published: (2024)
GPUOS: A GPU Operating System Primitive for Transparent Operation Fusion
by: Yang, Yiwei, et al.
Published: (2026)
by: Yang, Yiwei, et al.
Published: (2026)
Work-in-Progress: Multi-Deadline DAG Scheduling Model for Autonomous Driving Systems
by: Yano, Atsushi, et al.
Published: (2025)
by: Yano, Atsushi, et al.
Published: (2025)
RTP-LLM: High-Performance Alibaba LLM Inference Engine
by: Tan, Boyu, et al.
Published: (2026)
by: Tan, Boyu, et al.
Published: (2026)
Similar Items
-
AgenTEE: Confidential LLM Agent Execution on Edge Devices
by: Abdollahi, Sina, et al.
Published: (2026) -
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
by: Song, Yixin, et al.
Published: (2023) -
Efficient Function-as-a-Service for Large Language Models with TIDAL
by: Cui, Weihao, et al.
Published: (2025) -
NCCLbpf: Verified, Composable Policy Execution for GPU Collective Communication
by: Zheng, Yusheng
Published: (2026) -
BYOS: Knowledge-driven Large Language Models Bring Your Own Operating System More Excellent
by: Lin, Hongyu, et al.
Published: (2025)