AgentFlux: Decoupled Fine-Tuning & Inference for On-Device Agentic Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Kadekodi, Rohan, Jin, Zhan, Kamahori, Keisuke, Gu, Yile, Khatiri, Sean, Bayindirli, Noah H., Gorbunov, Sergey, Kasikci, Baris |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
by: Gu, Yile, et al.
Published: (2025)
by: Gu, Yile, et al.
Published: (2025)
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
by: Kamahori, Keisuke, et al.
Published: (2024)
by: Kamahori, Keisuke, et al.
Published: (2024)
VoxServe: Streaming-Centric Serving System for Speech Language Models
by: Kamahori, Keisuke, et al.
Published: (2026)
by: Kamahori, Keisuke, et al.
Published: (2026)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
by: Kamahori, Keisuke, et al.
Published: (2026)
by: Kamahori, Keisuke, et al.
Published: (2026)
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
by: Kamahori, Keisuke, et al.
Published: (2025)
by: Kamahori, Keisuke, et al.
Published: (2025)
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
by: Lin, Chien-Yu, et al.
Published: (2025)
by: Lin, Chien-Yu, et al.
Published: (2025)
Jenga: Responsive Tiered Memory Management without Thrashing
by: Kadekodi, Rohan, et al.
Published: (2025)
by: Kadekodi, Rohan, et al.
Published: (2025)
Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMs
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
Scalable and Accurate Application-Level Crash-Consistency Testing via Representative Testing
by: Gu, Yile, et al.
Published: (2025)
by: Gu, Yile, et al.
Published: (2025)
Argos: Agentic Time-Series Anomaly Detection with Autonomous Rule Generation via Large Language Models
by: Gu, Yile, et al.
Published: (2025)
by: Gu, Yile, et al.
Published: (2025)
DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling
by: Pan, Yi, et al.
Published: (2026)
by: Pan, Yi, et al.
Published: (2026)
NanoFlow: Towards Optimal Large Language Model Serving Throughput
by: Zhu, Kan, et al.
Published: (2024)
by: Zhu, Kan, et al.
Published: (2024)
Order isomorphisms in $C^*$-algebras
by: Khatiri, Youssef El
Published: (2026)
by: Khatiri, Youssef El
Published: (2026)
Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
by: Tang, Jiaming, et al.
Published: (2024)
by: Tang, Jiaming, et al.
Published: (2024)
Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
by: Hahm, Dongyoon, et al.
Published: (2025)
by: Hahm, Dongyoon, et al.
Published: (2025)
Quantum Machine Learning for Secondary Frequency Control
by: Jahed, Younes Ghazagh, et al.
Published: (2025)
by: Jahed, Younes Ghazagh, et al.
Published: (2025)
ANCA‐Negative Granulomatosis With Polyangiitis Mimicking Sinusitis and Rhinoscleroma: A Case Report
by: Sergey Gorbunov, et al.
Published: (2026)
by: Sergey Gorbunov, et al.
Published: (2026)
OMG-Agent: Toward Robust Missing Modality Generation with Decoupled Coarse-to-Fine Agentic Workflows
by: Dai, Ruiting, et al.
Published: (2026)
by: Dai, Ruiting, et al.
Published: (2026)
Removing RLHF Protections in GPT-4 via Fine-Tuning
by: Zhan, Qiusi, et al.
Published: (2023)
by: Zhan, Qiusi, et al.
Published: (2023)
Fake Runs, Real Fixes -- Analyzing xPU Performance Through Simulation
by: Zarkadas, Ioannis, et al.
Published: (2025)
by: Zarkadas, Ioannis, et al.
Published: (2025)
Visual Agentic Reinforcement Fine-Tuning
by: Liu, Ziyu, et al.
Published: (2025)
by: Liu, Ziyu, et al.
Published: (2025)
BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching
by: Zhao, Yilong, et al.
Published: (2024)
by: Zhao, Yilong, et al.
Published: (2024)
Ravan: Multi-Head Low-Rank Adaptation for Federated Fine-Tuning
by: Raje, Arian, et al.
Published: (2025)
by: Raje, Arian, et al.
Published: (2025)
Design of an Electrochemical Device for the Detection of Alkaline Phosphatase Inhibitors in Seawater
by: María J. Sáenz‐Espinar, et al.
Published: (2024)
by: María J. Sáenz‐Espinar, et al.
Published: (2024)
Gendering starvation: Women's experiences of the Kazakh famine, 1930–1933
by: Mehmet Volkan Kaşıkçı
Published: (2024)
by: Mehmet Volkan Kaşıkçı
Published: (2024)
Bridging Structural Causal Inference and Machine Learning The S-DIDML Estimator for Heterogeneous Treatment Effects
by: Yu, Yile, et al.
Published: (2025)
by: Yu, Yile, et al.
Published: (2025)
Initialization using Update Approximation is a Silver Bullet for Extremely Efficient Low-Rank Fine-Tuning
by: Ponkshe, Kaustubh, et al.
Published: (2024)
by: Ponkshe, Kaustubh, et al.
Published: (2024)
Decoupled Training with Local Reinforcement Fine-Tuning in Federated Learning
by: Ma, Yuting, et al.
Published: (2026)
by: Ma, Yuting, et al.
Published: (2026)
gigiProfiler: Diagnosing Performance Issues by Uncovering Application Resource Bottlenecks
by: Hu, Yigong, et al.
Published: (2025)
by: Hu, Yigong, et al.
Published: (2025)
Magneton: Optimizing Energy Efficiency of ML Systems via Differential Energy Debugging
by: Pan, Yi, et al.
Published: (2025)
by: Pan, Yi, et al.
Published: (2025)
Don't Just Fine-tune the Agent, Tune the Environment
by: Lu, Siyuan, et al.
Published: (2025)
by: Lu, Siyuan, et al.
Published: (2025)
Decoupled Multi-Predictor Optimization for Inference-Efficient Model Tuning
by: Luo, Liwei, et al.
Published: (2025)
by: Luo, Liwei, et al.
Published: (2025)
Manticore: Hardware-Accelerated RTL Simulation with Static Bulk-Synchronous Parallelism
by: Emami, Mahyar, et al.
Published: (2023)
by: Emami, Mahyar, et al.
Published: (2023)
Kraken: Inherently Parallel Transformers For Efficient Multi-Device Inference
by: Prabhakar, Rohan Baskar, et al.
Published: (2024)
by: Prabhakar, Rohan Baskar, et al.
Published: (2024)
Assessment of Generative Named Entity Recognition in the Era of Large Language Models
by: Zhan, Qi, et al.
Published: (2026)
by: Zhan, Qi, et al.
Published: (2026)
Neural Network Optimization Reimagined: Decoupled Techniques for Scratch and Fine-Tuning
by: Ning, Xin, et al.
Published: (2026)
by: Ning, Xin, et al.
Published: (2026)
Relating Wigner's Friend Scenarios to Nonclassical Causal Compatibility, Monogamy Relations, and Fine Tuning
by: Yīng, Yìlè, et al.
Published: (2023)
by: Yīng, Yìlè, et al.
Published: (2023)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
by: Ye, Zihao, et al.
Published: (2025)
by: Ye, Zihao, et al.
Published: (2025)
Upshot of Some Bioactive Compounds on Angiogenesis in Retinal Pigment Epithelial Cells
by: Serkan Sen, et al.
Published: (2025)
by: Serkan Sen, et al.
Published: (2025)
XBRLTagRec: Domain-Specific Fine-Tuning and Zero-Shot Re-Ranking with LLMs for Extreme Financial Numeral Labeling
by: Hu, Gang, et al.
Published: (2026)
by: Hu, Gang, et al.
Published: (2026)
Similar Items
-
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
by: Gu, Yile, et al.
Published: (2025) -
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
by: Kamahori, Keisuke, et al.
Published: (2024) -
VoxServe: Streaming-Centric Serving System for Speech Language Models
by: Kamahori, Keisuke, et al.
Published: (2026) -
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
by: Kamahori, Keisuke, et al.
Published: (2026) -
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
by: Kamahori, Keisuke, et al.
Published: (2025)