Low-resource domain adaptation while minimizing energy and hardware resource consumption
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Maina, Hernán, Wolovick, Nicolás, Benotti, Luciana |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards cost-effective and resource-aware aggregation at Edge for Federated Learning
von: Khan, Ahmad Faraz, et al.
Veröffentlicht: (2022)
von: Khan, Ahmad Faraz, et al.
Veröffentlicht: (2022)
Salted Inference: Enhancing Privacy while Maintaining Efficiency of Split Inference in Mobile Computing
von: Malekzadeh, Mohammad, et al.
Veröffentlicht: (2023)
von: Malekzadeh, Mohammad, et al.
Veröffentlicht: (2023)
Bridging Emotions and Architecture: Sentiment Analysis in Modern Distributed Systems
von: Shah, Mahak, et al.
Veröffentlicht: (2025)
von: Shah, Mahak, et al.
Veröffentlicht: (2025)
Ladder-residual: parallelism-aware architecture for accelerating large model inference with communication overlapping
von: Zhang, Muru, et al.
Veröffentlicht: (2025)
von: Zhang, Muru, et al.
Veröffentlicht: (2025)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
von: Li, Wenxuan, et al.
Veröffentlicht: (2025)
von: Li, Wenxuan, et al.
Veröffentlicht: (2025)
Unlocking Full Efficiency of Token Filtering in Large Language Model Training
von: Chai, Di, et al.
Veröffentlicht: (2025)
von: Chai, Di, et al.
Veröffentlicht: (2025)
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
von: Hong, Ke, et al.
Veröffentlicht: (2025)
von: Hong, Ke, et al.
Veröffentlicht: (2025)
CAFL-L: Constraint-Aware Federated Learning with Lagrangian Dual Optimization for On-Device Language Models
von: Zheng, Dongqi, et al.
Veröffentlicht: (2025)
von: Zheng, Dongqi, et al.
Veröffentlicht: (2025)
Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
von: Charles, Zachary, et al.
Veröffentlicht: (2025)
von: Charles, Zachary, et al.
Veröffentlicht: (2025)
$K^4$: Online Log Anomaly Detection Via Unsupervised Typicality Learning
von: Chen, Weicong, et al.
Veröffentlicht: (2025)
von: Chen, Weicong, et al.
Veröffentlicht: (2025)
Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding
von: Jin, Tian, et al.
Veröffentlicht: (2025)
von: Jin, Tian, et al.
Veröffentlicht: (2025)
Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models
von: Pan, Xuchen, et al.
Veröffentlicht: (2025)
von: Pan, Xuchen, et al.
Veröffentlicht: (2025)
Alchemist: Towards the Design of Efficient Online Continual Learning System
von: Huang, Yuyang, et al.
Veröffentlicht: (2025)
von: Huang, Yuyang, et al.
Veröffentlicht: (2025)
P/D-Device: Disaggregated Large Language Model between Cloud and Devices
von: Jin, Yibo, et al.
Veröffentlicht: (2025)
von: Jin, Yibo, et al.
Veröffentlicht: (2025)
X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
von: Yuan, Yueming, et al.
Veröffentlicht: (2025)
von: Yuan, Yueming, et al.
Veröffentlicht: (2025)
Efficient and Adaptable Overlapping for Computation and Communication via Signaling and Reordering
von: Hong, Ke, et al.
Veröffentlicht: (2025)
von: Hong, Ke, et al.
Veröffentlicht: (2025)
Optimizing RLHF Training for Large Language Models with Stage Fusion
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
Re-evaluating the Memory-balanced Pipeline Parallelism: BPipe
von: Huang, Mincong, et al.
Veröffentlicht: (2024)
von: Huang, Mincong, et al.
Veröffentlicht: (2024)
Federated Learning of Large Language Models with Parameter-Efficient Prompt Tuning and Adaptive Optimization
von: Che, Tianshi, et al.
Veröffentlicht: (2023)
von: Che, Tianshi, et al.
Veröffentlicht: (2023)
Towards Resiliency in Large Language Model Serving with KevlarFlow
von: Qian, Shangshu, et al.
Veröffentlicht: (2026)
von: Qian, Shangshu, et al.
Veröffentlicht: (2026)
Scalable Training of Mixture-of-Experts Models with Megatron Core
von: Yan, Zijie, et al.
Veröffentlicht: (2026)
von: Yan, Zijie, et al.
Veröffentlicht: (2026)
JORA: JAX Tensor-Parallel LoRA Library for Retrieval Augmented Fine-Tuning
von: Tahir, Anique, et al.
Veröffentlicht: (2024)
von: Tahir, Anique, et al.
Veröffentlicht: (2024)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
Fast Matrix Multiplications for Lookup Table-Quantized LLMs
von: Guo, Han, et al.
Veröffentlicht: (2024)
von: Guo, Han, et al.
Veröffentlicht: (2024)
Towards Federated RLHF with Aggregated Client Preference for LLMs
von: Wu, Feijie, et al.
Veröffentlicht: (2024)
von: Wu, Feijie, et al.
Veröffentlicht: (2024)
LLM-Pilot: Characterize and Optimize Performance of your LLM Inference Services
von: Łazuka, Małgorzata, et al.
Veröffentlicht: (2024)
von: Łazuka, Małgorzata, et al.
Veröffentlicht: (2024)
InferCept: Efficient Intercept Support for Augmented Large Language Model Inference
von: Abhyankar, Reyna, et al.
Veröffentlicht: (2024)
von: Abhyankar, Reyna, et al.
Veröffentlicht: (2024)
SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models
von: Lin, Zheng, et al.
Veröffentlicht: (2024)
von: Lin, Zheng, et al.
Veröffentlicht: (2024)
PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
von: Butler, Branden, et al.
Veröffentlicht: (2024)
von: Butler, Branden, et al.
Veröffentlicht: (2024)
Conveyor: Efficient Tool-aware LLM Serving with Tool Partial Execution
von: Xu, Yechen, et al.
Veröffentlicht: (2024)
von: Xu, Yechen, et al.
Veröffentlicht: (2024)
Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow
von: Mei, Yixuan, et al.
Veröffentlicht: (2024)
von: Mei, Yixuan, et al.
Veröffentlicht: (2024)
Optimizing Cross-Client Domain Coverage for Federated Instruction Tuning of Large Language Models
von: Wang, Zezhou, et al.
Veröffentlicht: (2024)
von: Wang, Zezhou, et al.
Veröffentlicht: (2024)
Queue management for slo-oriented large language model serving
von: Patke, Archit, et al.
Veröffentlicht: (2024)
von: Patke, Archit, et al.
Veröffentlicht: (2024)
PAAC: Privacy-Aware Agentic Device-Cloud Collaboration
von: Yuan, Liangqi, et al.
Veröffentlicht: (2026)
von: Yuan, Liangqi, et al.
Veröffentlicht: (2026)
Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction
von: Qiu, Haoran, et al.
Veröffentlicht: (2024)
von: Qiu, Haoran, et al.
Veröffentlicht: (2024)
Pipeline Parallelism with Controllable Memory
von: Qi, Penghui, et al.
Veröffentlicht: (2024)
von: Qi, Penghui, et al.
Veröffentlicht: (2024)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
von: Jin, Yibo, et al.
Veröffentlicht: (2024)
von: Jin, Yibo, et al.
Veröffentlicht: (2024)
FedBiOT: LLM Local Fine-tuning in Federated Learning without Full Model
von: Wu, Feijie, et al.
Veröffentlicht: (2024)
von: Wu, Feijie, et al.
Veröffentlicht: (2024)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards cost-effective and resource-aware aggregation at Edge for Federated Learning
von: Khan, Ahmad Faraz, et al.
Veröffentlicht: (2022) -
Salted Inference: Enhancing Privacy while Maintaining Efficiency of Split Inference in Mobile Computing
von: Malekzadeh, Mohammad, et al.
Veröffentlicht: (2023) -
Bridging Emotions and Architecture: Sentiment Analysis in Modern Distributed Systems
von: Shah, Mahak, et al.
Veröffentlicht: (2025) -
Ladder-residual: parallelism-aware architecture for accelerating large model inference with communication overlapping
von: Zhang, Muru, et al.
Veröffentlicht: (2025) -
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
von: Li, Wenxuan, et al.
Veröffentlicht: (2025)