Tiny but Mighty: A Software-Hardware Co-Design Approach for Efficient Multimodal Inference on Battery-Powered Small Devices
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Yilong, Zhang, Shuai, Zeng, Yijing, Zhang, Hao, Xiong, Xinmiao, Liu, Jingyu, Hu, Pan, Banerjee, Suman |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Benchmarking Compound AI Applications for Hardware-Software Co-Design
par: Samuthrsindh, Paramuth, et autres
Publié: (2026)
par: Samuthrsindh, Paramuth, et autres
Publié: (2026)
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
par: Wu, Feiyang, et autres
Publié: (2025)
par: Wu, Feiyang, et autres
Publié: (2025)
PICO: Pipeline Inference Framework for Versatile CNNs on Diverse Mobile Devices
par: Yang, Xiang, et autres
Publié: (2022)
par: Yang, Xiang, et autres
Publié: (2022)
Distributed On-Device LLM Inference With Over-the-Air Computation
par: Zhang, Kai, et autres
Publié: (2025)
par: Zhang, Kai, et autres
Publié: (2025)
Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical
par: Behera, Adarsh Prasad, et autres
Publié: (2024)
par: Behera, Adarsh Prasad, et autres
Publié: (2024)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
par: Li, Haley, et autres
Publié: (2026)
par: Li, Haley, et autres
Publié: (2026)
Multi-core & GPU-based Balanced Butterfly Counting in Signed Bipartite Graphs
par: Kiran, Mekala, et autres
Publié: (2026)
par: Kiran, Mekala, et autres
Publié: (2026)
AdaOper: Energy-efficient and Responsive Concurrent DNN Inference on Mobile Devices
par: Lin, Zheng, et autres
Publié: (2024)
par: Lin, Zheng, et autres
Publié: (2024)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
par: Sun, Mingyu, et autres
Publié: (2025)
par: Sun, Mingyu, et autres
Publié: (2025)
Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning
par: An, Wei, et autres
Publié: (2024)
par: An, Wei, et autres
Publié: (2024)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
par: Wu, Yu, et autres
Publié: (2025)
par: Wu, Yu, et autres
Publié: (2025)
Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets
par: Raj, Suman, et autres
Publié: (2024)
par: Raj, Suman, et autres
Publié: (2024)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
par: Ahmad, Sohaib, et autres
Publié: (2024)
par: Ahmad, Sohaib, et autres
Publié: (2024)
FlexPie: Accelerate Distributed Inference on Edge Devices with Flexible Combinatorial Optimization[Technical Report]
par: Zhang, Runhua, et autres
Publié: (2025)
par: Zhang, Runhua, et autres
Publié: (2025)
Pico-Cloud: Cloud Infrastructure for Tiny Edge Devices
par: Guri, Mordechai
Publié: (2025)
par: Guri, Mordechai
Publié: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
par: Huang, Jinqi, et autres
Publié: (2025)
par: Huang, Jinqi, et autres
Publié: (2025)
Power Aware Dynamic Reallocation For Inference
par: Jiang, Yiwei, et autres
Publié: (2026)
par: Jiang, Yiwei, et autres
Publié: (2026)
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
par: Arima, Eishi, et autres
Publié: (2024)
par: Arima, Eishi, et autres
Publié: (2024)
CoCoI: Distributed Coded Inference System for Straggler Mitigation
par: Liu, Xing, et autres
Publié: (2025)
par: Liu, Xing, et autres
Publié: (2025)
GPUArmor: A Hardware-Software Co-design for Efficient and Scalable Memory Safety on GPUs
par: Ziad, Mohamed Tarek Ibn, et autres
Publié: (2025)
par: Ziad, Mohamed Tarek Ibn, et autres
Publié: (2025)
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
par: Tayal, Mumuksh, et autres
Publié: (2025)
par: Tayal, Mumuksh, et autres
Publié: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
par: Chow, Will
Publié: (2025)
par: Chow, Will
Publié: (2025)
Energy-aware Incremental OTA Update for Flash-based Batteryless IoT Devices
par: Wei, Wei, et autres
Publié: (2024)
par: Wei, Wei, et autres
Publié: (2024)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
par: Arya, Mayank, et autres
Publié: (2025)
par: Arya, Mayank, et autres
Publié: (2025)
xNVMe: Unleashing Storage Hardware-Software Co-design
par: Lund, Simon A. F., et autres
Publié: (2024)
par: Lund, Simon A. F., et autres
Publié: (2024)
Improved Decision Module Selection for Hierarchical Inference in Resource-Constrained Edge Devices
par: Behera, Adarsh Prasad, et autres
Publié: (2024)
par: Behera, Adarsh Prasad, et autres
Publié: (2024)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
par: Chen, Jiu, et autres
Publié: (2026)
par: Chen, Jiu, et autres
Publié: (2026)
Extracting the Potential of Emerging Hardware Accelerators for Symmetric Eigenvalue Decomposition
par: Wang, Hansheng, et autres
Publié: (2024)
par: Wang, Hansheng, et autres
Publié: (2024)
Cicada: A Pipeline-Efficient Approach to Serverless Inference with Decoupled Management
par: Wu, Z., et autres
Publié: (2025)
par: Wu, Z., et autres
Publié: (2025)
DynaSplit: A Hardware-Software Co-Design Framework for Energy-Aware Inference on Edge
par: May, Daniel, et autres
Publié: (2024)
par: May, Daniel, et autres
Publié: (2024)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
par: Yang, Zheming, et autres
Publié: (2025)
par: Yang, Zheming, et autres
Publié: (2025)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
par: Xu, Yaodan, et autres
Publié: (2025)
par: Xu, Yaodan, et autres
Publié: (2025)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
par: Kong, Jie, et autres
Publié: (2026)
par: Kong, Jie, et autres
Publié: (2026)
Adaptive Device-Edge Collaboration on DNN Inference in AIoT: A Digital Twin-Assisted Approach
par: Hu, Shisheng, et autres
Publié: (2024)
par: Hu, Shisheng, et autres
Publié: (2024)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
par: Ma, Bin, et autres
Publié: (2026)
par: Ma, Bin, et autres
Publié: (2026)
PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies
par: Lee, Sunjung, et autres
Publié: (2026)
par: Lee, Sunjung, et autres
Publié: (2026)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
par: Kwak, Hyunseok, et autres
Publié: (2025)
par: Kwak, Hyunseok, et autres
Publié: (2025)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
par: Wang, Yingping, et autres
Publié: (2026)
par: Wang, Yingping, et autres
Publié: (2026)
Future-Proofing IoT: Unleashing the Power of AWS Greengrass in Propelling Smart Devices to New Heights
par: Kokkula, Sahasra, et autres
Publié: (2024)
par: Kokkula, Sahasra, et autres
Publié: (2024)
SimDC: A High-Fidelity Device Simulation Platform for Device-Cloud Collaborative Computing
par: Pei, Ruiguang, et autres
Publié: (2025)
par: Pei, Ruiguang, et autres
Publié: (2025)
Documents similaires
-
Benchmarking Compound AI Applications for Hardware-Software Co-Design
par: Samuthrsindh, Paramuth, et autres
Publié: (2026) -
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
par: Wu, Feiyang, et autres
Publié: (2025) -
PICO: Pipeline Inference Framework for Versatile CNNs on Diverse Mobile Devices
par: Yang, Xiang, et autres
Publié: (2022) -
Distributed On-Device LLM Inference With Over-the-Air Computation
par: Zhang, Kai, et autres
Publié: (2025) -
Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical
par: Behera, Adarsh Prasad, et autres
Publié: (2024)