Tiny but Mighty: A Software-Hardware Co-Design Approach for Efficient Multimodal Inference on Battery-Powered Small Devices
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Yilong, Zhang, Shuai, Zeng, Yijing, Zhang, Hao, Xiong, Xinmiao, Liu, Jingyu, Hu, Pan, Banerjee, Suman |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Benchmarking Compound AI Applications for Hardware-Software Co-Design
por: Samuthrsindh, Paramuth, et al.
Publicado: (2026)
por: Samuthrsindh, Paramuth, et al.
Publicado: (2026)
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
por: Wu, Feiyang, et al.
Publicado: (2025)
por: Wu, Feiyang, et al.
Publicado: (2025)
PICO: Pipeline Inference Framework for Versatile CNNs on Diverse Mobile Devices
por: Yang, Xiang, et al.
Publicado: (2022)
por: Yang, Xiang, et al.
Publicado: (2022)
Distributed On-Device LLM Inference With Over-the-Air Computation
por: Zhang, Kai, et al.
Publicado: (2025)
por: Zhang, Kai, et al.
Publicado: (2025)
Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical
por: Behera, Adarsh Prasad, et al.
Publicado: (2024)
por: Behera, Adarsh Prasad, et al.
Publicado: (2024)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
por: Li, Haley, et al.
Publicado: (2026)
por: Li, Haley, et al.
Publicado: (2026)
Multi-core & GPU-based Balanced Butterfly Counting in Signed Bipartite Graphs
por: Kiran, Mekala, et al.
Publicado: (2026)
por: Kiran, Mekala, et al.
Publicado: (2026)
AdaOper: Energy-efficient and Responsive Concurrent DNN Inference on Mobile Devices
por: Lin, Zheng, et al.
Publicado: (2024)
por: Lin, Zheng, et al.
Publicado: (2024)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
por: Sun, Mingyu, et al.
Publicado: (2025)
por: Sun, Mingyu, et al.
Publicado: (2025)
Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning
por: An, Wei, et al.
Publicado: (2024)
por: An, Wei, et al.
Publicado: (2024)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
por: Wu, Yu, et al.
Publicado: (2025)
por: Wu, Yu, et al.
Publicado: (2025)
Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets
por: Raj, Suman, et al.
Publicado: (2024)
por: Raj, Suman, et al.
Publicado: (2024)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
por: Ahmad, Sohaib, et al.
Publicado: (2024)
por: Ahmad, Sohaib, et al.
Publicado: (2024)
FlexPie: Accelerate Distributed Inference on Edge Devices with Flexible Combinatorial Optimization[Technical Report]
por: Zhang, Runhua, et al.
Publicado: (2025)
por: Zhang, Runhua, et al.
Publicado: (2025)
Pico-Cloud: Cloud Infrastructure for Tiny Edge Devices
por: Guri, Mordechai
Publicado: (2025)
por: Guri, Mordechai
Publicado: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
por: Huang, Jinqi, et al.
Publicado: (2025)
por: Huang, Jinqi, et al.
Publicado: (2025)
Power Aware Dynamic Reallocation For Inference
por: Jiang, Yiwei, et al.
Publicado: (2026)
por: Jiang, Yiwei, et al.
Publicado: (2026)
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
por: Arima, Eishi, et al.
Publicado: (2024)
por: Arima, Eishi, et al.
Publicado: (2024)
CoCoI: Distributed Coded Inference System for Straggler Mitigation
por: Liu, Xing, et al.
Publicado: (2025)
por: Liu, Xing, et al.
Publicado: (2025)
GPUArmor: A Hardware-Software Co-design for Efficient and Scalable Memory Safety on GPUs
por: Ziad, Mohamed Tarek Ibn, et al.
Publicado: (2025)
por: Ziad, Mohamed Tarek Ibn, et al.
Publicado: (2025)
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
por: Tayal, Mumuksh, et al.
Publicado: (2025)
por: Tayal, Mumuksh, et al.
Publicado: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
por: Chow, Will
Publicado: (2025)
por: Chow, Will
Publicado: (2025)
Energy-aware Incremental OTA Update for Flash-based Batteryless IoT Devices
por: Wei, Wei, et al.
Publicado: (2024)
por: Wei, Wei, et al.
Publicado: (2024)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
por: Arya, Mayank, et al.
Publicado: (2025)
por: Arya, Mayank, et al.
Publicado: (2025)
xNVMe: Unleashing Storage Hardware-Software Co-design
por: Lund, Simon A. F., et al.
Publicado: (2024)
por: Lund, Simon A. F., et al.
Publicado: (2024)
Improved Decision Module Selection for Hierarchical Inference in Resource-Constrained Edge Devices
por: Behera, Adarsh Prasad, et al.
Publicado: (2024)
por: Behera, Adarsh Prasad, et al.
Publicado: (2024)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
por: Chen, Jiu, et al.
Publicado: (2026)
por: Chen, Jiu, et al.
Publicado: (2026)
Extracting the Potential of Emerging Hardware Accelerators for Symmetric Eigenvalue Decomposition
por: Wang, Hansheng, et al.
Publicado: (2024)
por: Wang, Hansheng, et al.
Publicado: (2024)
Cicada: A Pipeline-Efficient Approach to Serverless Inference with Decoupled Management
por: Wu, Z., et al.
Publicado: (2025)
por: Wu, Z., et al.
Publicado: (2025)
DynaSplit: A Hardware-Software Co-Design Framework for Energy-Aware Inference on Edge
por: May, Daniel, et al.
Publicado: (2024)
por: May, Daniel, et al.
Publicado: (2024)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
por: Yang, Zheming, et al.
Publicado: (2025)
por: Yang, Zheming, et al.
Publicado: (2025)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
por: Xu, Yaodan, et al.
Publicado: (2025)
por: Xu, Yaodan, et al.
Publicado: (2025)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
por: Kong, Jie, et al.
Publicado: (2026)
por: Kong, Jie, et al.
Publicado: (2026)
Adaptive Device-Edge Collaboration on DNN Inference in AIoT: A Digital Twin-Assisted Approach
por: Hu, Shisheng, et al.
Publicado: (2024)
por: Hu, Shisheng, et al.
Publicado: (2024)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
por: Ma, Bin, et al.
Publicado: (2026)
por: Ma, Bin, et al.
Publicado: (2026)
PIM-SHERPA: Software Method for On-device LLM Inference by Resolving PIM Memory Attribute and Layout Inconsistencies
por: Lee, Sunjung, et al.
Publicado: (2026)
por: Lee, Sunjung, et al.
Publicado: (2026)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
por: Kwak, Hyunseok, et al.
Publicado: (2025)
por: Kwak, Hyunseok, et al.
Publicado: (2025)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
por: Wang, Yingping, et al.
Publicado: (2026)
por: Wang, Yingping, et al.
Publicado: (2026)
Future-Proofing IoT: Unleashing the Power of AWS Greengrass in Propelling Smart Devices to New Heights
por: Kokkula, Sahasra, et al.
Publicado: (2024)
por: Kokkula, Sahasra, et al.
Publicado: (2024)
SimDC: A High-Fidelity Device Simulation Platform for Device-Cloud Collaborative Computing
por: Pei, Ruiguang, et al.
Publicado: (2025)
por: Pei, Ruiguang, et al.
Publicado: (2025)
Ejemplares similares
-
Benchmarking Compound AI Applications for Hardware-Software Co-Design
por: Samuthrsindh, Paramuth, et al.
Publicado: (2026) -
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
por: Wu, Feiyang, et al.
Publicado: (2025) -
PICO: Pipeline Inference Framework for Versatile CNNs on Diverse Mobile Devices
por: Yang, Xiang, et al.
Publicado: (2022) -
Distributed On-Device LLM Inference With Over-the-Air Computation
por: Zhang, Kai, et al.
Publicado: (2025) -
Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical
por: Behera, Adarsh Prasad, et al.
Publicado: (2024)