Benchmarking Edge AI Platforms for High-Performance ML Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Jayanth, Rakshith, Gupta, Neelesh, Prasanna, Viktor |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
por: Gupta, Neelesh, et al.
Publicado: (2025)
por: Gupta, Neelesh, et al.
Publicado: (2025)
FAST-Prefill: FPGA Accelerated Sparse Attention for Long Context LLM Prefill
por: Jayanth, Rakshith, et al.
Publicado: (2026)
por: Jayanth, Rakshith, et al.
Publicado: (2026)
AI/ML in 3GPP 5G Advanced -- Services and Architecture
por: Taksande, Pradnya, et al.
Publicado: (2025)
por: Taksande, Pradnya, et al.
Publicado: (2025)
TabConv: Low-Computation CNN Inference via Table Lookups
por: Gupta, Neelesh, et al.
Publicado: (2024)
por: Gupta, Neelesh, et al.
Publicado: (2024)
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare
por: Desikan, Prasanna, et al.
Publicado: (2026)
por: Desikan, Prasanna, et al.
Publicado: (2026)
Enabling Long FFT Convolutions on Memory-Constrained FPGAs via Chunking
por: Wang, Peter, et al.
Publicado: (2025)
por: Wang, Peter, et al.
Publicado: (2025)
On the Sustainability of AI Inferences in the Edge
por: Sobhani, Ghazal, et al.
Publicado: (2025)
por: Sobhani, Ghazal, et al.
Publicado: (2025)
Embedded AI Companion System on Edge Devices
por: Gupta, Rahul, et al.
Publicado: (2026)
por: Gupta, Rahul, et al.
Publicado: (2026)
ML Research Benchmark
por: Kenney, Matthew
Publicado: (2024)
por: Kenney, Matthew
Publicado: (2024)
Beyond Benchmarks: The Economics of AI Inference
por: Zhuang, Boqin, et al.
Publicado: (2025)
por: Zhuang, Boqin, et al.
Publicado: (2025)
The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization
por: Chung, Jae-Won, et al.
Publicado: (2025)
por: Chung, Jae-Won, et al.
Publicado: (2025)
AnveshanaAI: A Multimodal Platform for Adaptive AI/ML Education through Automated Question Generation and Interactive Assessment
por: Thakur, Rakesh, et al.
Publicado: (2025)
por: Thakur, Rakesh, et al.
Publicado: (2025)
Dora: QoE-Aware Hybrid Parallelism for Distributed Edge AI
por: Jin, Jianli, et al.
Publicado: (2025)
por: Jin, Jianli, et al.
Publicado: (2025)
Model Recovery at the Edge under Resource Constraints for Physical AI
por: Xu, Bin, et al.
Publicado: (2025)
por: Xu, Bin, et al.
Publicado: (2025)
Using AI/ML to Find and Remediate Enterprise Secrets in Code & Document Sharing Platforms
por: Kerr, Gregor, et al.
Publicado: (2024)
por: Kerr, Gregor, et al.
Publicado: (2024)
JobMatchAI An Intelligent Job Matching Platform Using Knowledge Graphs, Semantic Search and Explainable AI
por: Vyas, Mayank, et al.
Publicado: (2026)
por: Vyas, Mayank, et al.
Publicado: (2026)
AI Safeguards, Generative AI and the Pandora Box: AI Safety Measures to Protect Businesses and Personal Reputation
por: Kumar, Prasanna
Publicado: (2026)
por: Kumar, Prasanna
Publicado: (2026)
TypeBandit: Type-Level Context Allocation and Reweighting for Effective Attribute Completion in Heterogeneous Graph Neural Networks
por: Wang, Ta-Yang, et al.
Publicado: (2026)
por: Wang, Ta-Yang, et al.
Publicado: (2026)
TIGER-MARL: Enhancing Multi-Agent Reinforcement Learning with Temporal Information through Graph-based Embeddings and Representations
por: Gupta, Nikunj, et al.
Publicado: (2025)
por: Gupta, Nikunj, et al.
Publicado: (2025)
LoCoML: A Framework for Real-World ML Inference Pipelines
por: Maddireddy, Kritin, et al.
Publicado: (2025)
por: Maddireddy, Kritin, et al.
Publicado: (2025)
Intra-DP: A High Performance Collaborative Inference System for Mobile Edge Computing
por: Sun, Zekai, et al.
Publicado: (2025)
por: Sun, Zekai, et al.
Publicado: (2025)
Effective ML Model Versioning in Edge Networks
por: Gentzen, Fin, et al.
Publicado: (2024)
por: Gentzen, Fin, et al.
Publicado: (2024)
AI-Powered Annotation Pipelines for Stabilizing Large Language Models: A Human-AI Synergy Approach
por: Pathak, Gangesh, et al.
Publicado: (2025)
por: Pathak, Gangesh, et al.
Publicado: (2025)
Towards Transparent AI Grading: Semantic Entropy as a Signal for Human-AI Disagreement
por: Iyer, Karrtik, et al.
Publicado: (2025)
por: Iyer, Karrtik, et al.
Publicado: (2025)
An Edge AI System Based on FPGA Platform for Railway Fault Detection
por: Li, Jiale, et al.
Publicado: (2024)
por: Li, Jiale, et al.
Publicado: (2024)
HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI
por: Pricope, Tidor-Vlad
Publicado: (2025)
por: Pricope, Tidor-Vlad
Publicado: (2025)
HiDP: Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge Platforms
por: Taufique, Zain, et al.
Publicado: (2024)
por: Taufique, Zain, et al.
Publicado: (2024)
The Dark Side of AI Transformers: Sentiment Polarization & the Loss of Business Neutrality by NLP Transformers
por: Kumar, Prasanna
Publicado: (2026)
por: Kumar, Prasanna
Publicado: (2026)
Intanify AI Platform: Embedded AI for Automated IP Audit and Due Diligence
por: Dorfler, Viktor, et al.
Publicado: (2025)
por: Dorfler, Viktor, et al.
Publicado: (2025)
ShrutiSense: Microtonal Modeling and Correction in Indian Classical Music
por: Ghosh, Rajarshi, et al.
Publicado: (2025)
por: Ghosh, Rajarshi, et al.
Publicado: (2025)
PaCKD: Pattern-Clustered Knowledge Distillation for Compressing Memory Access Prediction Models
por: Gupta, Neelesh, et al.
Publicado: (2024)
por: Gupta, Neelesh, et al.
Publicado: (2024)
A Hybrid Edge Classifier: Combining TinyML-Optimised CNN with RRAM-CMOS ACAM for Energy-Efficient Inference
por: Woodward, Kieran, et al.
Publicado: (2025)
por: Woodward, Kieran, et al.
Publicado: (2025)
MHDash: An Online Platform for Benchmarking Mental Health-Aware AI Assistants
por: Zhang, Yihe, et al.
Publicado: (2026)
por: Zhang, Yihe, et al.
Publicado: (2026)
FlashMoE: Reducing SSD I/O Bottlenecks via ML-Based Cache Replacement for Mixture-of-Experts Inference on Edge Devices
por: Kim, Byeongju, et al.
Publicado: (2026)
por: Kim, Byeongju, et al.
Publicado: (2026)
Scalable AI Inference: Performance Analysis and Optimization of AI Model Serving
por: Pham, Hung Cuong, et al.
Publicado: (2026)
por: Pham, Hung Cuong, et al.
Publicado: (2026)
Collaborative Edge AI Inference over Cloud-RAN
por: Zhang, Pengfei, et al.
Publicado: (2024)
por: Zhang, Pengfei, et al.
Publicado: (2024)
Enabling Physical AI at the Edge: Hardware-Accelerated Recovery of System Dynamics
por: Xu, Bin, et al.
Publicado: (2025)
por: Xu, Bin, et al.
Publicado: (2025)
Agentic Educational Content Generation for African Languages on Edge Devices
por: Gupta, Ravi, et al.
Publicado: (2025)
por: Gupta, Ravi, et al.
Publicado: (2025)
oneDAL Optimization for ARM Scalable Vector Extension: Maximizing Efficiency for High-Performance Data Science
por: Sharma, Chandan, et al.
Publicado: (2025)
por: Sharma, Chandan, et al.
Publicado: (2025)
ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows
por: Padigela, Harshith, et al.
Publicado: (2025)
por: Padigela, Harshith, et al.
Publicado: (2025)
Ejemplares similares
-
Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
por: Gupta, Neelesh, et al.
Publicado: (2025) -
FAST-Prefill: FPGA Accelerated Sparse Attention for Long Context LLM Prefill
por: Jayanth, Rakshith, et al.
Publicado: (2026) -
AI/ML in 3GPP 5G Advanced -- Services and Architecture
por: Taksande, Pradnya, et al.
Publicado: (2025) -
TabConv: Low-Computation CNN Inference via Table Lookups
por: Gupta, Neelesh, et al.
Publicado: (2024) -
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare
por: Desikan, Prasanna, et al.
Publicado: (2026)