Work-Efficient Parallel Non-Maximum Suppression Kernels
Fuente:
arXiv
Saved in:
| Main Authors: | Oro, David, Fernández, Carles, Martorell, Xavier, Hernando, Javier |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The High Cost of Keeping Warm: Characterizing Overhead in Serverless Autoscaling Policies
by: Kondrashov, Leonid, et al.
Published: (2025)
by: Kondrashov, Leonid, et al.
Published: (2025)
Sky$^ε$-Tree: Embracing the Batch Updates of B$^ε$-trees through Access Port Parallelism on Skyrmion Racetrack Memory
by: Tsai, Yu-Shiang, et al.
Published: (2024)
by: Tsai, Yu-Shiang, et al.
Published: (2024)
Trident: Adaptive Scheduling for Heterogeneous Multimodal Data Pipelines
by: Pan, Ding, et al.
Published: (2026)
by: Pan, Ding, et al.
Published: (2026)
SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
by: Lysenstøen, Christian
Published: (2026)
by: Lysenstøen, Christian
Published: (2026)
A Methodology to Assess Power Modeling in Energy-Aware Federated Learning on Heterogeneous Mobile Devices
by: Jallouli, Chaimae, et al.
Published: (2026)
by: Jallouli, Chaimae, et al.
Published: (2026)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
by: Li, Xiangchen, et al.
Published: (2026)
by: Li, Xiangchen, et al.
Published: (2026)
N2N: A Parallel Framework for Large-Scale MILP under Distributed Memory
by: Wang, Longfei, et al.
Published: (2025)
by: Wang, Longfei, et al.
Published: (2025)
Aethon: A Reference-Based Replication Primitive for Constant-Time Instantiation of Stateful AI Agents
by: Rao, Swanand, et al.
Published: (2026)
by: Rao, Swanand, et al.
Published: (2026)
Serving LLMs in HPC Clusters: A Comparative Study of Qualcomm Cloud AI 100 Ultra and NVIDIA Data Center GPUs
by: Sada, Mohammad Firas, et al.
Published: (2025)
by: Sada, Mohammad Firas, et al.
Published: (2025)
INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference
by: Šabanović, Ahmed, et al.
Published: (2026)
by: Šabanović, Ahmed, et al.
Published: (2026)
NBI-Slurm: Simplified submission of Slurm jobs with energy saving mode
by: Telatin, Andrea
Published: (2026)
by: Telatin, Andrea
Published: (2026)
Vertical Federated Image Segmentation
by: Mandal, Paul K., et al.
Published: (2024)
by: Mandal, Paul K., et al.
Published: (2024)
Horizontal Federated Computer Vision
by: Mandal, Paul K., et al.
Published: (2023)
by: Mandal, Paul K., et al.
Published: (2023)
Heterogeneous Model Fusion for Privacy-Aware Multi-Camera Surveillance via Synthetic Domain Adaptation
by: Lu, Peggy Joy, et al.
Published: (2026)
by: Lu, Peggy Joy, et al.
Published: (2026)
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
by: Li, Shigang, et al.
Published: (2020)
by: Li, Shigang, et al.
Published: (2020)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
by: Li, Xiangchen, et al.
Published: (2026)
by: Li, Xiangchen, et al.
Published: (2026)
Optimizing Multi-DNN Inference on Mobile Devices through Heterogeneous Processor Co-Execution
by: Gao, Yunquan, et al.
Published: (2025)
by: Gao, Yunquan, et al.
Published: (2025)
Efficient Construction of Large Search Spaces for Auto-Tuning
by: Willemsen, Floris-Jan, et al.
Published: (2025)
by: Willemsen, Floris-Jan, et al.
Published: (2025)
Learning Interpretable Scheduling Algorithms for Data Processing Clusters
by: Hu, Zhibo, et al.
Published: (2024)
by: Hu, Zhibo, et al.
Published: (2024)
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
by: Wang, Wenyi, et al.
Published: (2025)
by: Wang, Wenyi, et al.
Published: (2025)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
by: Shakeri, Heman, et al.
Published: (2026)
by: Shakeri, Heman, et al.
Published: (2026)
Next-Generation Event-Driven Architectures: Performance, Scalability, and Intelligent Orchestration Across Messaging Frameworks
by: Arafat, Jahidul, et al.
Published: (2025)
by: Arafat, Jahidul, et al.
Published: (2025)
An Evaluation of Massively Parallel Algorithms for DFA Minimization
by: Martens, Jan, et al.
Published: (2024)
by: Martens, Jan, et al.
Published: (2024)
Stream-K++: Adaptive GPU GEMM Kernel Scheduling and Selection using Bloom Filters
by: Sadasivan, Harisankar, et al.
Published: (2024)
by: Sadasivan, Harisankar, et al.
Published: (2024)
Cost-Aware Logging: Measuring the Financial Impact of Excessive Log Retention in Small-Scale Cloud Deployments
by: Putra, Jody Almaida
Published: (2026)
by: Putra, Jody Almaida
Published: (2026)
Static Batching of Irregular Workloads on GPUs: Framework and Application to Efficient MoE Model Inference
by: Li, Yinghan, et al.
Published: (2025)
by: Li, Yinghan, et al.
Published: (2025)
Optimizing edge AI models on HPC systems with the edge in the loop
by: Aach, Marcel, et al.
Published: (2025)
by: Aach, Marcel, et al.
Published: (2025)
Deploy, Calibrate, Monitor, Heal -- No Human Required: An Autonomous AI SRE Agent for Elasticsearch
by: Mukkolakkal, Muhamed Ramees Cheriya
Published: (2026)
by: Mukkolakkal, Muhamed Ramees Cheriya
Published: (2026)
Deadline-Aware Joint Task Scheduling and Offloading in Mobile Edge Computing Systems
by: Nguyen, Ngoc Hung, et al.
Published: (2025)
by: Nguyen, Ngoc Hung, et al.
Published: (2025)
Agent-based modeling for realistic reproduction of human mobility and contact behavior to evaluate test and isolation strategies in epidemic infectious disease spread
by: Kerkmann, David, et al.
Published: (2024)
by: Kerkmann, David, et al.
Published: (2024)
Hector: An Efficient Programming and Compilation Framework for Implementing Relational Graph Neural Networks in GPU Architectures
by: Wu, Kun, et al.
Published: (2023)
by: Wu, Kun, et al.
Published: (2023)
PIM-STM: Software Transactional Memory for Processing-In-Memory Systems
by: Lopes, André, et al.
Published: (2024)
by: Lopes, André, et al.
Published: (2024)
Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
by: Will, Jonathan, et al.
Published: (2025)
by: Will, Jonathan, et al.
Published: (2025)
VSS Challenge Problem: Verifying the Correctness of AllReduce Algorithms in the MPICH Implementation of MPI
by: Hovland, Paul D.
Published: (2025)
by: Hovland, Paul D.
Published: (2025)
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
by: Li, Chendi, et al.
Published: (2022)
by: Li, Chendi, et al.
Published: (2022)
Enabling Practical Transparent Checkpointing for MPI: A Topological Sort Approach
by: Xu, Yao, et al.
Published: (2024)
by: Xu, Yao, et al.
Published: (2024)
FLEdge: Benchmarking Federated Machine Learning Applications in Edge Computing Systems
by: Woisetschläger, Herbert, et al.
Published: (2023)
by: Woisetschläger, Herbert, et al.
Published: (2023)
TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
by: Bian, Zhuohang, et al.
Published: (2025)
by: Bian, Zhuohang, et al.
Published: (2025)
Spark-LLM-Eval: A Distributed Framework for Statistically Rigorous Large Language Model Evaluation
by: Mitra, Subhadip
Published: (2026)
by: Mitra, Subhadip
Published: (2026)
Operational Memory Architecture for Kubernetes:Preserving Causal Context Across the Evidence Horizon
by: Khan, Shamsher
Published: (2026)
by: Khan, Shamsher
Published: (2026)
Similar Items
-
The High Cost of Keeping Warm: Characterizing Overhead in Serverless Autoscaling Policies
by: Kondrashov, Leonid, et al.
Published: (2025) -
Sky$^ε$-Tree: Embracing the Batch Updates of B$^ε$-trees through Access Port Parallelism on Skyrmion Racetrack Memory
by: Tsai, Yu-Shiang, et al.
Published: (2024) -
Trident: Adaptive Scheduling for Heterogeneous Multimodal Data Pipelines
by: Pan, Ding, et al.
Published: (2026) -
SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
by: Lysenstøen, Christian
Published: (2026) -
A Methodology to Assess Power Modeling in Energy-Aware Federated Learning on Heterogeneous Mobile Devices
by: Jallouli, Chaimae, et al.
Published: (2026)