Deep Recommender Models Inference: Automatic Asymmetric Data Flow Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Ruggeri, Giuseppe, Andri, Renzo, Pagliari, Daniele Jahier, Cavigelli, Lukas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Passing the Baton: High Throughput Distributed Disk-Based Vector Search with BatANN
by: Dang, Nam Anh, et al.
Published: (2025)
by: Dang, Nam Anh, et al.
Published: (2025)
A framework to reason about consistency and atomicity guarantees in a sparsely-connected, partially-replicated peer-to-peer system
by: Nair, Sreeja S., et al.
Published: (2026)
by: Nair, Sreeja S., et al.
Published: (2026)
LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?
by: Takehi, Rikiya, et al.
Published: (2024)
by: Takehi, Rikiya, et al.
Published: (2024)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
by: Potocnik, Viviane, et al.
Published: (2024)
by: Potocnik, Viviane, et al.
Published: (2024)
HierarchicalKV: A GPU Hash Table with Cache Semantics for Continuous Online Embedding Storage
by: Rong, Haidong, et al.
Published: (2026)
by: Rong, Haidong, et al.
Published: (2026)
Toward a Universal GPU Instruction Set Architecture: A Cross-Vendor Analysis of Hardware-Invariant Computational Primitives in Parallel Processors
by: Abraham, Ojima, et al.
Published: (2026)
by: Abraham, Ojima, et al.
Published: (2026)
AWARE: Evaluating PriorityFresh Caching for Offline Emergency Warning Systems
by: Melvin, Charles, et al.
Published: (2025)
by: Melvin, Charles, et al.
Published: (2025)
Behavior-Aware Dual-Channel Preference Learning for Heterogeneous Sequential Recommendation
by: Xiao, Jing, et al.
Published: (2026)
by: Xiao, Jing, et al.
Published: (2026)
Unlocking Python's Cores: Hardware Usage and Energy Implications of Removing the GIL
by: Salazar, José Daniel Montoya
Published: (2026)
by: Salazar, José Daniel Montoya
Published: (2026)
From BM25 to Corrective RAG: Benchmarking Retrieval Strategies for Text-and-Table Documents
by: Akarsu, Meftun, et al.
Published: (2026)
by: Akarsu, Meftun, et al.
Published: (2026)
The Treatment of Ties in Rank-Biased Overlap
by: Corsi, Matteo, et al.
Published: (2024)
by: Corsi, Matteo, et al.
Published: (2024)
HTVM: Efficient Neural Network Deployment On Heterogeneous TinyML Platforms
by: Van Delm, Josse, et al.
Published: (2024)
by: Van Delm, Josse, et al.
Published: (2024)
Addressing tokens dynamic generation, propagation, storage and renewal to secure the GlideinWMS pilot based jobs and system
by: Coimbra, Bruno Moreira, et al.
Published: (2025)
by: Coimbra, Bruno Moreira, et al.
Published: (2025)
Using Containers to Speed Up Development, to Run Integration Tests and to Teach About Distributed Systems
by: Mambelli, Marco, et al.
Published: (2025)
by: Mambelli, Marco, et al.
Published: (2025)
GlideinBenchmark: collecting resource information to optimize provisioning
by: Mambelli, Marco, et al.
Published: (2025)
by: Mambelli, Marco, et al.
Published: (2025)
Neural Router: Semantic Content Matching for Agentic AI
by: Lovén, Lauri, et al.
Published: (2026)
by: Lovén, Lauri, et al.
Published: (2026)
flexvec: SQL Vector Retrieval with Programmatic Embedding Modulation
by: Delmas, Damian
Published: (2026)
by: Delmas, Damian
Published: (2026)
Diversification as Risk Minimization
by: Takehi, Rikiya, et al.
Published: (2025)
by: Takehi, Rikiya, et al.
Published: (2025)
Stop Using the Wilcoxon Test: Myth, Misconception and Misuse in IR Research
by: Urbano, Julián
Published: (2026)
by: Urbano, Julián
Published: (2026)
AlayaDB: The Data Foundation for Efficient and Effective Long-context LLM Inference
by: Deng, Yangshen, et al.
Published: (2025)
by: Deng, Yangshen, et al.
Published: (2025)
Inside VOLT: Designing an Open-Source GPU Compiler
by: Jeong, Shinnung, et al.
Published: (2025)
by: Jeong, Shinnung, et al.
Published: (2025)
SLA Management in Reconfigurable Multi-Agent RAG: A Systems Approach to Question Answering
by: Iannelli, Michael, et al.
Published: (2024)
by: Iannelli, Michael, et al.
Published: (2024)
Siren Federate: Bridging document, relational, and graph models for exploratory graph analysis
by: Bordea, Georgeta, et al.
Published: (2025)
by: Bordea, Georgeta, et al.
Published: (2025)
PeakNetFP: Peak-based Neural Audio Fingerprinting Robust to Extreme Time Stretching
by: Cortès-Sebastià, Guillem, et al.
Published: (2025)
by: Cortès-Sebastià, Guillem, et al.
Published: (2025)
I/O in Machine Learning Applications on HPC Systems: A 360-degree Survey
by: Lewis, Noah, et al.
Published: (2024)
by: Lewis, Noah, et al.
Published: (2024)
A Recommender System Based on Binary Matrix Representations for Cognitive Disorders
by: Kutil, Raoul H., et al.
Published: (2025)
by: Kutil, Raoul H., et al.
Published: (2025)
GPU-Augmented OLAP Execution Engine: GPU Offloading
by: Chang, Ilsun
Published: (2025)
by: Chang, Ilsun
Published: (2025)
DISTRIBUTEDANN: Efficient Scaling of a Single DISKANN Graph Across Thousands of Computers
by: Adams, Philip, et al.
Published: (2025)
by: Adams, Philip, et al.
Published: (2025)
De-DSI: Decentralised Differentiable Search Index
by: Neague, Petru, et al.
Published: (2024)
by: Neague, Petru, et al.
Published: (2024)
Bhakti: A Lightweight Vector Database Management System for Endowing Large Language Models with Semantic Search Capabilities and Memory
by: Wu, Zihao
Published: (2025)
by: Wu, Zihao
Published: (2025)
Valori: A Deterministic Memory Substrate for AI Systems
by: Gudur, Varshith
Published: (2025)
by: Gudur, Varshith
Published: (2025)
Deploy, Calibrate, Monitor, Heal -- No Human Required: An Autonomous AI SRE Agent for Elasticsearch
by: Mukkolakkal, Muhamed Ramees Cheriya
Published: (2026)
by: Mukkolakkal, Muhamed Ramees Cheriya
Published: (2026)
Scaling and Load-Balancing Equi-Joins
by: Metwally, Ahmed
Published: (2022)
by: Metwally, Ahmed
Published: (2022)
EnterpriseRAG-Bench: A RAG Benchmark for Company Internal Knowledge
by: Sun, Yuhong, et al.
Published: (2026)
by: Sun, Yuhong, et al.
Published: (2026)
Cost Trade-offs of Reasoning and Non-Reasoning Large Language Models in Text-to-SQL
by: Deochake, Saurabh, et al.
Published: (2025)
by: Deochake, Saurabh, et al.
Published: (2025)
Beyond Similarity Search: A Unified Data Layer for Production RAG Systems
by: Budigi, Venkata Krishna Prasanth, et al.
Published: (2026)
by: Budigi, Venkata Krishna Prasanth, et al.
Published: (2026)
PromptChain: A Decentralized Web3 Architecture for Managing AI Prompts as Digital Assets
by: Bara, Marc
Published: (2025)
by: Bara, Marc
Published: (2025)
Splitwise: Efficient generative LLM inference using phase splitting
by: Patel, Pratyush, et al.
Published: (2023)
by: Patel, Pratyush, et al.
Published: (2023)
Aggregating Digital Identities through Bridging. An Integration of Open Authentication Protocols for Web3 Identifiers
by: Biedermann, Ben, et al.
Published: (2025)
by: Biedermann, Ben, et al.
Published: (2025)
Serverless GPU Architecture for Enterprise HR Analytics: A Production-Scale BDaaS Implementation
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
Similar Items
-
Passing the Baton: High Throughput Distributed Disk-Based Vector Search with BatANN
by: Dang, Nam Anh, et al.
Published: (2025) -
A framework to reason about consistency and atomicity guarantees in a sparsely-connected, partially-replicated peer-to-peer system
by: Nair, Sreeja S., et al.
Published: (2026) -
LLM-Assisted Relevance Assessments: When Should We Ask LLMs for Help?
by: Takehi, Rikiya, et al.
Published: (2024) -
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
by: Potocnik, Viviane, et al.
Published: (2024) -
HierarchicalKV: A GPU Hash Table with Cache Semantics for Continuous Online Embedding Storage
by: Rong, Haidong, et al.
Published: (2026)