On-device Semantic Selection Made Low Latency and Memory Efficient with Monolithic Forwarding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Jiahao, Lin, Chengliang, Li, Dingji, Dong, Mingkai, Chen, Haibo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Thinking Forward: Memory-Efficient Federated Finetuning of Language Models
by: Panchal, Kunjal, et al.
Published: (2024)
by: Panchal, Kunjal, et al.
Published: (2024)
Latency-aware Human-in-the-Loop Reinforcement Learning for Semantic Communications
by: Li, Peizheng, et al.
Published: (2026)
by: Li, Peizheng, et al.
Published: (2026)
Low-Rank Adaptation of Evolutionary Deep Neural Networks for Efficient Learning of Time-Dependent PDEs
by: Zhang, Jiahao, et al.
Published: (2025)
by: Zhang, Jiahao, et al.
Published: (2025)
LoLaFL: Low-Latency Federated Learning via Forward-only Propagation
by: Zhang, Jierui, et al.
Published: (2024)
by: Zhang, Jierui, et al.
Published: (2024)
Outcome-Aware Tool Selection for Semantic Routers: Latency-Constrained Learning Without LLM Inference
by: Chen, Huamin, et al.
Published: (2026)
by: Chen, Huamin, et al.
Published: (2026)
LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning
by: Lin, Xiaotian, et al.
Published: (2025)
by: Lin, Xiaotian, et al.
Published: (2025)
bispectrum: Selective $G$-Bispectra Made Practical
by: Mathe, Johan, et al.
Published: (2026)
by: Mathe, Johan, et al.
Published: (2026)
SNAP: Low-Latency Test-Time Adaptation with Sparse Updates
by: Cha, Hyeongheon, et al.
Published: (2025)
by: Cha, Hyeongheon, et al.
Published: (2025)
Latent Diffusion Model-Enabled Low-Latency Semantic Communication in the Presence of Semantic Ambiguities and Wireless Channel Noises
by: Pei, Jianhua, et al.
Published: (2024)
by: Pei, Jianhua, et al.
Published: (2024)
PAUSE: Low-Latency and Privacy-Aware Active User Selection for Federated Learning
by: Peleg, Ori, et al.
Published: (2025)
by: Peleg, Ori, et al.
Published: (2025)
Threshold Neuron: A Brain-inspired Artificial Neuron for Efficient On-device Inference
by: Zheng, Zihao, et al.
Published: (2024)
by: Zheng, Zihao, et al.
Published: (2024)
Towards Deep Encrypted Training: Low-Latency, Memory-Efficient, and High-Throughput Inference for Privacy-Preserving Neural Networks
by: Njungle, Nges Brian, et al.
Published: (2026)
by: Njungle, Nges Brian, et al.
Published: (2026)
Match Made with Matrix Completion: Efficient Learning under Matching Interference
by: Tang, Zhiyuan, et al.
Published: (2026)
by: Tang, Zhiyuan, et al.
Published: (2026)
MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers
by: Jaiswal, Ajay, et al.
Published: (2026)
by: Jaiswal, Ajay, et al.
Published: (2026)
Memory-Efficient LLM Training by Various-Grained Low-Rank Projection of Gradients
by: Wang, Yezhen, et al.
Published: (2025)
by: Wang, Yezhen, et al.
Published: (2025)
GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching
by: Regmi, Sajal, et al.
Published: (2024)
by: Regmi, Sajal, et al.
Published: (2024)
PACIFIER: Pacing Opinion Depolarization via a Unified Graph Learning Framework
by: Liao, Mingkai
Published: (2026)
by: Liao, Mingkai
Published: (2026)
Digital Twin-Driven Zero-Shot Fault Diagnosis of Axial Piston Pumps Using Fluid-Borne Noise Signals
by: Dong, Chang, et al.
Published: (2025)
by: Dong, Chang, et al.
Published: (2025)
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
by: Li, Bangzheng, et al.
Published: (2025)
by: Li, Bangzheng, et al.
Published: (2025)
MicroNAS: Memory and Latency Constrained Hardware-Aware Neural Architecture Search for Time Series Classification on Microcontrollers
by: King, Tobias, et al.
Published: (2023)
by: King, Tobias, et al.
Published: (2023)
Low-rank Momentum Factorization for Memory Efficient Training
by: Mahdavinia, Pouria, et al.
Published: (2025)
by: Mahdavinia, Pouria, et al.
Published: (2025)
Client Selection in Federated Learning with Data Heterogeneity and Network Latencies
by: Vardhan, Harsh, et al.
Published: (2025)
by: Vardhan, Harsh, et al.
Published: (2025)
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
Mono-Forward: Revisiting Forward-Forward through Objective-Locality Decomposition
by: Gong, James, et al.
Published: (2025)
by: Gong, James, et al.
Published: (2025)
Memory-Efficient Fine-Tuning via Low-Rank Activation Compression
by: Shi, Jiang-Xin, et al.
Published: (2025)
by: Shi, Jiang-Xin, et al.
Published: (2025)
Grass: Compute Efficient Low-Memory LLM Training with Structured Sparse Gradients
by: Muhamed, Aashiq, et al.
Published: (2024)
by: Muhamed, Aashiq, et al.
Published: (2024)
Traces Propagation: Memory-Efficient and Scalable Forward-Only Learning in Spiking Neural Networks
by: Pes, Lorenzo, et al.
Published: (2025)
by: Pes, Lorenzo, et al.
Published: (2025)
Stateful Inference for Low-Latency Multi-Agent Tool Calling
by: Norgren, Victor
Published: (2026)
by: Norgren, Victor
Published: (2026)
Adaptive Regret for Bandits Made Possible: Two Queries Suffice
by: Lu, Zhou, et al.
Published: (2024)
by: Lu, Zhou, et al.
Published: (2024)
Low Latency Transformer Inference on FPGAs for Physics Applications with hls4ml
by: Jiang, Zhixing, et al.
Published: (2024)
by: Jiang, Zhixing, et al.
Published: (2024)
Muon with Spectral Guidance: Efficient Optimization for Scientific Machine Learning
by: Lu, Binghang, et al.
Published: (2026)
by: Lu, Binghang, et al.
Published: (2026)
Forward-Forward Autoencoder Architectures for Energy-Efficient Wireless Communications
by: Seifert, Daniel, et al.
Published: (2025)
by: Seifert, Daniel, et al.
Published: (2025)
Fast Forwarding Low-Rank Training
by: Rahamim, Adir, et al.
Published: (2024)
by: Rahamim, Adir, et al.
Published: (2024)
Adaptive Decentralized Federated Learning in Energy and Latency Constrained Wireless Networks
by: Yan, Zhigang, et al.
Published: (2024)
by: Yan, Zhigang, et al.
Published: (2024)
TrajMamba: An Efficient and Semantic-rich Vehicle Trajectory Pre-training Model
by: Liu, Yichen, et al.
Published: (2025)
by: Liu, Yichen, et al.
Published: (2025)
Selectivity and Shape in the Design of Forward-Forward Goodness Functions
by: Akkus, Talha Ruzgar, et al.
Published: (2026)
by: Akkus, Talha Ruzgar, et al.
Published: (2026)
The Mirrored Influence Hypothesis: Efficient Data Influence Estimation by Harnessing Forward Passes
by: Ko, Myeongseob, et al.
Published: (2024)
by: Ko, Myeongseob, et al.
Published: (2024)
Memory-Efficient Acceleration of Block Low-Rank Foundation Models on Resource Constrained GPUs
by: Abillama, Pierre, et al.
Published: (2025)
by: Abillama, Pierre, et al.
Published: (2025)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
by: Xiao, Kaicheng, et al.
Published: (2026)
by: Xiao, Kaicheng, et al.
Published: (2026)
An Adaptive Clustering Scheme for Client Selections in Communication-Efficient Federated Learning
by: Chen, Yan-Ann, et al.
Published: (2025)
by: Chen, Yan-Ann, et al.
Published: (2025)
Similar Items
-
Thinking Forward: Memory-Efficient Federated Finetuning of Language Models
by: Panchal, Kunjal, et al.
Published: (2024) -
Latency-aware Human-in-the-Loop Reinforcement Learning for Semantic Communications
by: Li, Peizheng, et al.
Published: (2026) -
Low-Rank Adaptation of Evolutionary Deep Neural Networks for Efficient Learning of Time-Dependent PDEs
by: Zhang, Jiahao, et al.
Published: (2025) -
LoLaFL: Low-Latency Federated Learning via Forward-only Propagation
by: Zhang, Jierui, et al.
Published: (2024) -
Outcome-Aware Tool Selection for Semantic Routers: Latency-Constrained Learning Without LLM Inference
by: Chen, Huamin, et al.
Published: (2026)