MMEdge: Accelerating On-device Multimodal Inference via Pipelined Sensing and Encoding
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Runxi, Yu, Mingxuan, Tsoi, Mingyu, Ouyang, Xiaomin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoViD: View-Invariant 3D Human Pose Estimation via Motion-View Disentanglement
by: Liu, Yejia, et al.
Published: (2026)
by: Liu, Yejia, et al.
Published: (2026)
Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta Compression
by: Wang, Xiaohui, et al.
Published: (2025)
by: Wang, Xiaohui, et al.
Published: (2025)
A Wireless Collaborated Inference Acceleration Framework for Plant Disease Recognition
by: Zhu, Hele, et al.
Published: (2025)
by: Zhu, Hele, et al.
Published: (2025)
Learning to Inference Adaptively for Multimodal Large Language Models
by: Xu, Zhuoyan, et al.
Published: (2025)
by: Xu, Zhuoyan, et al.
Published: (2025)
A Multimodal RAG Framework for Housing Damage Assessment: Collaborative Optimization of Image Encoding and Policy Vector Retrieval
by: Miao, Jiayi, et al.
Published: (2025)
by: Miao, Jiayi, et al.
Published: (2025)
Efficient Noise Mitigation for Enhancing Inference Accuracy in DNNs on Mixed-Signal Accelerators
by: Azizi, Seyedarmin, et al.
Published: (2024)
by: Azizi, Seyedarmin, et al.
Published: (2024)
MeanCache: From Instantaneous to Average Velocity for Accelerating Flow Matching Inference
by: Gao, Huanlin, et al.
Published: (2026)
by: Gao, Huanlin, et al.
Published: (2026)
Graph Alignment via Dual-Pass Spectral Encoding and Latent Space Communication
by: Behmanesh, Maysam, et al.
Published: (2025)
by: Behmanesh, Maysam, et al.
Published: (2025)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
CellCLIP -- Learning Perturbation Effects in Cell Painting via Text-Guided Contrastive Learning
by: Lu, Mingyu, et al.
Published: (2025)
by: Lu, Mingyu, et al.
Published: (2025)
LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model
by: Muhtar, Dilxat, et al.
Published: (2024)
by: Muhtar, Dilxat, et al.
Published: (2024)
FreqCa: Accelerating Diffusion Models via Frequency-Aware Caching
by: Liu, Jiacheng, et al.
Published: (2025)
by: Liu, Jiacheng, et al.
Published: (2025)
FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
by: Wang, Zehan, et al.
Published: (2024)
by: Wang, Zehan, et al.
Published: (2024)
MultiSense-Pneumo: A Multimodal Learning Framework for Pneumonia Screening in Resource-Constrained Settings
by: Jayakody, Dineth, et al.
Published: (2026)
by: Jayakody, Dineth, et al.
Published: (2026)
AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
by: Huang, Runhui, et al.
Published: (2026)
by: Huang, Runhui, et al.
Published: (2026)
FedAFD: Multimodal Federated Learning via Adversarial Fusion and Distillation
by: Tan, Min, et al.
Published: (2026)
by: Tan, Min, et al.
Published: (2026)
SafeFix: Targeted Model Repair via Controlled Image Generation
by: Xu, Ouyang, et al.
Published: (2025)
by: Xu, Ouyang, et al.
Published: (2025)
Multi-Head Encoding for Extreme Label Classification
by: Liang, Daojun, et al.
Published: (2024)
by: Liang, Daojun, et al.
Published: (2024)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
by: Xiao, Tong, et al.
Published: (2025)
by: Xiao, Tong, et al.
Published: (2025)
MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts
by: Zhu, Jie, et al.
Published: (2024)
by: Zhu, Jie, et al.
Published: (2024)
ATraDiff: Accelerating Online Reinforcement Learning with Imaginary Trajectories
by: Yang, Qianlan, et al.
Published: (2024)
by: Yang, Qianlan, et al.
Published: (2024)
Robustness to distribution shifts of compressed networks for edge devices
by: Shen, Lulan, et al.
Published: (2024)
by: Shen, Lulan, et al.
Published: (2024)
MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language Models
by: Shi, Yang, et al.
Published: (2026)
by: Shi, Yang, et al.
Published: (2026)
Accelerating Diffusion Transformers with Token-wise Feature Caching
by: Zou, Chang, et al.
Published: (2024)
by: Zou, Chang, et al.
Published: (2024)
Accelerating Large-Scale Dataset Distillation via Exploration-Exploitation Optimization
by: Alahmadi, Muhammad J., et al.
Published: (2026)
by: Alahmadi, Muhammad J., et al.
Published: (2026)
FlexLoc: Conditional Neural Networks for Zero-Shot Sensor Perspective Invariance in Object Localization with Distributed Multimodal Sensors
by: Wu, Jason, et al.
Published: (2024)
by: Wu, Jason, et al.
Published: (2024)
Does a Neural Network Really Encode Symbolic Concepts?
by: Li, Mingjie, et al.
Published: (2023)
by: Li, Mingjie, et al.
Published: (2023)
Transformers and Slot Encoding for Sample Efficient Physical World Modelling
by: Petri, Francesco, et al.
Published: (2024)
by: Petri, Francesco, et al.
Published: (2024)
FastVLM: Efficient Vision Encoding for Vision Language Models
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2024)
Optimizing Helmet Detection with Hybrid YOLO Pipelines: A Detailed Analysis
by: M, Vaikunth, et al.
Published: (2024)
by: M, Vaikunth, et al.
Published: (2024)
Vehicle-centric Perception via Multimodal Structured Pre-training
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
Robust Multimodal Learning via Cross-Modal Proxy Tokens
by: Reza, Md Kaykobad, et al.
Published: (2025)
by: Reza, Md Kaykobad, et al.
Published: (2025)
Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
by: Chang, Kai-Po, et al.
Published: (2025)
by: Chang, Kai-Po, et al.
Published: (2025)
Deep Dependency Networks and Advanced Inference Schemes for Multi-Label Classification
by: Arya, Shivvrat, et al.
Published: (2024)
by: Arya, Shivvrat, et al.
Published: (2024)
DCL-SE: Dynamic Curriculum Learning for Spatiotemporal Encoding of Brain Imaging
by: Zhou, Meihua, et al.
Published: (2025)
by: Zhou, Meihua, et al.
Published: (2025)
Extrapolation of Periodic Functions Using Binary Encoding of Continuous Numerical Values
by: Powell, Brian P., et al.
Published: (2025)
by: Powell, Brian P., et al.
Published: (2025)
When Graph meets Multimodal: Benchmarking and Meditating on Multimodal Attributed Graphs Learning
by: Yan, Hao, et al.
Published: (2024)
by: Yan, Hao, et al.
Published: (2024)
VIPaint: Image Inpainting with Pre-Trained Diffusion Models via Variational Inference
by: Agarwal, Sakshi, et al.
Published: (2024)
by: Agarwal, Sakshi, et al.
Published: (2024)
AllocMV: Optimal Resource Allocation for Music Video Generation via Structured Persistent State
by: Wang, Huimin, et al.
Published: (2026)
by: Wang, Huimin, et al.
Published: (2026)
UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience
by: Lin, Zichuan, et al.
Published: (2026)
by: Lin, Zichuan, et al.
Published: (2026)
Similar Items
-
MoViD: View-Invariant 3D Human Pose Estimation via Motion-View Disentanglement
by: Liu, Yejia, et al.
Published: (2026) -
Breaking the Compression Ceiling: Data-Free Pipeline for Ultra-Efficient Delta Compression
by: Wang, Xiaohui, et al.
Published: (2025) -
A Wireless Collaborated Inference Acceleration Framework for Plant Disease Recognition
by: Zhu, Hele, et al.
Published: (2025) -
Learning to Inference Adaptively for Multimodal Large Language Models
by: Xu, Zhuoyan, et al.
Published: (2025) -
A Multimodal RAG Framework for Housing Damage Assessment: Collaborative Optimization of Image Encoding and Policy Vector Retrieval
by: Miao, Jiayi, et al.
Published: (2025)