O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Peiran, Liu, Yunze, Wu, Chi-Hao, Chen, Chen, Shen, Junxiao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916047136751616
author Wu, Peiran
Liu, Yunze
Wu, Chi-Hao
Chen, Chen
Shen, Junxiao
author_facet Wu, Peiran
Liu, Yunze
Wu, Chi-Hao
Chen, Chen
Shen, Junxiao
contents Omnimodal large language models enable unified audio video understanding, but long joint token sequences make inference costly, and existing benchmarks do not fully isolate audio visual association in noisy user generated videos. We introduce UGC-AVQA, a public UGC benchmark with 1,000 videos and 4,816 QA pairs, where an audio removal test ensures that benchmark questions require both acoustic and visual evidence. To reduce inference cost, we propose OMAC, a training free plug in compression method that preserves salient visual memory and temporally grounded audio anchors. To further make compact models robust to compressed inputs, we introduce O-MARC, a compression distillation framework for learning with memory compressed multimodal contexts. On Qwen2.5-Omni-3B, O-MARC improves the average score across four benchmarks to 45.8, outperforming full token inference at 44.1 and OmniZip at 41.0. OMAC also keeps inference efficient, reducing latency by 34.6\% (1.53$\times$ speedup) and memory by 34.7\% compared with full token inference.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26584
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding
Wu, Peiran
Liu, Yunze
Wu, Chi-Hao
Chen, Chen
Shen, Junxiao
Computer Vision and Pattern Recognition
Omnimodal large language models enable unified audio video understanding, but long joint token sequences make inference costly, and existing benchmarks do not fully isolate audio visual association in noisy user generated videos. We introduce UGC-AVQA, a public UGC benchmark with 1,000 videos and 4,816 QA pairs, where an audio removal test ensures that benchmark questions require both acoustic and visual evidence. To reduce inference cost, we propose OMAC, a training free plug in compression method that preserves salient visual memory and temporally grounded audio anchors. To further make compact models robust to compressed inputs, we introduce O-MARC, a compression distillation framework for learning with memory compressed multimodal contexts. On Qwen2.5-Omni-3B, O-MARC improves the average score across four benchmarks to 45.8, outperforming full token inference at 44.1 and OmniZip at 41.0. OMAC also keeps inference efficient, reducing latency by 34.6\% (1.53$\times$ speedup) and memory by 34.7\% compared with full token inference.
title O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.26584