OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yunze, Wu, Chi-Hao, Zhou, Enmin, Shen, Junxiao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding
von: Wu, Peiran, et al.
Veröffentlicht: (2026)
von: Wu, Peiran, et al.
Veröffentlicht: (2026)
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
von: Li, Po-han, et al.
Veröffentlicht: (2024)
von: Li, Po-han, et al.
Veröffentlicht: (2024)
Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval
von: Ding, Yiming, et al.
Veröffentlicht: (2026)
von: Ding, Yiming, et al.
Veröffentlicht: (2026)
OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models
von: Chen, Jinshu, et al.
Veröffentlicht: (2025)
von: Chen, Jinshu, et al.
Veröffentlicht: (2025)
OmniControl: Control Any Joint at Any Time for Human Motion Generation
von: Xie, Yiming, et al.
Veröffentlicht: (2023)
von: Xie, Yiming, et al.
Veröffentlicht: (2023)
GAP: Gaussianize Any Point Clouds with Text Guidance
von: Zhang, Weiqi, et al.
Veröffentlicht: (2025)
von: Zhang, Weiqi, et al.
Veröffentlicht: (2025)
AniClipart: Clipart Animation with Text-to-Video Priors
von: Wu, Ronghuan, et al.
Veröffentlicht: (2024)
von: Wu, Ronghuan, et al.
Veröffentlicht: (2024)
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
von: Gu, Yuchao, et al.
Veröffentlicht: (2026)
von: Gu, Yuchao, et al.
Veröffentlicht: (2026)
VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining
von: Liu, Yunze, et al.
Veröffentlicht: (2025)
von: Liu, Yunze, et al.
Veröffentlicht: (2025)
ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction
von: Yeo, Juan, et al.
Veröffentlicht: (2025)
von: Yeo, Juan, et al.
Veröffentlicht: (2025)
Place Anything into Any Video
von: Liu, Ziling, et al.
Veröffentlicht: (2024)
von: Liu, Ziling, et al.
Veröffentlicht: (2024)
Distill Any Depth: Distillation Creates a Stronger Monocular Depth Estimator
von: He, Xiankang, et al.
Veröffentlicht: (2025)
von: He, Xiankang, et al.
Veröffentlicht: (2025)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
SwapAnyone: Consistent and Realistic Video Synthesis for Swapping Any Person into Any Video
von: Zhao, Chengshu, et al.
Veröffentlicht: (2025)
von: Zhao, Chengshu, et al.
Veröffentlicht: (2025)
Trace Anything: Representing Any Video in 4D via Trajectory Fields
von: Liu, Xinhang, et al.
Veröffentlicht: (2025)
von: Liu, Xinhang, et al.
Veröffentlicht: (2025)
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
von: Li, Shufan, et al.
Veröffentlicht: (2024)
von: Li, Shufan, et al.
Veröffentlicht: (2024)
Segment Any Motion in Videos
von: Huang, Nan, et al.
Veröffentlicht: (2025)
von: Huang, Nan, et al.
Veröffentlicht: (2025)
X2SAM: Any Segmentation in Images and Videos
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
von: Li, Yanlin, et al.
Veröffentlicht: (2026)
von: Li, Yanlin, et al.
Veröffentlicht: (2026)
SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA
von: Zheng, Xinyi, et al.
Veröffentlicht: (2026)
von: Zheng, Xinyi, et al.
Veröffentlicht: (2026)
Tracking Any Point with Frame-Event Fusion Network at High Frame Rate
von: Liu, Jiaxiong, et al.
Veröffentlicht: (2024)
von: Liu, Jiaxiong, et al.
Veröffentlicht: (2024)
Any-to-Any Learning in Computational Pathology via Triplet Multimodal Pretraining
von: Sun, Qichen, et al.
Veröffentlicht: (2025)
von: Sun, Qichen, et al.
Veröffentlicht: (2025)
Adversarial Video Promotion Against Text-to-Video Retrieval
von: Tian, Qiwei, et al.
Veröffentlicht: (2025)
von: Tian, Qiwei, et al.
Veröffentlicht: (2025)
Count Anything at Any Granularity
von: Liu, Chang, et al.
Veröffentlicht: (2026)
von: Liu, Chang, et al.
Veröffentlicht: (2026)
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
von: Jeong, Boseung, et al.
Veröffentlicht: (2025)
von: Jeong, Boseung, et al.
Veröffentlicht: (2025)
AnyI2V: Animating Any Conditional Image with Motion Control
von: Li, Ziye, et al.
Veröffentlicht: (2025)
von: Li, Ziye, et al.
Veröffentlicht: (2025)
Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?
von: Liang, Susan, et al.
Veröffentlicht: (2026)
von: Liang, Susan, et al.
Veröffentlicht: (2026)
OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation
von: Zhao, Lin, et al.
Veröffentlicht: (2026)
von: Zhao, Lin, et al.
Veröffentlicht: (2026)
AnyTrans: Translate AnyText in the Image with Large Scale Models
von: Qian, Zhipeng, et al.
Veröffentlicht: (2024)
von: Qian, Zhipeng, et al.
Veröffentlicht: (2024)
DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing
von: Wang, Weitao, et al.
Veröffentlicht: (2025)
von: Wang, Weitao, et al.
Veröffentlicht: (2025)
AnyText: Multilingual Visual Text Generation And Editing
von: Tuo, Yuxiang, et al.
Veröffentlicht: (2023)
von: Tuo, Yuxiang, et al.
Veröffentlicht: (2023)
Referring to Any Person
von: Jiang, Qing, et al.
Veröffentlicht: (2025)
von: Jiang, Qing, et al.
Veröffentlicht: (2025)
Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval
von: Liu, Weijia, et al.
Veröffentlicht: (2025)
von: Liu, Weijia, et al.
Veröffentlicht: (2025)
Depth Any Video with Scalable Synthetic Data
von: Yang, Honghui, et al.
Veröffentlicht: (2024)
von: Yang, Honghui, et al.
Veröffentlicht: (2024)
Gesture2Text: A Generalizable Decoder for Word-Gesture Keyboards in XR Through Trajectory Coarse Discretization and Pre-training
von: Shen, Junxiao, et al.
Veröffentlicht: (2024)
von: Shen, Junxiao, et al.
Veröffentlicht: (2024)
Drive Any Mesh: 4D Latent Diffusion for Mesh Deformation from Video
von: Shi, Yahao, et al.
Veröffentlicht: (2025)
von: Shi, Yahao, et al.
Veröffentlicht: (2025)
AnyAD: Unified Any-Modality Anomaly Detection in Incomplete Multi-Sequence MRI
von: Wu, Changwei, et al.
Veröffentlicht: (2025)
von: Wu, Changwei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
von: Wu, Peiran, et al.
Veröffentlicht: (2025) -
O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding
von: Wu, Peiran, et al.
Veröffentlicht: (2026) -
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
von: Wu, Peiran, et al.
Veröffentlicht: (2025) -
ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos
von: Wu, Peiran, et al.
Veröffentlicht: (2025) -
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
von: Li, Po-han, et al.
Veröffentlicht: (2024)