Saved in:
| Main Authors: | Liu, Yunze, Wu, Chi-Hao, Zhou, Enmin, Shen, Junxiao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.26641 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding
by: Wu, Peiran, et al.
Published: (2026)
by: Wu, Peiran, et al.
Published: (2026)
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
by: Li, Po-han, et al.
Published: (2024)
by: Li, Po-han, et al.
Published: (2024)
Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval
by: Ding, Yiming, et al.
Published: (2026)
by: Ding, Yiming, et al.
Published: (2026)
OmniControl: Control Any Joint at Any Time for Human Motion Generation
by: Xie, Yiming, et al.
Published: (2023)
by: Xie, Yiming, et al.
Published: (2023)
OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models
by: Chen, Jinshu, et al.
Published: (2025)
by: Chen, Jinshu, et al.
Published: (2025)
VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining
by: Liu, Yunze, et al.
Published: (2025)
by: Liu, Yunze, et al.
Published: (2025)
GAP: Gaussianize Any Point Clouds with Text Guidance
by: Zhang, Weiqi, et al.
Published: (2025)
by: Zhang, Weiqi, et al.
Published: (2025)
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
by: Gu, Yuchao, et al.
Published: (2026)
by: Gu, Yuchao, et al.
Published: (2026)
AniClipart: Clipart Animation with Text-to-Video Priors
by: Wu, Ronghuan, et al.
Published: (2024)
by: Wu, Ronghuan, et al.
Published: (2024)
SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA
by: Zheng, Xinyi, et al.
Published: (2026)
by: Zheng, Xinyi, et al.
Published: (2026)
Gesture2Text: A Generalizable Decoder for Word-Gesture Keyboards in XR Through Trajectory Coarse Discretization and Pre-training
by: Shen, Junxiao, et al.
Published: (2024)
by: Shen, Junxiao, et al.
Published: (2024)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
ATAS: Any-to-Any Self-Distillation for Enhanced Open-Vocabulary Dense Prediction
by: Yeo, Juan, et al.
Published: (2025)
by: Yeo, Juan, et al.
Published: (2025)
Distill Any Depth: Distillation Creates a Stronger Monocular Depth Estimator
by: He, Xiankang, et al.
Published: (2025)
by: He, Xiankang, et al.
Published: (2025)
Place Anything into Any Video
by: Liu, Ziling, et al.
Published: (2024)
by: Liu, Ziling, et al.
Published: (2024)
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
by: Li, Shufan, et al.
Published: (2024)
by: Li, Shufan, et al.
Published: (2024)
X2SAM: Any Segmentation in Images and Videos
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
SwapAnyone: Consistent and Realistic Video Synthesis for Swapping Any Person into Any Video
by: Zhao, Chengshu, et al.
Published: (2025)
by: Zhao, Chengshu, et al.
Published: (2025)
Trace Anything: Representing Any Video in 4D via Trajectory Fields
by: Liu, Xinhang, et al.
Published: (2025)
by: Liu, Xinhang, et al.
Published: (2025)
Segment Any Motion in Videos
by: Huang, Nan, et al.
Published: (2025)
by: Huang, Nan, et al.
Published: (2025)
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
by: Li, Yanlin, et al.
Published: (2026)
by: Li, Yanlin, et al.
Published: (2026)
Any-to-Any Learning in Computational Pathology via Triplet Multimodal Pretraining
by: Sun, Qichen, et al.
Published: (2025)
by: Sun, Qichen, et al.
Published: (2025)
AnyFit: Controllable Virtual Try-on for Any Combination of Attire Across Any Scenario
by: Li, Yuhan, et al.
Published: (2024)
by: Li, Yuhan, et al.
Published: (2024)
Tracking Any Point with Frame-Event Fusion Network at High Frame Rate
by: Liu, Jiaxiong, et al.
Published: (2024)
by: Liu, Jiaxiong, et al.
Published: (2024)
AnyTrans: Translate AnyText in the Image with Large Scale Models
by: Qian, Zhipeng, et al.
Published: (2024)
by: Qian, Zhipeng, et al.
Published: (2024)
Count Anything at Any Granularity
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
AnyI2V: Animating Any Conditional Image with Motion Control
by: Li, Ziye, et al.
Published: (2025)
by: Li, Ziye, et al.
Published: (2025)
Depth Any Video with Scalable Synthetic Data
by: Yang, Honghui, et al.
Published: (2024)
by: Yang, Honghui, et al.
Published: (2024)
Adversarial Video Promotion Against Text-to-Video Retrieval
by: Tian, Qiwei, et al.
Published: (2025)
by: Tian, Qiwei, et al.
Published: (2025)
DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing
by: Wang, Weitao, et al.
Published: (2025)
by: Wang, Weitao, et al.
Published: (2025)
Unaligning Everything: Or Aligning Any Text to Any Image in Multimodal Models
by: Salman, Shaeke, et al.
Published: (2024)
by: Salman, Shaeke, et al.
Published: (2024)
Referring to Any Person
by: Jiang, Qing, et al.
Published: (2025)
by: Jiang, Qing, et al.
Published: (2025)
AnyText: Multilingual Visual Text Generation And Editing
by: Tuo, Yuxiang, et al.
Published: (2023)
by: Tuo, Yuxiang, et al.
Published: (2023)
Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?
by: Liang, Susan, et al.
Published: (2026)
by: Liang, Susan, et al.
Published: (2026)
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
by: Jeong, Boseung, et al.
Published: (2025)
by: Jeong, Boseung, et al.
Published: (2025)
AnyAD: Unified Any-Modality Anomaly Detection in Incomplete Multi-Sequence MRI
by: Wu, Changwei, et al.
Published: (2025)
by: Wu, Changwei, et al.
Published: (2025)
OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation
by: Zhao, Lin, et al.
Published: (2026)
by: Zhao, Lin, et al.
Published: (2026)
Similar Items
-
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
by: Wu, Peiran, et al.
Published: (2025) -
O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding
by: Wu, Peiran, et al.
Published: (2026) -
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
by: Wu, Peiran, et al.
Published: (2025) -
ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos
by: Wu, Peiran, et al.
Published: (2025) -
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
by: Li, Po-han, et al.
Published: (2024)