OmniVid: A Generative Framework for Universal Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Junke, Chen, Dongdong, Luo, Chong, He, Bo, Yuan, Lu, Wu, Zuxuan, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OmniTracker: Unifying Object Tracking by Tracking-with-Detection
von: Wang, Junke, et al.
Veröffentlicht: (2023)
von: Wang, Junke, et al.
Veröffentlicht: (2023)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
von: Wang, Junke, et al.
Veröffentlicht: (2024)
von: Wang, Junke, et al.
Veröffentlicht: (2024)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026)
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026)
FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
von: Xie, Yiweng, et al.
Veröffentlicht: (2026)
von: Xie, Yiweng, et al.
Veröffentlicht: (2026)
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
von: He, Zhihao, et al.
Veröffentlicht: (2026)
von: He, Zhihao, et al.
Veröffentlicht: (2026)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
von: Xu, Yicheng, et al.
Veröffentlicht: (2025)
von: Xu, Yicheng, et al.
Veröffentlicht: (2025)
SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
von: Wang, Junke, et al.
Veröffentlicht: (2025)
von: Wang, Junke, et al.
Veröffentlicht: (2025)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
von: Wang, Yi, et al.
Veröffentlicht: (2023)
von: Wang, Yi, et al.
Veröffentlicht: (2023)
REDUCIO! Generating 1K Video within 16 Seconds using Extremely Compressed Motion Latents
von: Tian, Rui, et al.
Veröffentlicht: (2024)
von: Tian, Rui, et al.
Veröffentlicht: (2024)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
von: Zhou, Ziwei, et al.
Veröffentlicht: (2025)
von: Zhou, Ziwei, et al.
Veröffentlicht: (2025)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
von: Qiu, Zongyang, et al.
Veröffentlicht: (2025)
von: Qiu, Zongyang, et al.
Veröffentlicht: (2025)
GenRec: Unifying Video Generation and Recognition with Diffusion Models
von: Weng, Zejia, et al.
Veröffentlicht: (2024)
von: Weng, Zejia, et al.
Veröffentlicht: (2024)
StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation
von: Tu, Shuyuan, et al.
Veröffentlicht: (2025)
von: Tu, Shuyuan, et al.
Veröffentlicht: (2025)
Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning
von: You, Zuyao, et al.
Veröffentlicht: (2025)
von: You, Zuyao, et al.
Veröffentlicht: (2025)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
von: Yang, Zhoufaran, et al.
Veröffentlicht: (2025)
von: Yang, Zhoufaran, et al.
Veröffentlicht: (2025)
TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction
von: Ma, Yukuo, et al.
Veröffentlicht: (2025)
von: Ma, Yukuo, et al.
Veröffentlicht: (2025)
DeRA: Decoupled Representation Alignment for Video Tokenization
von: Guo, Pengbo, et al.
Veröffentlicht: (2025)
von: Guo, Pengbo, et al.
Veröffentlicht: (2025)
TextVidBench: A Benchmark for Long Video Scene Text Understanding
von: Zhong, Yangyang, et al.
Veröffentlicht: (2025)
von: Zhong, Yangyang, et al.
Veröffentlicht: (2025)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
von: Xi, Dianbing, et al.
Veröffentlicht: (2025)
von: Xi, Dianbing, et al.
Veröffentlicht: (2025)
DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data
von: Jin, Wonjoon, et al.
Veröffentlicht: (2026)
von: Jin, Wonjoon, et al.
Veröffentlicht: (2026)
Learning Accurate Segmentation Purely from Self-Supervision
von: You, Zuyao, et al.
Veröffentlicht: (2026)
von: You, Zuyao, et al.
Veröffentlicht: (2026)
DCDM: Divide-and-Conquer Diffusion Models for Consistency-Preserving Video Generation
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2026)
Omni-Video: Democratizing Unified Video Understanding and Generation
von: Tan, Zhiyu, et al.
Veröffentlicht: (2025)
von: Tan, Zhiyu, et al.
Veröffentlicht: (2025)
OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer
von: Zhang, Pengze, et al.
Veröffentlicht: (2026)
von: Zhang, Pengze, et al.
Veröffentlicht: (2026)
Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
von: Zheng, Peng, et al.
Veröffentlicht: (2025)
von: Zheng, Peng, et al.
Veröffentlicht: (2025)
UniVid: The Open-Source Unified Video Model
von: Luo, Jiabin, et al.
Veröffentlicht: (2025)
von: Luo, Jiabin, et al.
Veröffentlicht: (2025)
Multi-Prompt Alignment for Multi-Source Unsupervised Domain Adaptation
von: Chen, Haoran, et al.
Veröffentlicht: (2022)
von: Chen, Haoran, et al.
Veröffentlicht: (2022)
CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization
von: Chen, Yitong, et al.
Veröffentlicht: (2026)
von: Chen, Yitong, et al.
Veröffentlicht: (2026)
VidSketch: Hand-drawn Sketch-Driven Video Generation with Diffusion Control
von: Jiang, Lifan, et al.
Veröffentlicht: (2025)
von: Jiang, Lifan, et al.
Veröffentlicht: (2025)
O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding
von: Wu, Peiran, et al.
Veröffentlicht: (2026)
von: Wu, Peiran, et al.
Veröffentlicht: (2026)
RoboOmni: Proactive Robot Manipulation in Omni-modal Context
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing
von: Zhao, Haoyu, et al.
Veröffentlicht: (2023)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2023)
Learning to Rank Patches for Unbiased Image Redundancy Reduction
von: Luo, Yang, et al.
Veröffentlicht: (2024)
von: Luo, Yang, et al.
Veröffentlicht: (2024)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
von: Lan, Xiaohan, et al.
Veröffentlicht: (2024)
von: Lan, Xiaohan, et al.
Veröffentlicht: (2024)
Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval
von: Zou, Zichen, et al.
Veröffentlicht: (2026)
von: Zou, Zichen, et al.
Veröffentlicht: (2026)
GeoGS3D: Single-view 3D Reconstruction via Geometric-aware Diffusion Model and Gaussian Splatting
von: Feng, Qijun, et al.
Veröffentlicht: (2024)
von: Feng, Qijun, et al.
Veröffentlicht: (2024)
Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation
von: Tu, Shuyuan, et al.
Veröffentlicht: (2026)
von: Tu, Shuyuan, et al.
Veröffentlicht: (2026)
Adaptive Retention & Correction: Test-Time Training for Continual Learning
von: Chen, Haoran, et al.
Veröffentlicht: (2024)
von: Chen, Haoran, et al.
Veröffentlicht: (2024)
GenVidBench: A 6-Million Benchmark for AI-Generated Video Detection
von: Ni, Zhenliang, et al.
Veröffentlicht: (2025)
von: Ni, Zhenliang, et al.
Veröffentlicht: (2025)
CoMP: Continual Multimodal Pre-training for Vision Foundation Models
von: Chen, Yitong, et al.
Veröffentlicht: (2025)
von: Chen, Yitong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
OmniTracker: Unifying Object Tracking by Tracking-with-Detection
von: Wang, Junke, et al.
Veröffentlicht: (2023) -
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
von: Wang, Junke, et al.
Veröffentlicht: (2024) -
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
von: Shi, Jiapeng, et al.
Veröffentlicht: (2026) -
FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
von: Xie, Yiweng, et al.
Veröffentlicht: (2026) -
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
von: He, Zhihao, et al.
Veröffentlicht: (2026)