Saved in:
| Main Authors: | Kirillov, Ivan, Parkhomenko, Denis, Chernyshev, Kirill, Pletnev, Alexander, Shi, Yibo, Lin, Kai, Babin, Dmitry |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2406.16544 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Alchemist: Turning Public Text-to-Image Data into Generative Gold
by: Startsev, Valerii, et al.
Published: (2025)
by: Startsev, Valerii, et al.
Published: (2025)
UCVC: A Unified Contextual Video Compression Framework with Joint P-frame and B-frame Coding
by: Yang, Jiayu, et al.
Published: (2024)
by: Yang, Jiayu, et al.
Published: (2024)
Motion Free B-frame Coding for Neural Video Compression
by: Nguyen, Van Thang
Published: (2024)
by: Nguyen, Van Thang
Published: (2024)
Differentiable Rendering with Reparameterized Volume Sampling
by: Morozov, Nikita, et al.
Published: (2023)
by: Morozov, Nikita, et al.
Published: (2023)
Deepfake detection in videos with multiple faces using geometric-fakeness features
by: Vyshegorodtsev, Kirill, et al.
Published: (2024)
by: Vyshegorodtsev, Kirill, et al.
Published: (2024)
Hierarchical Memory for Long Video QA
by: Wang, Yiqin, et al.
Published: (2024)
by: Wang, Yiqin, et al.
Published: (2024)
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
by: Zou, Kai, et al.
Published: (2026)
by: Zou, Kai, et al.
Published: (2026)
Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
by: Arkhipkin, Vladimir, et al.
Published: (2025)
by: Arkhipkin, Vladimir, et al.
Published: (2025)
PRISM-Loc: a Lightweight Long-range LiDAR Localization in Urban Environments with Topological Maps
by: Muravyev, Kirill, et al.
Published: (2025)
by: Muravyev, Kirill, et al.
Published: (2025)
IBVC: Interpolation-driven B-frame Video Compression
by: Xu, Chenming, et al.
Published: (2023)
by: Xu, Chenming, et al.
Published: (2023)
CasTex: Cascaded Text-to-Texture Synthesis via Explicit Texture Maps and Physically-Based Shading
by: Aliev, Mishan, et al.
Published: (2025)
by: Aliev, Mishan, et al.
Published: (2025)
Regularized Distribution Matching Distillation for One-step Unpaired Image-to-Image Translation
by: Rakitin, Denis, et al.
Published: (2024)
by: Rakitin, Denis, et al.
Published: (2024)
Generative inpainting of incomplete Euclidean distance matrices of trajectories generated by a fractional Brownian motion
by: Lobashev, Alexander, et al.
Published: (2024)
by: Lobashev, Alexander, et al.
Published: (2024)
Global Motion Understanding in Large-Scale Video Object Segmentation
by: Fedynyak, Volodymyr, et al.
Published: (2024)
by: Fedynyak, Volodymyr, et al.
Published: (2024)
MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
by: Wu, Weijia, et al.
Published: (2024)
by: Wu, Weijia, et al.
Published: (2024)
Neural B-frame Video Compression with Bi-directional Reference Harmonization
by: Liu, Yuxi, et al.
Published: (2025)
by: Liu, Yuxi, et al.
Published: (2025)
WVSC: Wireless Video Semantic Communication with Multi-frame Compensation
by: Xie, Bingyan, et al.
Published: (2025)
by: Xie, Bingyan, et al.
Published: (2025)
Extending Video Masked Autoencoders to 128 frames
by: Gundavarapu, Nitesh Bharadwaj, et al.
Published: (2024)
by: Gundavarapu, Nitesh Bharadwaj, et al.
Published: (2024)
Hierarchical Patch Diffusion Models for High-Resolution Video Generation
by: Skorokhodov, Ivan, et al.
Published: (2024)
by: Skorokhodov, Ivan, et al.
Published: (2024)
HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding
by: Shi, Mengqi, et al.
Published: (2026)
by: Shi, Mengqi, et al.
Published: (2026)
Bridging the Gap Between Saliency Prediction and Image Quality Assessment
by: Alexey, Kirillov, et al.
Published: (2024)
by: Alexey, Kirillov, et al.
Published: (2024)
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
by: Yang, Xuyi, et al.
Published: (2025)
by: Yang, Xuyi, et al.
Published: (2025)
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
Post-Training Quantization for Diffusion Transformer via Hierarchical Timestep Grouping
by: Ding, Ning, et al.
Published: (2025)
by: Ding, Ning, et al.
Published: (2025)
ARC-Chapter: Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries
by: Pu, Junfu, et al.
Published: (2025)
by: Pu, Junfu, et al.
Published: (2025)
Variable-frame CNNLSTM for Breast Nodule Classification using Ultrasound Videos
by: Cui, Xiangxiang, et al.
Published: (2025)
by: Cui, Xiangxiang, et al.
Published: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
by: Fang, Xinyu, et al.
Published: (2024)
by: Fang, Xinyu, et al.
Published: (2024)
DeVOS: Flow-Guided Deformable Transformer for Video Object Segmentation
by: Fedynyak, Volodymyr, et al.
Published: (2024)
by: Fedynyak, Volodymyr, et al.
Published: (2024)
OmniRoam: World Wandering via Long-Horizon Panoramic Video Generation
by: Liu, Yuheng, et al.
Published: (2026)
by: Liu, Yuheng, et al.
Published: (2026)
End-to-End Facial Expression Detection in Long Videos
by: Fang, Yini, et al.
Published: (2025)
by: Fang, Yini, et al.
Published: (2025)
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
by: Li, Jungang, et al.
Published: (2024)
by: Li, Jungang, et al.
Published: (2024)
HAtt-Flow: Hierarchical Attention-Flow Mechanism for Group Activity Scene Graph Generation in Videos
by: Chappa, Naga VS Raviteja, et al.
Published: (2023)
by: Chappa, Naga VS Raviteja, et al.
Published: (2023)
LOGO: A Long-Form Video Dataset for Group Action Quality Assessment
by: Zhang, Shiyi, et al.
Published: (2024)
by: Zhang, Shiyi, et al.
Published: (2024)
MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
by: Jia, Weinan, et al.
Published: (2025)
by: Jia, Weinan, et al.
Published: (2025)
IQA-Adapter: Exploring Knowledge Transfer from Image Quality Assessment to Diffusion-based Generative Models
by: Abud, Khaled, et al.
Published: (2024)
by: Abud, Khaled, et al.
Published: (2024)
RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
by: Xun, Shuhang, et al.
Published: (2025)
by: Xun, Shuhang, et al.
Published: (2025)
Long Context Tuning for Video Generation
by: Guo, Yuwei, et al.
Published: (2025)
by: Guo, Yuwei, et al.
Published: (2025)
FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance
by: Feng, Jiasong, et al.
Published: (2024)
by: Feng, Jiasong, et al.
Published: (2024)
A Neural-network Enhanced Video Coding Framework beyond ECM
by: Zhao, Yanchen, et al.
Published: (2024)
by: Zhao, Yanchen, et al.
Published: (2024)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
by: Yin, Yufei, et al.
Published: (2025)
by: Yin, Yufei, et al.
Published: (2025)
Similar Items
-
Alchemist: Turning Public Text-to-Image Data into Generative Gold
by: Startsev, Valerii, et al.
Published: (2025) -
UCVC: A Unified Contextual Video Compression Framework with Joint P-frame and B-frame Coding
by: Yang, Jiayu, et al.
Published: (2024) -
Motion Free B-frame Coding for Neural Video Compression
by: Nguyen, Van Thang
Published: (2024) -
Differentiable Rendering with Reparameterized Volume Sampling
by: Morozov, Nikita, et al.
Published: (2023) -
Deepfake detection in videos with multiple faces using geometric-fakeness features
by: Vyshegorodtsev, Kirill, et al.
Published: (2024)