Saved in:
| Main Authors: | Sheng, Xihua, Zhang, Yingwen, Xu, Long, Wang, Shiqi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.03922 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Information-Theoretic Regularizer for Lossy Neural Image Compression
by: Zhang, Yingwen, et al.
Published: (2024)
by: Zhang, Yingwen, et al.
Published: (2024)
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026)
by: Zhen, Haoyu, et al.
Published: (2026)
Bi-Directional Deep Contextual Video Compression
by: Sheng, Xihua, et al.
Published: (2024)
by: Sheng, Xihua, et al.
Published: (2024)
Recent Advances of End-to-End Video Coding Technologies for AVS Standard Development
by: Sheng, Xihua, et al.
Published: (2026)
by: Sheng, Xihua, et al.
Published: (2026)
Fine-Grained Motion Compression and Selective Temporal Fusion for Neural B-Frame Video Coding
by: Sheng, Xihua, et al.
Published: (2025)
by: Sheng, Xihua, et al.
Published: (2025)
An End-to-End Framework for Video Multi-Person Pose Estimation
by: Wei, Zhihong
Published: (2025)
by: Wei, Zhihong
Published: (2025)
CADC: Content Adaptive Diffusion-Based Generative Image Compression
by: Sheng, Xihua, et al.
Published: (2026)
by: Sheng, Xihua, et al.
Published: (2026)
End-to-End Streaming Video Temporal Action Segmentation with Reinforce Learning
by: Zhang, Jinrong, et al.
Published: (2023)
by: Zhang, Jinrong, et al.
Published: (2023)
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
by: Wang, Linhan, et al.
Published: (2026)
by: Wang, Linhan, et al.
Published: (2026)
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
by: Zhang, Yaolun, et al.
Published: (2026)
by: Zhang, Yaolun, et al.
Published: (2026)
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
by: Zhang, Chen-Lin, et al.
Published: (2025)
by: Zhang, Chen-Lin, et al.
Published: (2025)
MoCha:End-to-End Video Character Replacement without Structural Guidance
by: Xu, Zhengbo, et al.
Published: (2026)
by: Xu, Zhengbo, et al.
Published: (2026)
End-to-End Dense Video Grounding via Parallel Regression
by: Shi, Fengyuan, et al.
Published: (2021)
by: Shi, Fengyuan, et al.
Published: (2021)
NVC-1B: A Large Neural Video Coding Model
by: Sheng, Xihua, et al.
Published: (2024)
by: Sheng, Xihua, et al.
Published: (2024)
GenAD: Generative End-to-End Autonomous Driving
by: Zheng, Wenzhao, et al.
Published: (2024)
by: Zheng, Wenzhao, et al.
Published: (2024)
End-to-End Facial Expression Detection in Long Videos
by: Fang, Yini, et al.
Published: (2025)
by: Fang, Yini, et al.
Published: (2025)
End-To-End Underwater Video Enhancement: Dataset and Model
by: Du, Dazhao, et al.
Published: (2024)
by: Du, Dazhao, et al.
Published: (2024)
Prediction and Reference Quality Adaptation for Learned Video Compression
by: Sheng, Xihua, et al.
Published: (2024)
by: Sheng, Xihua, et al.
Published: (2024)
End-to-End Unmixing with Material Prompts for Hyperspectral Object Tracking
by: Han, Xu, et al.
Published: (2026)
by: Han, Xu, et al.
Published: (2026)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
by: Shi, Yudi, et al.
Published: (2026)
by: Shi, Yudi, et al.
Published: (2026)
ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
by: Li, Yongkang, et al.
Published: (2025)
by: Li, Yongkang, et al.
Published: (2025)
ImmersePro: End-to-End Stereo Video Synthesis Via Implicit Disparity Learning
by: Shi, Jian, et al.
Published: (2024)
by: Shi, Jian, et al.
Published: (2024)
CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal
by: He, Qingdong, et al.
Published: (2026)
by: He, Qingdong, et al.
Published: (2026)
Matching Distance and Geometric Distribution Aided Learning Multiview Point Cloud Registration
by: Li, Shiqi, et al.
Published: (2025)
by: Li, Shiqi, et al.
Published: (2025)
End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer
by: Yu, Yonghui, et al.
Published: (2025)
by: Yu, Yonghui, et al.
Published: (2025)
From Category to Scenery: An End-to-End Framework for Multi-Person Human-Object Interaction Recognition in Videos
by: Qiao, Tanqiu, et al.
Published: (2024)
by: Qiao, Tanqiu, et al.
Published: (2024)
Spatial Decomposition and Temporal Fusion based Inter Prediction for Learned Video Compression
by: Sheng, Xihua, et al.
Published: (2024)
by: Sheng, Xihua, et al.
Published: (2024)
CoopTrack: Exploring End-to-End Learning for Efficient Cooperative Sequential Perception
by: Zhong, Jiaru, et al.
Published: (2025)
by: Zhong, Jiaru, et al.
Published: (2025)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
by: Zhang, Pingping, et al.
Published: (2024)
by: Zhang, Pingping, et al.
Published: (2024)
End-to-End Vision Tokenizer Tuning
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Interactive Face Video Coding: A Generative Compression Framework
by: Chen, Bolin, et al.
Published: (2023)
by: Chen, Bolin, et al.
Published: (2023)
Active Learning from Scene Embeddings for End-to-End Autonomous Driving
by: Jiang, Wenhao, et al.
Published: (2025)
by: Jiang, Wenhao, et al.
Published: (2025)
MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
by: Jia, Weinan, et al.
Published: (2025)
by: Jia, Weinan, et al.
Published: (2025)
On the Pitfalls of Batch Normalization for End-to-End Video Learning: A Study on Surgical Workflow Analysis
by: Rivoir, Dominik, et al.
Published: (2022)
by: Rivoir, Dominik, et al.
Published: (2022)
VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework
by: Huang, Donglin, et al.
Published: (2025)
by: Huang, Donglin, et al.
Published: (2025)
Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation
by: Ma, Enhui, et al.
Published: (2024)
by: Ma, Enhui, et al.
Published: (2024)
EVE: Towards End-to-End Video Subtitle Extraction with Vision-Language Models
by: Yu, Haiyang, et al.
Published: (2025)
by: Yu, Haiyang, et al.
Published: (2025)
Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives
by: Wang, Junli, et al.
Published: (2026)
by: Wang, Junli, et al.
Published: (2026)
Enhancing Traffic Safety with Parallel Dense Video Captioning for End-to-End Event Analysis
by: Shoman, Maged, et al.
Published: (2024)
by: Shoman, Maged, et al.
Published: (2024)
Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
Similar Items
-
An Information-Theoretic Regularizer for Lossy Neural Image Compression
by: Zhang, Yingwen, et al.
Published: (2024) -
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026) -
Bi-Directional Deep Contextual Video Compression
by: Sheng, Xihua, et al.
Published: (2024) -
Recent Advances of End-to-End Video Coding Technologies for AVS Standard Development
by: Sheng, Xihua, et al.
Published: (2026) -
Fine-Grained Motion Compression and Selective Temporal Fusion for Neural B-Frame Video Coding
by: Sheng, Xihua, et al.
Published: (2025)