Saved in:
| Main Authors: | Zhang, Qikang, Lei, Yingjie, Liu, Wei, Liu, Daochang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.13438 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Test Time Training for 4D Medical Image Interpolation
by: Zhang, Qikang, et al.
Published: (2025)
by: Zhang, Qikang, et al.
Published: (2025)
Multi-Level Heterogeneous Knowledge Transfer Network on Forward Scattering Center Model for Limited Samples SAR ATR
by: Zhao, Chenxi, et al.
Published: (2025)
by: Zhao, Chenxi, et al.
Published: (2025)
Surgical Triplet Recognition via Diffusion Model
by: Liu, Daochang, et al.
Published: (2024)
by: Liu, Daochang, et al.
Published: (2024)
Generative Physical AI in Vision: A Survey
by: Liu, Daochang, et al.
Published: (2025)
by: Liu, Daochang, et al.
Published: (2025)
CollagePrompt: A Benchmark for Budget-Friendly Visual Recognition with GPT-4V
by: Xu, Siyu, et al.
Published: (2024)
by: Xu, Siyu, et al.
Published: (2024)
Sample-efficient Unsupervised Policy Cloning from Ensemble Self-supervised Labeled Videos
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling
by: Yan, Jiebin, et al.
Published: (2025)
by: Yan, Jiebin, et al.
Published: (2025)
DraftAttention: Fast Video Diffusion via Low-Resolution Attention Guidance
by: Shen, Xuan, et al.
Published: (2025)
by: Shen, Xuan, et al.
Published: (2025)
Not All Tokens are Guided Equal: Improving Guidance in Visual Autoregressive Models
by: Nguyen, Ky Dan, et al.
Published: (2025)
by: Nguyen, Ky Dan, et al.
Published: (2025)
Eyes on Target: Gaze-Aware Object Detection in Egocentric Video
by: Lall, Vishakha, et al.
Published: (2025)
by: Lall, Vishakha, et al.
Published: (2025)
Spatia: Video Generation with Updatable Spatial Memory
by: Zhao, Jinjing, et al.
Published: (2025)
by: Zhao, Jinjing, et al.
Published: (2025)
EgoVLM: Policy Optimization for Egocentric Video Understanding
by: Vinod, Ashwin, et al.
Published: (2025)
by: Vinod, Ashwin, et al.
Published: (2025)
Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
by: Liu, Yuhan, et al.
Published: (2025)
by: Liu, Yuhan, et al.
Published: (2025)
On Equivariance and Fast Sampling in Video Diffusion Models Trained with Warped Noise
by: Liu, Chao, et al.
Published: (2025)
by: Liu, Chao, et al.
Published: (2025)
A Comprehensive Survey on Human Video Generation: Challenges, Methods, and Insights
by: Lei, Wentao, et al.
Published: (2024)
by: Lei, Wentao, et al.
Published: (2024)
Video-T1: Test-Time Scaling for Video Generation
by: Liu, Fangfu, et al.
Published: (2025)
by: Liu, Fangfu, et al.
Published: (2025)
Hyperbolic Contrastive Learning for Hierarchical 3D Point Cloud Embedding
by: Liu, Yingjie, et al.
Published: (2025)
by: Liu, Yingjie, et al.
Published: (2025)
Fine-gained Zero-shot Video Sampling
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
PMG: Progressive Motion Generation via Sparse Anchor Postures Curriculum Learning
by: Xi, Yingjie, et al.
Published: (2025)
by: Xi, Yingjie, et al.
Published: (2025)
Video Generation with Consistency Tuning
by: Wang, Chaoyi, et al.
Published: (2024)
by: Wang, Chaoyi, et al.
Published: (2024)
Video-Infinity: Distributed Long Video Generation
by: Tan, Zhenxiong, et al.
Published: (2024)
by: Tan, Zhenxiong, et al.
Published: (2024)
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning
by: Ou, Siqu, et al.
Published: (2025)
by: Ou, Siqu, et al.
Published: (2025)
DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
by: Jiang, Dongzhi, et al.
Published: (2025)
by: Jiang, Dongzhi, et al.
Published: (2025)
Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation
by: Li, Baoteng, et al.
Published: (2026)
by: Li, Baoteng, et al.
Published: (2026)
Targeted Downstream-Agnostic Attack
by: Lei, Zhuxin, et al.
Published: (2026)
by: Lei, Zhuxin, et al.
Published: (2026)
BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations
by: Feng, Weixi, et al.
Published: (2025)
by: Feng, Weixi, et al.
Published: (2025)
Flow-Based Generative Modeling for Optimizing Sampling Policies in Compressed Sensing Applications
by: Pavelkin, Roman, et al.
Published: (2026)
by: Pavelkin, Roman, et al.
Published: (2026)
Optical Flow Representation Alignment Mamba Diffusion Model for Medical Video Generation
by: Wang, Zhenbin, et al.
Published: (2024)
by: Wang, Zhenbin, et al.
Published: (2024)
KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding
by: Li, Zongyao, et al.
Published: (2025)
by: Li, Zongyao, et al.
Published: (2025)
One-Step Diffusion for Detail-Rich and Temporally Consistent Video Super-Resolution
by: Sun, Yujing, et al.
Published: (2025)
by: Sun, Yujing, et al.
Published: (2025)
RISE-Video: Can Video Generators Decode Implicit World Rules?
by: Liu, Mingxin, et al.
Published: (2026)
by: Liu, Mingxin, et al.
Published: (2026)
HARIVO: Harnessing Text-to-Image Models for Video Generation
by: Kwon, Mingi, et al.
Published: (2024)
by: Kwon, Mingi, et al.
Published: (2024)
AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
by: Li, Yiheng, et al.
Published: (2026)
by: Li, Yiheng, et al.
Published: (2026)
Consistent Video Editing as Flow-Driven Image-to-Video Generation
by: Wang, Ge, et al.
Published: (2025)
by: Wang, Ge, et al.
Published: (2025)
Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2025)
by: Gu, Xin, et al.
Published: (2025)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
Rectified SpaAttn: Revisiting Attention Sparsity for Efficient Video Generation
by: Liu, Xuewen, et al.
Published: (2025)
by: Liu, Xuewen, et al.
Published: (2025)
MAGI-1: Autoregressive Video Generation at Scale
by: ai, Sand., et al.
Published: (2025)
by: ai, Sand., et al.
Published: (2025)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
by: Liu, Wenqi, et al.
Published: (2026)
by: Liu, Wenqi, et al.
Published: (2026)
Similar Items
-
Test Time Training for 4D Medical Image Interpolation
by: Zhang, Qikang, et al.
Published: (2025) -
Multi-Level Heterogeneous Knowledge Transfer Network on Forward Scattering Center Model for Limited Samples SAR ATR
by: Zhao, Chenxi, et al.
Published: (2025) -
Surgical Triplet Recognition via Diffusion Model
by: Liu, Daochang, et al.
Published: (2024) -
Generative Physical AI in Vision: A Survey
by: Liu, Daochang, et al.
Published: (2025) -
CollagePrompt: A Benchmark for Budget-Friendly Visual Recognition with GPT-4V
by: Xu, Siyu, et al.
Published: (2024)