Cut2Next: Generating Next Shot via In-Context Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | He, Jingwen, Liu, Hongbo, Li, Jiajun, Huang, Ziqi, Qiao, Yu, Ouyang, Wanli, Liu, Ziwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
SalientFusion: Context-Aware Compositional Zero-Shot Food Recognition
by: Song, Jiajun, et al.
Published: (2025)
by: Song, Jiajun, et al.
Published: (2025)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
by: Ji, Longbin, et al.
Published: (2026)
by: Ji, Longbin, et al.
Published: (2026)
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
by: Liu, Hongbo, et al.
Published: (2025)
by: Liu, Hongbo, et al.
Published: (2025)
A Comprehensive Survey on 3D Content Generation
by: Liu, Jian, et al.
Published: (2024)
by: Liu, Jian, et al.
Published: (2024)
Playing with Transformer at 30+ FPS via Next-Frame Diffusion
by: Cheng, Xinle, et al.
Published: (2025)
by: Cheng, Xinle, et al.
Published: (2025)
Flow-GRPO: Training Flow Matching Models via Online RL
by: Liu, Jie, et al.
Published: (2025)
by: Liu, Jie, et al.
Published: (2025)
Video-GPT via Next Clip Diffusion
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
by: Zhang, Huichao, et al.
Published: (2026)
by: Zhang, Huichao, et al.
Published: (2026)
Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation
by: Inferix Team, et al.
Published: (2025)
by: Inferix Team, et al.
Published: (2025)
Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
by: Ren, Shuhuai, et al.
Published: (2025)
by: Ren, Shuhuai, et al.
Published: (2025)
Fostering Video Reasoning via Next-Event Prediction
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
by: Tian, Keyu, et al.
Published: (2024)
by: Tian, Keyu, et al.
Published: (2024)
ReCA: Multi-Shot Long Video Extrapolation via Recursive Context Allocation
by: Liu, Akide, et al.
Published: (2026)
by: Liu, Akide, et al.
Published: (2026)
Learning Multi-Modal Mobility Dynamics for Generalized Next Location Recommendation
by: Dai, Junshu, et al.
Published: (2025)
by: Dai, Junshu, et al.
Published: (2025)
Next Visual Granularity Generation
by: Wang, Yikai, et al.
Published: (2025)
by: Wang, Yikai, et al.
Published: (2025)
Semantic Surgery: Zero-Shot Concept Erasure in Diffusion Models
by: Xiong, Lexiang, et al.
Published: (2025)
by: Xiong, Lexiang, et al.
Published: (2025)
MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators
by: Zhang, Yaqi, et al.
Published: (2023)
by: Zhang, Yaqi, et al.
Published: (2023)
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
by: Zhuo, Le, et al.
Published: (2024)
by: Zhuo, Le, et al.
Published: (2024)
ContextDrag: Precise Drag-Based Image Editing via Context-Preserving Token Injection and Position-Aligned Attention
by: He, Huiguo, et al.
Published: (2025)
by: He, Huiguo, et al.
Published: (2025)
RealDPO: Real or Not Real, that is the Preference
by: Cheng, Guo, et al.
Published: (2025)
by: Cheng, Guo, et al.
Published: (2025)
Simulating the Visual World with Artificial Intelligence: A Roadmap
by: Yue, Jingtong, et al.
Published: (2025)
by: Yue, Jingtong, et al.
Published: (2025)
NeuRodin: A Two-stage Framework for High-Fidelity Neural Surface Reconstruction
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
Depth Any Video with Scalable Synthetic Data
by: Yang, Honghui, et al.
Published: (2024)
by: Yang, Honghui, et al.
Published: (2024)
Predicting the Next Action by Modeling the Abstract Goal
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
Fine-Tuning Next-Scale Visual Autoregressive Models with Group Relative Policy Optimization
by: Gallici, Matteo, et al.
Published: (2025)
by: Gallici, Matteo, et al.
Published: (2025)
Prompt Tuning with Soft Context Sharing for Vision-Language Models
by: Ding, Kun, et al.
Published: (2022)
by: Ding, Kun, et al.
Published: (2022)
Generalizable Object Re-Identification via Visual In-Context Prompting
by: Huang, Zhizhong, et al.
Published: (2025)
by: Huang, Zhizhong, et al.
Published: (2025)
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
by: Zhou, Chunting, et al.
Published: (2024)
by: Zhou, Chunting, et al.
Published: (2024)
VersusDebias: Universal Zero-Shot Debiasing for Text-to-Image Models via SLM-Based Prompt Engineering and Generative Adversary
by: Luo, Hanjun, et al.
Published: (2024)
by: Luo, Hanjun, et al.
Published: (2024)
Cut to the Chase: Training-free Multimodal Summarization via Chain-of-Events
by: You, Xiaoxing, et al.
Published: (2026)
by: You, Xiaoxing, et al.
Published: (2026)
Dense Video Captioning using Graph-based Sentence Summarization
by: Zhang, Zhiwang, et al.
Published: (2025)
by: Zhang, Zhiwang, et al.
Published: (2025)
Show, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization
by: Zhang, Zhiwang, et al.
Published: (2025)
by: Zhang, Zhiwang, et al.
Published: (2025)
Model Compression using Progressive Channel Pruning
by: Guo, Jinyang, et al.
Published: (2025)
by: Guo, Jinyang, et al.
Published: (2025)
Overcoming Semantic Dilution in Transformer-Based Next Frame Prediction
by: Nguyen, Hy, et al.
Published: (2025)
by: Nguyen, Hy, et al.
Published: (2025)
VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
by: Zheng, Dian, et al.
Published: (2025)
by: Zheng, Dian, et al.
Published: (2025)
Pix2Next: Leveraging Vision Foundation Models for RGB to NIR Image Translation
by: Jin, Youngwan, et al.
Published: (2024)
by: Jin, Youngwan, et al.
Published: (2024)
2DP-2MRC: 2-Dimensional Pointer-based Machine Reading Comprehension Method for Multimodal Moment Retrieval
by: He, Jiajun, et al.
Published: (2024)
by: He, Jiajun, et al.
Published: (2024)
EfficientFSL: Enhancing Few-Shot Classification via Query-Only Tuning in Vision Transformers
by: Liao, Wenwen, et al.
Published: (2026)
by: Liao, Wenwen, et al.
Published: (2026)
VEnhancer: Generative Space-Time Enhancement for Video Generation
by: He, Jingwen, et al.
Published: (2024)
by: He, Jingwen, et al.
Published: (2024)
Similar Items
-
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
by: Zhang, Fan, et al.
Published: (2024) -
SalientFusion: Context-Aware Compositional Zero-Shot Food Recognition
by: Song, Jiajun, et al.
Published: (2025) -
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
by: Ji, Longbin, et al.
Published: (2026) -
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
by: Liu, Hongbo, et al.
Published: (2025) -
A Comprehensive Survey on 3D Content Generation
by: Liu, Jian, et al.
Published: (2024)