Temporal In-Context Fine-Tuning with Temporal Reasoning for Versatile Control of Video Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Kinam, Hyung, Junha, Choo, Jaegul |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
by: Hyung, Junha, et al.
Published: (2024)
by: Hyung, Junha, et al.
Published: (2024)
Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models
by: Hwang, Sungwon, et al.
Published: (2025)
by: Hwang, Sungwon, et al.
Published: (2025)
EgoX: Egocentric Video Generation from a Single Exocentric Video
by: Kang, Taewoong, et al.
Published: (2025)
by: Kang, Taewoong, et al.
Published: (2025)
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
by: Jin, Hoiyeong, et al.
Published: (2025)
by: Jin, Hoiyeong, et al.
Published: (2025)
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
by: Kim, Min-Jung, et al.
Published: (2025)
by: Kim, Min-Jung, et al.
Published: (2025)
MagiCapture: High-Resolution Multi-Concept Portrait Customization
by: Hyung, Junha, et al.
Published: (2023)
by: Hyung, Junha, et al.
Published: (2023)
SelfSwapper: Self-Supervised Face Swapping via Shape Agnostic Masked AutoEncoder
by: Lee, Jaeseong, et al.
Published: (2024)
by: Lee, Jaeseong, et al.
Published: (2024)
TCAN: Animating Human Images with Temporally Consistent Pose Guidance using Diffusion Models
by: Kim, Jeongho, et al.
Published: (2024)
by: Kim, Jeongho, et al.
Published: (2024)
Is user feedback always informative? Retrieval Latent Defending for Semi-Supervised Domain Adaptation without Source Data
by: Song, Junha, et al.
Published: (2024)
by: Song, Junha, et al.
Published: (2024)
Effective Rank Analysis and Regularization for Enhanced 3D Gaussian Splatting
by: Hyung, Junha, et al.
Published: (2024)
by: Hyung, Junha, et al.
Published: (2024)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
by: Song, Junha, et al.
Published: (2026)
by: Song, Junha, et al.
Published: (2026)
Good Noise Makes Good Edits: A Training-Free Diffusion-Based Video Editing with Image and Text Prompts
by: Choi, Saemee, et al.
Published: (2025)
by: Choi, Saemee, et al.
Published: (2025)
RL makes MLLMs see better than SFT
by: Song, Junha, et al.
Published: (2025)
by: Song, Junha, et al.
Published: (2025)
Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention
by: Jo, Kyungmin, et al.
Published: (2025)
by: Jo, Kyungmin, et al.
Published: (2025)
What to Preserve and What to Transfer: Faithful, Identity-Preserving Diffusion-based Hairstyle Transfer
by: Chung, Chaeyeon, et al.
Published: (2024)
by: Chung, Chaeyeon, et al.
Published: (2024)
TV-LiVE: Training-Free, Text-Guided Video Editing via Layer Informed Vitality Exploitation
by: Kim, Min-Jung, et al.
Published: (2025)
by: Kim, Min-Jung, et al.
Published: (2025)
Vector Prism: Animating Vector Graphics by Stratifying Semantic Structure
by: Yun, Jooyeol, et al.
Published: (2025)
by: Yun, Jooyeol, et al.
Published: (2025)
Skip-and-Play: Depth-Driven Pose-Preserved Image Generation for Any Objects
by: Jo, Kyungmin, et al.
Published: (2024)
by: Jo, Kyungmin, et al.
Published: (2024)
Scaling Up Personalized Image Aesthetic Assessment via Task Vector Customization
by: Yun, Jooyeol, et al.
Published: (2024)
by: Yun, Jooyeol, et al.
Published: (2024)
MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
by: Song, Junha, et al.
Published: (2025)
by: Song, Junha, et al.
Published: (2025)
Memory-Efficient Fine-Tuning Diffusion Transformers via Dynamic Patch Sampling and Block Skipping
by: Park, Sunghyun, et al.
Published: (2026)
by: Park, Sunghyun, et al.
Published: (2026)
Towards Calibrated Robust Fine-Tuning of Vision-Language Models
by: Oh, Changdae, et al.
Published: (2023)
by: Oh, Changdae, et al.
Published: (2023)
Enabling Region-Specific Control via Lassos in Point-Based Colorization
by: Lee, Sanghyeon, et al.
Published: (2024)
by: Lee, Sanghyeon, et al.
Published: (2024)
Bones Can't Be Triangles: Accurate and Efficient Vertebrae Keypoint Estimation through Collaborative Error Revision
by: Kim, Jinhee, et al.
Published: (2024)
by: Kim, Jinhee, et al.
Published: (2024)
SurFhead: Affine Rig Blending for Geometrically Accurate 2D Gaussian Surfel Head Avatars
by: Lee, Jaeseong, et al.
Published: (2024)
by: Lee, Jaeseong, et al.
Published: (2024)
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
by: Qian, Long, et al.
Published: (2024)
by: Qian, Long, et al.
Published: (2024)
Image Diffusion Models Exhibit Emergent Temporal Propagation in Videos
by: Kim, Youngseo, et al.
Published: (2025)
by: Kim, Youngseo, et al.
Published: (2025)
DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
by: Wu, Ziyi, et al.
Published: (2025)
by: Wu, Ziyi, et al.
Published: (2025)
OPRO: Orthogonal Panel-Relative Operators for Panel-Aware In-Context Image Generation
by: Lee, Sanghyeon, et al.
Published: (2026)
by: Lee, Sanghyeon, et al.
Published: (2026)
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
by: Fateh, Fawad Javed, et al.
Published: (2024)
by: Fateh, Fawad Javed, et al.
Published: (2024)
Temporal Gains, Spatial Costs: Revisiting Video Fine-Tuning in Multimodal Large Language Models
by: Zhang, Linghao, et al.
Published: (2026)
by: Zhang, Linghao, et al.
Published: (2026)
PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask
by: Kim, Jeongho, et al.
Published: (2024)
by: Kim, Jeongho, et al.
Published: (2024)
Training Spatial-Frequency Visual Prompts and Probabilistic Clusters for Accurate Black-Box Transfer Learning
by: Cho, Wonwoo, et al.
Published: (2024)
by: Cho, Wonwoo, et al.
Published: (2024)
SphereDiff: Tuning-free 360° Static and Dynamic Panorama Generation via Spherical Latent Representation
by: Park, Minho, et al.
Published: (2025)
by: Park, Minho, et al.
Published: (2025)
GaussianMotion: End-to-End Learning of Animatable Gaussian Avatars with Pose Guidance from Text
by: Shim, Gyumin, et al.
Published: (2025)
by: Shim, Gyumin, et al.
Published: (2025)
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models
by: Souza, Rafael, et al.
Published: (2024)
by: Souza, Rafael, et al.
Published: (2024)
Emergent Temporal Correspondences from Video Diffusion Transformers
by: Nam, Jisu, et al.
Published: (2025)
by: Nam, Jisu, et al.
Published: (2025)
Efficient and Versatile Robust Fine-Tuning of Zero-shot Models
by: Kim, Sungyeon, et al.
Published: (2024)
by: Kim, Sungyeon, et al.
Published: (2024)
ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models
by: Ko, Dohwan, et al.
Published: (2025)
by: Ko, Dohwan, et al.
Published: (2025)
TPDiff: Temporal Pyramid Video Diffusion Model
by: Ran, Lingmin, et al.
Published: (2025)
by: Ran, Lingmin, et al.
Published: (2025)
Similar Items
-
Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
by: Hyung, Junha, et al.
Published: (2024) -
Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models
by: Hwang, Sungwon, et al.
Published: (2025) -
EgoX: Egocentric Video Generation from a Single Exocentric Video
by: Kang, Taewoong, et al.
Published: (2025) -
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
by: Jin, Hoiyeong, et al.
Published: (2025) -
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
by: Kim, Min-Jung, et al.
Published: (2025)