Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Hwang, Sungwon, Jang, Hyojin, Kim, Kinam, Park, Minho, Choo, Jaegul |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Temporal In-Context Fine-Tuning with Temporal Reasoning for Versatile Control of Video Diffusion Models
by: Kim, Kinam, et al.
Published: (2025)
by: Kim, Kinam, et al.
Published: (2025)
SphereDiff: Tuning-free 360° Static and Dynamic Panorama Generation via Spherical Latent Representation
by: Park, Minho, et al.
Published: (2025)
by: Park, Minho, et al.
Published: (2025)
EgoX: Egocentric Video Generation from a Single Exocentric Video
by: Kang, Taewoong, et al.
Published: (2025)
by: Kang, Taewoong, et al.
Published: (2025)
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
by: Jin, Hoiyeong, et al.
Published: (2025)
by: Jin, Hoiyeong, et al.
Published: (2025)
Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
by: Hyung, Junha, et al.
Published: (2024)
by: Hyung, Junha, et al.
Published: (2024)
Good Noise Makes Good Edits: A Training-Free Diffusion-Based Video Editing with Image and Text Prompts
by: Choi, Saemee, et al.
Published: (2025)
by: Choi, Saemee, et al.
Published: (2025)
CA-LoRA: Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation
by: Park, Minho, et al.
Published: (2025)
by: Park, Minho, et al.
Published: (2025)
Regularized Training with Generated Datasets for Name-Only Transfer of Vision-Language Models
by: Park, Minho, et al.
Published: (2024)
by: Park, Minho, et al.
Published: (2024)
Zero-Shot Head Swapping in Real-World Scenarios
by: Kang, Taewoong, et al.
Published: (2025)
by: Kang, Taewoong, et al.
Published: (2025)
VEGS: View Extrapolation of Urban Scenes in 3D Gaussian Splatting using Learned Priors
by: Hwang, Sungwon, et al.
Published: (2024)
by: Hwang, Sungwon, et al.
Published: (2024)
SurFhead: Affine Rig Blending for Geometrically Accurate 2D Gaussian Surfel Head Avatars
by: Lee, Jaeseong, et al.
Published: (2024)
by: Lee, Jaeseong, et al.
Published: (2024)
Effective Rank Analysis and Regularization for Enhanced 3D Gaussian Splatting
by: Hyung, Junha, et al.
Published: (2024)
by: Hyung, Junha, et al.
Published: (2024)
What to Preserve and What to Transfer: Faithful, Identity-Preserving Diffusion-based Hairstyle Transfer
by: Chung, Chaeyeon, et al.
Published: (2024)
by: Chung, Chaeyeon, et al.
Published: (2024)
Memory-Efficient Fine-Tuning Diffusion Transformers via Dynamic Patch Sampling and Block Skipping
by: Park, Sunghyun, et al.
Published: (2026)
by: Park, Sunghyun, et al.
Published: (2026)
Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention
by: Jo, Kyungmin, et al.
Published: (2025)
by: Jo, Kyungmin, et al.
Published: (2025)
TV-LiVE: Training-Free, Text-Guided Video Editing via Layer Informed Vitality Exploitation
by: Kim, Min-Jung, et al.
Published: (2025)
by: Kim, Min-Jung, et al.
Published: (2025)
DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion
by: Hwang, Geunmin, et al.
Published: (2025)
by: Hwang, Geunmin, et al.
Published: (2025)
TCAN: Animating Human Images with Temporally Consistent Pose Guidance using Diffusion Models
by: Kim, Jeongho, et al.
Published: (2024)
by: Kim, Jeongho, et al.
Published: (2024)
Vector Prism: Animating Vector Graphics by Stratifying Semantic Structure
by: Yun, Jooyeol, et al.
Published: (2025)
by: Yun, Jooyeol, et al.
Published: (2025)
Skip-and-Play: Depth-Driven Pose-Preserved Image Generation for Any Objects
by: Jo, Kyungmin, et al.
Published: (2024)
by: Jo, Kyungmin, et al.
Published: (2024)
Scaling Up Personalized Image Aesthetic Assessment via Task Vector Customization
by: Yun, Jooyeol, et al.
Published: (2024)
by: Yun, Jooyeol, et al.
Published: (2024)
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
by: Jang, Sangwon, et al.
Published: (2025)
by: Jang, Sangwon, et al.
Published: (2025)
Frame-wise Conditioning Adaptation for Fine-Tuning Diffusion Models in Text-to-Video Prediction
by: Liu, Zheyuan, et al.
Published: (2025)
by: Liu, Zheyuan, et al.
Published: (2025)
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
by: Kim, Min-Jung, et al.
Published: (2025)
by: Kim, Min-Jung, et al.
Published: (2025)
Towards Calibrated Robust Fine-Tuning of Vision-Language Models
by: Oh, Changdae, et al.
Published: (2023)
by: Oh, Changdae, et al.
Published: (2023)
SHIFT: Motion Alignment in Video Diffusion Models with Adversarial Hybrid Fine-Tuning
by: Ye, Xi, et al.
Published: (2026)
by: Ye, Xi, et al.
Published: (2026)
Bones Can't Be Triangles: Accurate and Efficient Vertebrae Keypoint Estimation through Collaborative Error Revision
by: Kim, Jinhee, et al.
Published: (2024)
by: Kim, Jinhee, et al.
Published: (2024)
PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask
by: Kim, Jeongho, et al.
Published: (2024)
by: Kim, Jeongho, et al.
Published: (2024)
Enhancing Intrinsic Features for Debiasing via Investigating Class-Discerning Common Attributes in Bias-Contrastive Pair
by: Park, Jeonghoon, et al.
Published: (2024)
by: Park, Jeonghoon, et al.
Published: (2024)
From Wardrobe to Canvas: Wardrobe Polyptych LoRA for Part-level Controllable Human Image Generation
by: Kim, Jeongho, et al.
Published: (2025)
by: Kim, Jeongho, et al.
Published: (2025)
Training Spatial-Frequency Visual Prompts and Probabilistic Clusters for Accurate Black-Box Transfer Learning
by: Cho, Wonwoo, et al.
Published: (2024)
by: Cho, Wonwoo, et al.
Published: (2024)
AHS: Adaptive Head Synthesis via Synthetic Data Augmentations
by: Kang, Taewoong, et al.
Published: (2026)
by: Kang, Taewoong, et al.
Published: (2026)
GaussianMotion: End-to-End Learning of Animatable Gaussian Avatars with Pose Guidance from Text
by: Shim, Gyumin, et al.
Published: (2025)
by: Shim, Gyumin, et al.
Published: (2025)
Generalizable Disaster Damage Assessment via Change Detection with Vision Foundation Model
by: Ahn, Kyeongjin, et al.
Published: (2024)
by: Ahn, Kyeongjin, et al.
Published: (2024)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
by: Hwang, Dongyoon, et al.
Published: (2024)
by: Hwang, Dongyoon, et al.
Published: (2024)
Event-Based Video Frame Interpolation With Cross-Modal Asymmetric Bidirectional Motion Fields
by: Kim, Taewoo, et al.
Published: (2025)
by: Kim, Taewoo, et al.
Published: (2025)
Enabling Region-Specific Control via Lassos in Point-Based Colorization
by: Lee, Sanghyeon, et al.
Published: (2024)
by: Lee, Sanghyeon, et al.
Published: (2024)
Investigating Pre-Training Objectives for Generalization in Vision-Based Reinforcement Learning
by: Kim, Donghu, et al.
Published: (2024)
by: Kim, Donghu, et al.
Published: (2024)
Advancing Cross-Domain Generalizability in Face Anti-Spoofing: Insights, Design, and Metrics
by: Kim, Hyojin, et al.
Published: (2024)
by: Kim, Hyojin, et al.
Published: (2024)
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
by: Kim, Jeonghyeon, et al.
Published: (2025)
by: Kim, Jeonghyeon, et al.
Published: (2025)
Similar Items
-
Temporal In-Context Fine-Tuning with Temporal Reasoning for Versatile Control of Video Diffusion Models
by: Kim, Kinam, et al.
Published: (2025) -
SphereDiff: Tuning-free 360° Static and Dynamic Panorama Generation via Spherical Latent Representation
by: Park, Minho, et al.
Published: (2025) -
EgoX: Egocentric Video Generation from a Single Exocentric Video
by: Kang, Taewoong, et al.
Published: (2025) -
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
by: Jin, Hoiyeong, et al.
Published: (2025) -
Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling
by: Hyung, Junha, et al.
Published: (2024)