Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Simon, Marcel, Kim, Tae-Ho, Yeom, Seul-Ki |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniForm: A Reuse Attention Mechanism Optimized for Efficient Vision Transformers on Edge Devices
by: Yeom, Seul-Ki, et al.
Published: (2024)
by: Yeom, Seul-Ki, et al.
Published: (2024)
VINO: Video-driven Invariance for Non-contextual Objects via Structural Prior Guided De-contextualization
by: Yeom, Seul-Ki, et al.
Published: (2026)
by: Yeom, Seul-Ki, et al.
Published: (2026)
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
by: Kim, Manjin, et al.
Published: (2026)
by: Kim, Manjin, et al.
Published: (2026)
Co-learning Single-Step Diffusion Upsampler and Downsampler with Two Discriminators and Distillation
by: Kim, Sohwi, et al.
Published: (2024)
by: Kim, Sohwi, et al.
Published: (2024)
Single-Step Bidirectional Unpaired Image Translation Using Implicit Bridge Consistency Distillation
by: Lee, Suhyeon, et al.
Published: (2025)
by: Lee, Suhyeon, et al.
Published: (2025)
UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
by: Liu, Ropeway, et al.
Published: (2025)
by: Liu, Ropeway, et al.
Published: (2025)
Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation
by: Hassan, Mariam, et al.
Published: (2026)
by: Hassan, Mariam, et al.
Published: (2026)
Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility
by: Hao, Yutong, et al.
Published: (2025)
by: Hao, Yutong, et al.
Published: (2025)
Real2Sim in HOI: Toward Physically Plausible HOI Reconstruction from Monocular Videos
by: Zhao, Yubo, et al.
Published: (2026)
by: Zhao, Yubo, et al.
Published: (2026)
Self-Distilled Masked Auto-Encoders are Efficient Video Anomaly Detectors
by: Ristea, Nicolae-Catalin, et al.
Published: (2023)
by: Ristea, Nicolae-Catalin, et al.
Published: (2023)
Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image Generation
by: Dao, Quan, et al.
Published: (2024)
by: Dao, Quan, et al.
Published: (2024)
VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior
by: Yang, Xindi, et al.
Published: (2025)
by: Yang, Xindi, et al.
Published: (2025)
Activation Quantization of Vision Encoders Needs Prefixing Registers
by: Kim, Seunghyeon, et al.
Published: (2025)
by: Kim, Seunghyeon, et al.
Published: (2025)
PhyCAGE: Physically Plausible Compositional 3D Asset Generation from a Single Image
by: Yan, Han, et al.
Published: (2024)
by: Yan, Han, et al.
Published: (2024)
One-Step Diffusion-based Real-World Image Super-Resolution with Visual Perception Distillation
by: Wu, Xue, et al.
Published: (2025)
by: Wu, Xue, et al.
Published: (2025)
Single Trajectory Distillation for Accelerating Image and Video Style Transfer
by: Xu, Sijie, et al.
Published: (2024)
by: Xu, Sijie, et al.
Published: (2024)
MMPhysVideo: Scaling Physical Plausibility in Video Generation via Joint Multimodal Modeling
by: Lin, Shubo, et al.
Published: (2026)
by: Lin, Shubo, et al.
Published: (2026)
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
Chain of Event-Centric Causal Thought for Physically Plausible Video Generation
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
BiTT: Bi-directional Texture Reconstruction of Interacting Two Hands from a Single Image
by: Kim, Minje, et al.
Published: (2024)
by: Kim, Minje, et al.
Published: (2024)
PhySIC: Physically Plausible 3D Human-Scene Interaction and Contact from a Single Image
by: Muralidhar, Pradyumna Yalandur, et al.
Published: (2025)
by: Muralidhar, Pradyumna Yalandur, et al.
Published: (2025)
TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in CLIP to Alleviate Single Tag Bias
by: Jo, Sanghyun, et al.
Published: (2024)
by: Jo, Sanghyun, et al.
Published: (2024)
From Generated Human Videos to Physically Plausible Robot Trajectories
by: Ni, James, et al.
Published: (2025)
by: Ni, James, et al.
Published: (2025)
Distilled Pooling Transformer Encoder for Efficient Realistic Image Dehazing
by: Tran, Le-Anh, et al.
Published: (2024)
by: Tran, Le-Anh, et al.
Published: (2024)
OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance
by: Wang, Cong, et al.
Published: (2026)
by: Wang, Cong, et al.
Published: (2026)
Towards Anatomically Plausible Human Image Generation via Synthetic Localized Preferences
by: Li, Bao, et al.
Published: (2026)
by: Li, Bao, et al.
Published: (2026)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
by: Maaz, Muhammad, et al.
Published: (2024)
by: Maaz, Muhammad, et al.
Published: (2024)
A Self-Supervised Approach on Motion Calibration for Enhancing Physical Plausibility in Text-to-Motion
by: Shim, Gahyeon, et al.
Published: (2026)
by: Shim, Gahyeon, et al.
Published: (2026)
MagicDistillation: Weak-to-Strong Video Distillation for Large-Scale Few-Step Synthesis
by: Shao, Shitong, et al.
Published: (2025)
by: Shao, Shitong, et al.
Published: (2025)
Improving Human Motion Plausibility with Body Momentum
by: Nguyen, Ha Linh, et al.
Published: (2025)
by: Nguyen, Ha Linh, et al.
Published: (2025)
Self-Improving 4D Perception via Self-Distillation
by: Huang, Nan, et al.
Published: (2026)
by: Huang, Nan, et al.
Published: (2026)
Towards One-step Causal Video Generation via Adversarial Self-Distillation
by: Yang, Yongqi, et al.
Published: (2025)
by: Yang, Yongqi, et al.
Published: (2025)
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
by: Jiang, Dengyang, et al.
Published: (2026)
by: Jiang, Dengyang, et al.
Published: (2026)
Prompt Augmentation for Self-supervised Text-guided Image Manipulation
by: Bodur, Rumeysa, et al.
Published: (2024)
by: Bodur, Rumeysa, et al.
Published: (2024)
As-Plausible-As-Possible: Plausibility-Aware Mesh Deformation Using 2D Diffusion Priors
by: Yoo, Seungwoo, et al.
Published: (2023)
by: Yoo, Seungwoo, et al.
Published: (2023)
Efficient Universal Perception Encoder
by: Zhu, Chenchen, et al.
Published: (2026)
by: Zhu, Chenchen, et al.
Published: (2026)
Distill Video Datasets into Images
by: Zhao, Zhenghao, et al.
Published: (2025)
by: Zhao, Zhenghao, et al.
Published: (2025)
UM-Depth : Uncertainty Masked Self-Supervised Monocular Depth Estimation with Visual Odometry
by: Um, Tae-Wook, et al.
Published: (2025)
by: Um, Tae-Wook, et al.
Published: (2025)
THOM: Generating Physically Plausible Hand-Object Meshes From Text
by: Jeong, Uyoung, et al.
Published: (2026)
by: Jeong, Uyoung, et al.
Published: (2026)
Similar Items
-
UniForm: A Reuse Attention Mechanism Optimized for Efficient Vision Transformers on Edge Devices
by: Yeom, Seul-Ki, et al.
Published: (2024) -
VINO: Video-driven Invariance for Non-contextual Objects via Structural Prior Guided De-contextualization
by: Yeom, Seul-Ki, et al.
Published: (2026) -
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
by: Kim, Manjin, et al.
Published: (2026) -
Co-learning Single-Step Diffusion Upsampler and Downsampler with Two Discriminators and Distillation
by: Kim, Sohwi, et al.
Published: (2024) -
Single-Step Bidirectional Unpaired Image Translation Using Implicit Bridge Consistency Distillation
by: Lee, Suhyeon, et al.
Published: (2025)