PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse Motion
Fuente:
arXiv
Saved in:
| Main Authors: | Choi, Jaehyun, Hur, Jiwan, Han, Gyojin, Yu, Jaemyung, Kim, Junmo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DAM: Domain-Aware Module for Multi-Domain Dataset Condensation
by: Choi, Jaehyun, et al.
Published: (2025)
by: Choi, Jaehyun, et al.
Published: (2025)
Frequency-Aware Token Reduction for Efficient Vision Transformer
by: Lee, Dong-Jae, et al.
Published: (2025)
by: Lee, Dong-Jae, et al.
Published: (2025)
Learning Neural Deformation Representation for 4D Dynamic Shape Generation
by: Han, Gyojin, et al.
Published: (2026)
by: Han, Gyojin, et al.
Published: (2026)
Cross-Axis Feature Fusion with Joint-Wise Motion Difference Prediction for Text-Based 3D Human Motion Editing
by: Han, Gyojin, et al.
Published: (2026)
by: Han, Gyojin, et al.
Published: (2026)
Self-supervised Transformation Learning for Equivariant Representations
by: Yu, Jaemyung, et al.
Published: (2025)
by: Yu, Jaemyung, et al.
Published: (2025)
Unlocking the Capabilities of Masked Generative Models for Image Synthesis via Self-Guidance
by: Hur, Jiwan, et al.
Published: (2024)
by: Hur, Jiwan, et al.
Published: (2024)
DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization
by: Lee, Dongyeun, et al.
Published: (2025)
by: Lee, Dongyeun, et al.
Published: (2025)
Inlier-Centric Post-Training Quantization for Object Detection Models
by: Kim, Minsu, et al.
Published: (2026)
by: Kim, Minsu, et al.
Published: (2026)
IMSE: Intrinsic Mixture of Spectral Experts Fine-tuning for Test-Time Adaptation
by: Baek, Sunghyun, et al.
Published: (2026)
by: Baek, Sunghyun, et al.
Published: (2026)
AH-OCDA: Amplitude-based Curriculum Learning and Hopfield Segmentation Model for Open Compound Domain Adaptation
by: Choi, Jaehyun, et al.
Published: (2024)
by: Choi, Jaehyun, et al.
Published: (2024)
SSG: Scaled Spatial Guidance for Multi-Scale Visual Autoregressive Generation
by: Shin, Youngwoo, et al.
Published: (2026)
by: Shin, Youngwoo, et al.
Published: (2026)
Beyond Semantics: Disentangling Information Scope in Sparse Autoencoders for CLIP
by: Ro, Yusung, et al.
Published: (2026)
by: Ro, Yusung, et al.
Published: (2026)
PRISM: Diversifying Dataset Distillation by Decoupling Architectural Priors
by: Moser, Brian B., et al.
Published: (2025)
by: Moser, Brian B., et al.
Published: (2025)
CVA: Context-aware Video-text Alignment for Video Temporal Grounding
by: Moon, Sungho, et al.
Published: (2026)
by: Moon, Sungho, et al.
Published: (2026)
Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!
by: Chung, Jiwan, et al.
Published: (2024)
by: Chung, Jiwan, et al.
Published: (2024)
PISCO: Precise Video Instance Insertion with Sparse Control
by: Gao, Xiangbo, et al.
Published: (2026)
by: Gao, Xiangbo, et al.
Published: (2026)
Dataset Condensation with Color Compensation
by: Wu, Huyu, et al.
Published: (2025)
by: Wu, Huyu, et al.
Published: (2025)
Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
by: Kwon, Minchan, et al.
Published: (2026)
by: Kwon, Minchan, et al.
Published: (2026)
Elucidating the Design Space of Dataset Condensation
by: Shao, Shitong, et al.
Published: (2024)
by: Shao, Shitong, et al.
Published: (2024)
Dataset Condensation with Latent Quantile Matching
by: Wei, Wei, et al.
Published: (2024)
by: Wei, Wei, et al.
Published: (2024)
B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding
by: Choi, Changho, et al.
Published: (2025)
by: Choi, Changho, et al.
Published: (2025)
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
by: Jin, Hoiyeong, et al.
Published: (2025)
by: Jin, Hoiyeong, et al.
Published: (2025)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
by: Li, Quanhao, et al.
Published: (2025)
by: Li, Quanhao, et al.
Published: (2025)
SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion
by: Chen, Xinyu, et al.
Published: (2026)
by: Chen, Xinyu, et al.
Published: (2026)
PRISM: Distributed Inference for Foundation Models at Edge
by: Qazi, Muhammad Azlan, et al.
Published: (2025)
by: Qazi, Muhammad Azlan, et al.
Published: (2025)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
by: Kim, Youngmin, et al.
Published: (2025)
by: Kim, Youngmin, et al.
Published: (2025)
Global Geometry Is Not Enough for Vision Representations
by: Chung, Jiwan, et al.
Published: (2026)
by: Chung, Jiwan, et al.
Published: (2026)
MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching
by: Wu, Yen-Siang, et al.
Published: (2025)
by: Wu, Yen-Siang, et al.
Published: (2025)
ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic Object
by: Zhang, Chenshuang, et al.
Published: (2024)
by: Zhang, Chenshuang, et al.
Published: (2024)
Text-to-image Diffusion Models in Generative AI: A Survey
by: Zhang, Chenshuang, et al.
Published: (2023)
by: Zhang, Chenshuang, et al.
Published: (2023)
Progressive Fourier Neural Representation for Sequential Video Compilation
by: Kang, Haeyong, et al.
Published: (2023)
by: Kang, Haeyong, et al.
Published: (2023)
Motion meets Attention: Video Motion Prompts
by: Chen, Qixiang, et al.
Published: (2024)
by: Chen, Qixiang, et al.
Published: (2024)
Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers
by: Chen, Pengtao, et al.
Published: (2025)
by: Chen, Pengtao, et al.
Published: (2025)
PR-MaGIC: Prompt Refinement Via Mask Decoder Gradient Flow For In-Context Segmentation
by: Lee, Minjae, et al.
Published: (2026)
by: Lee, Minjae, et al.
Published: (2026)
RefineNet: Enhancing Text-to-Image Conversion with High-Resolution and Detail Accuracy through Hierarchical Transformers and Progressive Refinement
by: Shi, Fan
Published: (2023)
by: Shi, Fan
Published: (2023)
PRISM: A Multi-View Multi-Capability Retail Video Dataset for Embodied Vision-Language Models
by: Rouhi, Amirreza, et al.
Published: (2026)
by: Rouhi, Amirreza, et al.
Published: (2026)
Diffusion-Based Offline RL for Improved Decision-Making in Augmented ARC Task
by: Kim, Yunho, et al.
Published: (2024)
by: Kim, Yunho, et al.
Published: (2024)
Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition
by: Oh, Youngtaek, et al.
Published: (2024)
by: Oh, Youngtaek, et al.
Published: (2024)
Towards Visual Text Design Transfer Across Languages
by: Choi, Yejin, et al.
Published: (2024)
by: Choi, Yejin, et al.
Published: (2024)
PMG: Progressive Motion Generation via Sparse Anchor Postures Curriculum Learning
by: Xi, Yingjie, et al.
Published: (2025)
by: Xi, Yingjie, et al.
Published: (2025)
Similar Items
-
DAM: Domain-Aware Module for Multi-Domain Dataset Condensation
by: Choi, Jaehyun, et al.
Published: (2025) -
Frequency-Aware Token Reduction for Efficient Vision Transformer
by: Lee, Dong-Jae, et al.
Published: (2025) -
Learning Neural Deformation Representation for 4D Dynamic Shape Generation
by: Han, Gyojin, et al.
Published: (2026) -
Cross-Axis Feature Fusion with Joint-Wise Motion Difference Prediction for Text-Based 3D Human Motion Editing
by: Han, Gyojin, et al.
Published: (2026) -
Self-supervised Transformation Learning for Equivariant Representations
by: Yu, Jaemyung, et al.
Published: (2025)