Progressive Growing of Video Tokenizers for Temporally Compact Latent Spaces
Fuente:
arXiv
Saved in:
| Main Authors: | Mahapatra, Aniruddha, Mai, Long, Bourgin, David, Zhang, Yitian, Liu, Feng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder
by: Zhang, Yitian, et al.
Published: (2025)
by: Zhang, Yitian, et al.
Published: (2025)
Multiple Latent Space Mapping for Compressed Dark Image Enhancement
by: Zeng, Yi, et al.
Published: (2024)
by: Zeng, Yi, et al.
Published: (2024)
Optimal Transport Driven Asymmetric Image-to-Image Translation for Nuclei Segmentation of Histological Images
by: Mahapatra, Suman, et al.
Published: (2025)
by: Mahapatra, Suman, et al.
Published: (2025)
SiliCoN: Simultaneous Nuclei Segmentation and Color Normalization of Histological Images
by: Mahapatra, Suman, et al.
Published: (2025)
by: Mahapatra, Suman, et al.
Published: (2025)
CV-VAE: A Compatible Video VAE for Latent Generative Video Models
by: Zhao, Sijie, et al.
Published: (2024)
by: Zhao, Sijie, et al.
Published: (2024)
Latent Space Consistency for Sparse-View CT Reconstruction
by: Chen, Duoyou, et al.
Published: (2025)
by: Chen, Duoyou, et al.
Published: (2025)
Uncertainty-Aware Post-Detection Framework for Enhanced Fire and Smoke Detection in Compact Deep Learning Models
by: Joshi, Aniruddha Srinivas, et al.
Published: (2025)
by: Joshi, Aniruddha Srinivas, et al.
Published: (2025)
SegReg: Latent Space Regularization for Improved Medical Image Segmentation
by: Vaish, Puru, et al.
Published: (2026)
by: Vaish, Puru, et al.
Published: (2026)
Latent Space Synergy: Text-Guided Data Augmentation for Direct Diffusion Biomedical Segmentation
by: Aqeel, Muhammad, et al.
Published: (2025)
by: Aqeel, Muhammad, et al.
Published: (2025)
Generating Progressive Images from Pathological Transitions via Diffusion Model
by: Liu, Zeyu, et al.
Published: (2023)
by: Liu, Zeyu, et al.
Published: (2023)
EchoLVFM: One-Step Video Generation via Latent Flow Matching for Echocardiogram Synthesis
by: Oladokun, Emmanuel, et al.
Published: (2026)
by: Oladokun, Emmanuel, et al.
Published: (2026)
Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization
by: Liang, Jingyun, et al.
Published: (2026)
by: Liang, Jingyun, et al.
Published: (2026)
Energy-Based Prior Latent Space Diffusion model for Reconstruction of Lumbar Vertebrae from Thick Slice MRI
by: Wang, Yanke, et al.
Published: (2024)
by: Wang, Yanke, et al.
Published: (2024)
GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion
by: Kim, Gwanghyun, et al.
Published: (2025)
by: Kim, Gwanghyun, et al.
Published: (2025)
RAISE: Realness Assessment for Image Synthesis and Evaluation
by: Mukherjee, Aniruddha, et al.
Published: (2025)
by: Mukherjee, Aniruddha, et al.
Published: (2025)
Semantic and Temporal Integration in Latent Diffusion Space for High-Fidelity Video Super-Resolution
by: Wang, Yiwen, et al.
Published: (2025)
by: Wang, Yiwen, et al.
Published: (2025)
LatentArtiFusion: An Effective and Efficient Histological Artifacts Restoration Framework
by: He, Zhenqi, et al.
Published: (2024)
by: He, Zhenqi, et al.
Published: (2024)
Modulo Video Recovery via Selective Spatiotemporal Vision Transformer
by: Geng, Tianyu, et al.
Published: (2025)
by: Geng, Tianyu, et al.
Published: (2025)
Multiscale Structure-Guided Latent Diffusion for Multimodal MRI Translation
by: Lin, Jianqiang, et al.
Published: (2026)
by: Lin, Jianqiang, et al.
Published: (2026)
Beyond the Eye: A Relational Model for Early Dementia Detection Using Retinal OCTA Images
by: Liu, Shouyue, et al.
Published: (2024)
by: Liu, Shouyue, et al.
Published: (2024)
Are Compact Rationales Free? Measuring Tile Selection Headroom in Frozen WSI-MIL
by: Jung, Hyun Do, et al.
Published: (2026)
by: Jung, Hyun Do, et al.
Published: (2026)
3DPX: Progressive 2D-to-3D Oral Image Reconstruction with Hybrid MLP-CNN Networks
by: Li, Xiaoshuang, et al.
Published: (2024)
by: Li, Xiaoshuang, et al.
Published: (2024)
Progressive Alignment Degradation Learning for Pansharpening
by: Zhao, Enzhe, et al.
Published: (2025)
by: Zhao, Enzhe, et al.
Published: (2025)
Progressive Curriculum Learning with Scale-Enhanced U-Net for Continuous Airway Segmentation
by: Yang, Bingyu, et al.
Published: (2024)
by: Yang, Bingyu, et al.
Published: (2024)
Geospatial-Temporal Sensemaking of Remote Sensing Activity Detections with Multimodal Large Language Model
by: Ramirez, David F., et al.
Published: (2026)
by: Ramirez, David F., et al.
Published: (2026)
T1-contrast Enhanced MRI Generation from Multi-parametric MRI for Glioma Patients with Latent Tumor Conditioning
by: Eidex, Zach, et al.
Published: (2024)
by: Eidex, Zach, et al.
Published: (2024)
CIResDiff: A Clinically-Informed Residual Diffusion Model for Predicting Idiopathic Pulmonary Fibrosis Progression
by: Jiang, Caiwen, et al.
Published: (2024)
by: Jiang, Caiwen, et al.
Published: (2024)
ST-NeRP: Spatial-Temporal Neural Representation Learning with Prior Embedding for Patient-specific Imaging Study
by: Qiu, Liang, et al.
Published: (2024)
by: Qiu, Liang, et al.
Published: (2024)
st-DTPM: Spatial-Temporal Guided Diffusion Transformer Probabilistic Model for Delayed Scan PET Image Prediction
by: Hong, Ran, et al.
Published: (2024)
by: Hong, Ran, et al.
Published: (2024)
Synthetic Generation and Latent Projection Denoising of Rim Lesions in Multiple Sclerosis
by: Roberts, Alexandra G., et al.
Published: (2025)
by: Roberts, Alexandra G., et al.
Published: (2025)
STACT-Time: Spatio-Temporal Cross Attention for Cine Thyroid Ultrasound Time Series Classification
by: Adam, Irsyad, et al.
Published: (2025)
by: Adam, Irsyad, et al.
Published: (2025)
Visual Question Answering in Ophthalmology: A Progressive and Practical Perspective
by: Chen, Xiaolan, et al.
Published: (2024)
by: Chen, Xiaolan, et al.
Published: (2024)
Progressive Transfer Learning for Multi-Pass Fundus Image Restoration
by: Phan, Uyen, et al.
Published: (2025)
by: Phan, Uyen, et al.
Published: (2025)
Multimodal Fusion and Coherence Modeling for Video Topic Segmentation
by: Yu, Hai, et al.
Published: (2024)
by: Yu, Hai, et al.
Published: (2024)
Adaptive Learning Strategies for Mitotic Figure Classification in MIDOG2025 Challenge
by: Meng, Biwen, et al.
Published: (2025)
by: Meng, Biwen, et al.
Published: (2025)
Latent Anomaly Detection: Masked VQ-GAN for Unsupervised Segmentation in Medical CBCT
by: Wang, Pengwei
Published: (2025)
by: Wang, Pengwei
Published: (2025)
LSA: Latent Style Augmentation Towards Stain-Agnostic Cervical Cancer Screening
by: Cai, Jiangdong, et al.
Published: (2025)
by: Cai, Jiangdong, et al.
Published: (2025)
Generative Artificial Intelligence in Medical Imaging: Foundations, Progress, and Clinical Translation
by: Zhou, Xuanru, et al.
Published: (2025)
by: Zhou, Xuanru, et al.
Published: (2025)
VideoQA-SC: Adaptive Semantic Communication for Video Question Answering
by: Guo, Jiangyuan, et al.
Published: (2024)
by: Guo, Jiangyuan, et al.
Published: (2024)
Uncovering Latent Pathological Signatures in Pulmonary CT via Cross-Window Knowledge Distillation
by: Peng, Bo, et al.
Published: (2026)
by: Peng, Bo, et al.
Published: (2026)
Similar Items
-
REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder
by: Zhang, Yitian, et al.
Published: (2025) -
Multiple Latent Space Mapping for Compressed Dark Image Enhancement
by: Zeng, Yi, et al.
Published: (2024) -
Optimal Transport Driven Asymmetric Image-to-Image Translation for Nuclei Segmentation of Histological Images
by: Mahapatra, Suman, et al.
Published: (2025) -
SiliCoN: Simultaneous Nuclei Segmentation and Color Normalization of Histological Images
by: Mahapatra, Suman, et al.
Published: (2025) -
CV-VAE: A Compatible Video VAE for Latent Generative Video Models
by: Zhao, Sijie, et al.
Published: (2024)