MTC-VAE: Multi-Level Temporal Compression with Content Awareness
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dong, Yubo, Zhu, Linchao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
von: Xu, Yunqiu, et al.
Veröffentlicht: (2024)
von: Xu, Yunqiu, et al.
Veröffentlicht: (2024)
FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention
von: Lu, Yu, et al.
Veröffentlicht: (2024)
von: Lu, Yu, et al.
Veröffentlicht: (2024)
H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
von: Jia, Heng, et al.
Veröffentlicht: (2025)
von: Jia, Heng, et al.
Veröffentlicht: (2025)
3DID: Direct 3D Inverse Design for Aerodynamics with Physics-Aware Optimization
von: Hao, Yuze, et al.
Veröffentlicht: (2025)
von: Hao, Yuze, et al.
Veröffentlicht: (2025)
EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing
von: Yang, Xiangpeng, et al.
Veröffentlicht: (2024)
von: Yang, Xiangpeng, et al.
Veröffentlicht: (2024)
Content-Aware Mamba for Learned Image Compression
von: Chen, Yunuo, et al.
Veröffentlicht: (2025)
von: Chen, Yunuo, et al.
Veröffentlicht: (2025)
VideoGrain: Modulating Space-Time Attention for Multi-grained Video Editing
von: Yang, Xiangpeng, et al.
Veröffentlicht: (2025)
von: Yang, Xiangpeng, et al.
Veröffentlicht: (2025)
DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignment
von: Cai, Xin, et al.
Veröffentlicht: (2026)
von: Cai, Xin, et al.
Veröffentlicht: (2026)
GPD: Guided Progressive Distillation for Fast and High-Quality Video Generation
von: Liang, Xiao, et al.
Veröffentlicht: (2026)
von: Liang, Xiao, et al.
Veröffentlicht: (2026)
CADC: Content Adaptive Diffusion-Based Generative Image Compression
von: Sheng, Xihua, et al.
Veröffentlicht: (2026)
von: Sheng, Xihua, et al.
Veröffentlicht: (2026)
MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization
von: Jia, Mingkai, et al.
Veröffentlicht: (2025)
von: Jia, Mingkai, et al.
Veröffentlicht: (2025)
EO-VAE: Towards A Multi-sensor Tokenizer for Earth Observation Data
von: Lehmann, Nils, et al.
Veröffentlicht: (2026)
von: Lehmann, Nils, et al.
Veröffentlicht: (2026)
RayFormer: Modeling Inter- and Intra-Ray Similarity for NeRF-Based Video Snapshot Compressive Imaging
von: Dong, Yubo, et al.
Veröffentlicht: (2026)
von: Dong, Yubo, et al.
Veröffentlicht: (2026)
ViBE: Visual-to-M/EEG Brain Encoding via Spatio-Temporal VAE and Distribution-Aligned Projection
von: Xu, Ganxi, et al.
Veröffentlicht: (2026)
von: Xu, Ganxi, et al.
Veröffentlicht: (2026)
Combating Label Noise With A General Surrogate Model For Sample Selection
von: Liang, Chao, et al.
Veröffentlicht: (2023)
von: Liang, Chao, et al.
Veröffentlicht: (2023)
Slimmable Networks for Contrastive Self-supervised Learning
von: Zhao, Shuai, et al.
Veröffentlicht: (2022)
von: Zhao, Shuai, et al.
Veröffentlicht: (2022)
DGL: Dynamic Global-Local Prompt Tuning for Text-Video Retrieval
von: Yang, Xiangpeng, et al.
Veröffentlicht: (2024)
von: Yang, Xiangpeng, et al.
Veröffentlicht: (2024)
CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
Knowledge-Enhanced Dual-stream Zero-shot Composed Image Retrieval
von: Suo, Yucheng, et al.
Veröffentlicht: (2024)
von: Suo, Yucheng, et al.
Veröffentlicht: (2024)
Collaborative Group: Composed Image Retrieval via Consensus Learning from Noisy Annotations
von: Zhang, Xu, et al.
Veröffentlicht: (2023)
von: Zhang, Xu, et al.
Veröffentlicht: (2023)
Pose-Aware Multi-Level Motion Parsing for Action Quality Assessment
von: Zhu, Shuaikang, et al.
Veröffentlicht: (2025)
von: Zhu, Shuaikang, et al.
Veröffentlicht: (2025)
Flip Distribution Alignment VAE for Multi-Phase MRI Synthesis
von: Kui, Xiaoyan, et al.
Veröffentlicht: (2025)
von: Kui, Xiaoyan, et al.
Veröffentlicht: (2025)
Towards 1000-fold Electron Microscopy Image Compression for Connectomics via VQ-VAE with Transformer Prior
von: Yang, Fuming, et al.
Veröffentlicht: (2025)
von: Yang, Fuming, et al.
Veröffentlicht: (2025)
TVRN: Invertible Neural Networks for Compression-Aware Temporal Video Rescaling
von: Feng, Xinmin, et al.
Veröffentlicht: (2026)
von: Feng, Xinmin, et al.
Veröffentlicht: (2026)
Video Compression with Hierarchical Temporal Neural Representation
von: Zhu, Jun, et al.
Veröffentlicht: (2026)
von: Zhu, Jun, et al.
Veröffentlicht: (2026)
Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object Tracking
von: Wang, Shilei, et al.
Veröffentlicht: (2026)
von: Wang, Shilei, et al.
Veröffentlicht: (2026)
CANeRV: Content Adaptive Neural Representation for Video Compression
von: Tang, Lv, et al.
Veröffentlicht: (2025)
von: Tang, Lv, et al.
Veröffentlicht: (2025)
DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs
von: Park, Minyoung, et al.
Veröffentlicht: (2026)
von: Park, Minyoung, et al.
Veröffentlicht: (2026)
Spatial Degradation-Aware and Temporal Consistent Diffusion Model for Compressed Video Super-Resolution
von: An, Hongyu, et al.
Veröffentlicht: (2025)
von: An, Hongyu, et al.
Veröffentlicht: (2025)
UniGarmentManip: A Unified Framework for Category-Level Garment Manipulation via Dense Visual Correspondence
von: Wu, Ruihai, et al.
Veröffentlicht: (2024)
von: Wu, Ruihai, et al.
Veröffentlicht: (2024)
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
von: Zhou, Yupeng, et al.
Veröffentlicht: (2025)
von: Zhou, Yupeng, et al.
Veröffentlicht: (2025)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving
von: Chen, Zeming, et al.
Veröffentlicht: (2025)
von: Chen, Zeming, et al.
Veröffentlicht: (2025)
AudioScenic: Audio-Driven Video Scene Editing
von: Shen, Kaixin, et al.
Veröffentlicht: (2024)
von: Shen, Kaixin, et al.
Veröffentlicht: (2024)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
Ultron: Enabling Temporal Geometry Compression of 3D Mesh Sequences using Temporal Correspondence and Mesh Deformation
von: Zhu, Haichao
Veröffentlicht: (2024)
von: Zhu, Haichao
Veröffentlicht: (2024)
CFNet: Optimizing Remote Sensing Change Detection through Content-Aware Enhancement
von: Wu, Fan, et al.
Veröffentlicht: (2025)
von: Wu, Fan, et al.
Veröffentlicht: (2025)
Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
von: Zhu, Lunjie, et al.
Veröffentlicht: (2026)
von: Zhu, Lunjie, et al.
Veröffentlicht: (2026)
PIFu for the Real World: A Self-supervised Framework to Reconstruct Dressed Human from Single-view Images
von: Xiong, Zhangyang, et al.
Veröffentlicht: (2022)
von: Xiong, Zhangyang, et al.
Veröffentlicht: (2022)
MultiHateLoc: Towards Temporal Localisation of Multimodal Hate Content in Online Videos
von: Sun, Qiyue, et al.
Veröffentlicht: (2025)
von: Sun, Qiyue, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
von: Xu, Yunqiu, et al.
Veröffentlicht: (2024) -
FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention
von: Lu, Yu, et al.
Veröffentlicht: (2024) -
H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
von: Jia, Heng, et al.
Veröffentlicht: (2025) -
3DID: Direct 3D Inverse Design for Aerodynamics with Physics-Aware Optimization
von: Hao, Yuze, et al.
Veröffentlicht: (2025) -
EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing
von: Yang, Xiangpeng, et al.
Veröffentlicht: (2024)