Saved in:
| Main Authors: | Dong, Yubo, Zhu, Linchao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.01340 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention
by: Lu, Yu, et al.
Published: (2024)
by: Lu, Yu, et al.
Published: (2024)
MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
by: Xu, Yunqiu, et al.
Published: (2024)
by: Xu, Yunqiu, et al.
Published: (2024)
H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
by: Jia, Heng, et al.
Published: (2025)
by: Jia, Heng, et al.
Published: (2025)
3DID: Direct 3D Inverse Design for Aerodynamics with Physics-Aware Optimization
by: Hao, Yuze, et al.
Published: (2025)
by: Hao, Yuze, et al.
Published: (2025)
EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing
by: Yang, Xiangpeng, et al.
Published: (2024)
by: Yang, Xiangpeng, et al.
Published: (2024)
VideoGrain: Modulating Space-Time Attention for Multi-grained Video Editing
by: Yang, Xiangpeng, et al.
Published: (2025)
by: Yang, Xiangpeng, et al.
Published: (2025)
GPD: Guided Progressive Distillation for Fast and High-Quality Video Generation
by: Liang, Xiao, et al.
Published: (2026)
by: Liang, Xiao, et al.
Published: (2026)
Content-Aware Mamba for Learned Image Compression
by: Chen, Yunuo, et al.
Published: (2025)
by: Chen, Yunuo, et al.
Published: (2025)
DA-VAE: Plug-in Latent Compression for Diffusion via Detail Alignment
by: Cai, Xin, et al.
Published: (2026)
by: Cai, Xin, et al.
Published: (2026)
CADC: Content Adaptive Diffusion-Based Generative Image Compression
by: Sheng, Xihua, et al.
Published: (2026)
by: Sheng, Xihua, et al.
Published: (2026)
MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization
by: Jia, Mingkai, et al.
Published: (2025)
by: Jia, Mingkai, et al.
Published: (2025)
RayFormer: Modeling Inter- and Intra-Ray Similarity for NeRF-Based Video Snapshot Compressive Imaging
by: Dong, Yubo, et al.
Published: (2026)
by: Dong, Yubo, et al.
Published: (2026)
EO-VAE: Towards A Multi-sensor Tokenizer for Earth Observation Data
by: Lehmann, Nils, et al.
Published: (2026)
by: Lehmann, Nils, et al.
Published: (2026)
Combating Label Noise With A General Surrogate Model For Sample Selection
by: Liang, Chao, et al.
Published: (2023)
by: Liang, Chao, et al.
Published: (2023)
Slimmable Networks for Contrastive Self-supervised Learning
by: Zhao, Shuai, et al.
Published: (2022)
by: Zhao, Shuai, et al.
Published: (2022)
DGL: Dynamic Global-Local Prompt Tuning for Text-Video Retrieval
by: Yang, Xiangpeng, et al.
Published: (2024)
by: Yang, Xiangpeng, et al.
Published: (2024)
CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
Knowledge-Enhanced Dual-stream Zero-shot Composed Image Retrieval
by: Suo, Yucheng, et al.
Published: (2024)
by: Suo, Yucheng, et al.
Published: (2024)
Collaborative Group: Composed Image Retrieval via Consensus Learning from Noisy Annotations
by: Zhang, Xu, et al.
Published: (2023)
by: Zhang, Xu, et al.
Published: (2023)
ViBE: Visual-to-M/EEG Brain Encoding via Spatio-Temporal VAE and Distribution-Aligned Projection
by: Xu, Ganxi, et al.
Published: (2026)
by: Xu, Ganxi, et al.
Published: (2026)
TVRN: Invertible Neural Networks for Compression-Aware Temporal Video Rescaling
by: Feng, Xinmin, et al.
Published: (2026)
by: Feng, Xinmin, et al.
Published: (2026)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
Pose-Aware Multi-Level Motion Parsing for Action Quality Assessment
by: Zhu, Shuaikang, et al.
Published: (2025)
by: Zhu, Shuaikang, et al.
Published: (2025)
Flip Distribution Alignment VAE for Multi-Phase MRI Synthesis
by: Kui, Xiaoyan, et al.
Published: (2025)
by: Kui, Xiaoyan, et al.
Published: (2025)
AudioScenic: Audio-Driven Video Scene Editing
by: Shen, Kaixin, et al.
Published: (2024)
by: Shen, Kaixin, et al.
Published: (2024)
UniGarmentManip: A Unified Framework for Category-Level Garment Manipulation via Dense Visual Correspondence
by: Wu, Ruihai, et al.
Published: (2024)
by: Wu, Ruihai, et al.
Published: (2024)
Towards 1000-fold Electron Microscopy Image Compression for Connectomics via VQ-VAE with Transformer Prior
by: Yang, Fuming, et al.
Published: (2025)
by: Yang, Fuming, et al.
Published: (2025)
Video Compression with Hierarchical Temporal Neural Representation
by: Zhu, Jun, et al.
Published: (2026)
by: Zhu, Jun, et al.
Published: (2026)
CANeRV: Content Adaptive Neural Representation for Video Compression
by: Tang, Lv, et al.
Published: (2025)
by: Tang, Lv, et al.
Published: (2025)
Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object Tracking
by: Wang, Shilei, et al.
Published: (2026)
by: Wang, Shilei, et al.
Published: (2026)
PIFu for the Real World: A Self-supervised Framework to Reconstruct Dressed Human from Single-view Images
by: Xiong, Zhangyang, et al.
Published: (2022)
by: Xiong, Zhangyang, et al.
Published: (2022)
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
by: Zhou, Yupeng, et al.
Published: (2025)
by: Zhou, Yupeng, et al.
Published: (2025)
DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs
by: Park, Minyoung, et al.
Published: (2026)
by: Park, Minyoung, et al.
Published: (2026)
Spatial Degradation-Aware and Temporal Consistent Diffusion Model for Compressed Video Super-Resolution
by: An, Hongyu, et al.
Published: (2025)
by: An, Hongyu, et al.
Published: (2025)
BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving
by: Chen, Zeming, et al.
Published: (2025)
by: Chen, Zeming, et al.
Published: (2025)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
by: Guo, Yanan, et al.
Published: (2025)
by: Guo, Yanan, et al.
Published: (2025)
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
by: Zhang, Yunzhu, et al.
Published: (2025)
by: Zhang, Yunzhu, et al.
Published: (2025)
MVP: Multiple View Prediction Improves GUI Grounding
by: Zhang, Yunzhu, et al.
Published: (2025)
by: Zhang, Yunzhu, et al.
Published: (2025)
Ultron: Enabling Temporal Geometry Compression of 3D Mesh Sequences using Temporal Correspondence and Mesh Deformation
by: Zhu, Haichao
Published: (2024)
by: Zhu, Haichao
Published: (2024)
Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
by: Zhu, Lunjie, et al.
Published: (2026)
by: Zhu, Lunjie, et al.
Published: (2026)
Similar Items
-
FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention
by: Lu, Yu, et al.
Published: (2024) -
MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
by: Xu, Yunqiu, et al.
Published: (2024) -
H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
by: Jia, Heng, et al.
Published: (2025) -
3DID: Direct 3D Inverse Design for Aerodynamics with Physics-Aware Optimization
by: Hao, Yuze, et al.
Published: (2025) -
EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing
by: Yang, Xiangpeng, et al.
Published: (2024)