Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yan, Tian, Changyao, Xia, Renqiu, Liao, Ning, Guo, Weiwei, Yan, Junchi, Li, Hongsheng, Dai, Jifeng, Li, Hao, Yang, Xue |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets
von: Liao, Ning, et al.
Veröffentlicht: (2023)
von: Liao, Ning, et al.
Veröffentlicht: (2023)
Learning 1D Causal Visual Representation with De-focus Attention Networks
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
von: Tao, Chenxin, et al.
Veröffentlicht: (2024)
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
von: Li, Yan, et al.
Veröffentlicht: (2026)
von: Li, Yan, et al.
Veröffentlicht: (2026)
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
ADDP: Learning General Representations for Image Recognition and Generation with Alternating Denoising Diffusion Process
von: Tian, Changyao, et al.
Veröffentlicht: (2023)
von: Tian, Changyao, et al.
Veröffentlicht: (2023)
Exploiting Unlabeled Data with Multiple Expert Teachers for Open Vocabulary Aerial Object Detection and Its Orientation Adaptation
von: Li, Yan, et al.
Veröffentlicht: (2024)
von: Li, Yan, et al.
Veröffentlicht: (2024)
Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy
von: Ye, Hancheng, et al.
Veröffentlicht: (2024)
von: Ye, Hancheng, et al.
Veröffentlicht: (2024)
Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving
von: Liu, Qi, et al.
Veröffentlicht: (2025)
von: Liu, Qi, et al.
Veröffentlicht: (2025)
Point2RBox: Combine Knowledge from Synthetic Visual Patterns for End-to-end Oriented Object Detection with Single Point Supervision
von: Yu, Yi, et al.
Veröffentlicht: (2023)
von: Yu, Yi, et al.
Veröffentlicht: (2023)
Progressive Growing of Video Tokenizers for Temporally Compact Latent Spaces
von: Mahapatra, Aniruddha, et al.
Veröffentlicht: (2025)
von: Mahapatra, Aniruddha, et al.
Veröffentlicht: (2025)
DriveVGGT: Calibration-Constrained Visual Geometry Transformers for Multi-Camera Autonomous Driving
von: Jia, Xiaosong, et al.
Veröffentlicht: (2025)
von: Jia, Xiaosong, et al.
Veröffentlicht: (2025)
Local Learning for Covariate Selection in Nonparametric Causal Effect Estimation with Latent Variables
von: Li, Zheng, et al.
Veröffentlicht: (2024)
von: Li, Zheng, et al.
Veröffentlicht: (2024)
Toward Open Vocabulary Aerial Object Detection with CLIP-Activated Student-Teacher Learning
von: Li, Yan, et al.
Veröffentlicht: (2023)
von: Li, Yan, et al.
Veröffentlicht: (2023)
NTKMTL: Mitigating Task Imbalance in Multi-Task Learning from Neural Tangent Kernel Perspective
von: Qin, Xiaohan, et al.
Veröffentlicht: (2025)
von: Qin, Xiaohan, et al.
Veröffentlicht: (2025)
Think2Drive: Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous Driving (in CARLA-v2)
von: Li, Qifeng, et al.
Veröffentlicht: (2024)
von: Li, Qifeng, et al.
Veröffentlicht: (2024)
LangBridge: Interpreting Image as a Combination of Language Embeddings
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liao, Jiaqi, et al.
Veröffentlicht: (2025)
StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding
von: Xia, Renqiu, et al.
Veröffentlicht: (2023)
von: Xia, Renqiu, et al.
Veröffentlicht: (2023)
How Transformers Learn to Plan via Multi-Token Prediction
von: Huang, Jianhao, et al.
Veröffentlicht: (2026)
von: Huang, Jianhao, et al.
Veröffentlicht: (2026)
ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuning
von: Qin, Xiaohan, et al.
Veröffentlicht: (2025)
von: Qin, Xiaohan, et al.
Veröffentlicht: (2025)
Trust Your Critic: Robust Reward Modeling and Reinforcement Learning for Faithful Image Editing and Generation
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2026)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2026)
Generative Latent Video Compression
von: Guo, Zongyu, et al.
Veröffentlicht: (2025)
von: Guo, Zongyu, et al.
Veröffentlicht: (2025)
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
von: Luo, Gen, et al.
Veröffentlicht: (2025)
von: Luo, Gen, et al.
Veröffentlicht: (2025)
Generative Human Motion Stylization in Latent Space
von: Guo, Chuan, et al.
Veröffentlicht: (2024)
von: Guo, Chuan, et al.
Veröffentlicht: (2024)
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
von: Tian, Changyao, et al.
Veröffentlicht: (2025)
von: Tian, Changyao, et al.
Veröffentlicht: (2025)
AMP: Autoregressive Motion Prediction Revisited with Next Token Prediction for Autonomous Driving
von: Jia, Xiaosong, et al.
Veröffentlicht: (2024)
von: Jia, Xiaosong, et al.
Veröffentlicht: (2024)
Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low Retention
von: Du, Junhao, et al.
Veröffentlicht: (2026)
von: Du, Junhao, et al.
Veröffentlicht: (2026)
FineRMoE: Dimension Expansion for Finer-Grained Expert with Its Upcycling Approach
von: Liao, Ning, et al.
Veröffentlicht: (2026)
von: Liao, Ning, et al.
Veröffentlicht: (2026)
M-Tuning: Prompt Tuning with Mitigated Label Bias in Open-Set Scenarios
von: Liao, Ning, et al.
Veröffentlicht: (2023)
von: Liao, Ning, et al.
Veröffentlicht: (2023)
ElasticTok: Adaptive Tokenization for Image and Video
von: Yan, Wilson, et al.
Veröffentlicht: (2024)
von: Yan, Wilson, et al.
Veröffentlicht: (2024)
Temporal Latent Variable Structural Causal Model for Causal Discovery under External Interferences
von: Cai, Ruichu, et al.
Veröffentlicht: (2025)
von: Cai, Ruichu, et al.
Veröffentlicht: (2025)
VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents
von: Wang, Feng, et al.
Veröffentlicht: (2026)
von: Wang, Feng, et al.
Veröffentlicht: (2026)
Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models
von: Xing, Wenbin, et al.
Veröffentlicht: (2026)
von: Xing, Wenbin, et al.
Veröffentlicht: (2026)
Local Path Optimization in The Latent Space Using Learned Distance Gradient
von: Zhang, Jiawei, et al.
Veröffentlicht: (2025)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2025)
SPOT: Scalable 3D Pre-training via Occupancy Prediction for Learning Transferable 3D Representations
von: Yan, Xiangchao, et al.
Veröffentlicht: (2023)
von: Yan, Xiangchao, et al.
Veröffentlicht: (2023)
Generative Video Compression with One-Dimensional Latent Representation
von: Zheng, Zihan, et al.
Veröffentlicht: (2026)
von: Zheng, Zihan, et al.
Veröffentlicht: (2026)
Causal Structure Representation Learning of Confounders in Latent Space for Recommendation
von: Xu, Hangtong, et al.
Veröffentlicht: (2023)
von: Xu, Hangtong, et al.
Veröffentlicht: (2023)
Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2025)
Rethinking Video Tokenization: A Conditioned Diffusion-based Approach
von: Yang, Nianzu, et al.
Veröffentlicht: (2025)
von: Yang, Nianzu, et al.
Veröffentlicht: (2025)
Local Causal Structure Learning in the Presence of Latent Variables
von: Xie, Feng, et al.
Veröffentlicht: (2024)
von: Xie, Feng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets
von: Liao, Ning, et al.
Veröffentlicht: (2023) -
Learning 1D Causal Visual Representation with De-focus Attention Networks
von: Tao, Chenxin, et al.
Veröffentlicht: (2024) -
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
von: Li, Yan, et al.
Veröffentlicht: (2026) -
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
von: Li, Hao, et al.
Veröffentlicht: (2024) -
ADDP: Learning General Representations for Image Recognition and Generation with Alternating Denoising Diffusion Process
von: Tian, Changyao, et al.
Veröffentlicht: (2023)