Aligning Latent Spaces with Flow Priors
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yizhuo, Ge, Yuying, Ge, Yixiao, Shan, Ying, Luo, Ping |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiCoDe: Diffusion-Compressed Deep Tokens for Autoregressive Video Generation with Language Models
by: Li, Yizhuo, et al.
Published: (2024)
by: Li, Yizhuo, et al.
Published: (2024)
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
by: Ge, Yuying, et al.
Published: (2024)
by: Ge, Yuying, et al.
Published: (2024)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
by: Chen, Yi, et al.
Published: (2024)
by: Chen, Yi, et al.
Published: (2024)
AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation
by: Qiu, Lu, et al.
Published: (2025)
by: Qiu, Lu, et al.
Published: (2025)
DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA
by: Chen, Yi, et al.
Published: (2026)
by: Chen, Yi, et al.
Published: (2026)
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
GrootVL: Tree Topology is All You Need in State Space Model
by: Xiao, Yicheng, et al.
Published: (2024)
by: Xiao, Yicheng, et al.
Published: (2024)
SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing
by: Ge, Yuying, et al.
Published: (2024)
by: Ge, Yuying, et al.
Published: (2024)
AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction
by: Cheng, Junhao, et al.
Published: (2025)
by: Cheng, Junhao, et al.
Published: (2025)
ARC-Chapter: Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries
by: Pu, Junfu, et al.
Published: (2025)
by: Pu, Junfu, et al.
Published: (2025)
SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
by: Li, Bohao, et al.
Published: (2024)
by: Li, Bohao, et al.
Published: (2024)
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
by: Cheng, Junhao, et al.
Published: (2025)
by: Cheng, Junhao, et al.
Published: (2025)
GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers
by: Ma, Shijie, et al.
Published: (2025)
by: Ma, Shijie, et al.
Published: (2025)
Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Learning Multimodal Latent Space with EBM Prior and MCMC Inference
by: Yuan, Shiyu, et al.
Published: (2024)
by: Yuan, Shiyu, et al.
Published: (2024)
SEED-Story: Multimodal Long Story Generation with Large Language Model
by: Yang, Shuai, et al.
Published: (2024)
by: Yang, Shuai, et al.
Published: (2024)
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
by: Ding, Xiaohan, et al.
Published: (2023)
by: Ding, Xiaohan, et al.
Published: (2023)
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios
by: Qiu, Lu, et al.
Published: (2024)
by: Qiu, Lu, et al.
Published: (2024)
Supervised Fine-tuning in turn Improves Visual Foundation Models
by: Jiang, Xiaohu, et al.
Published: (2024)
by: Jiang, Xiaohu, et al.
Published: (2024)
Neural Prior Estimation: Learning Class Priors from Latent Representations
by: Yavari, Masoud, et al.
Published: (2026)
by: Yavari, Masoud, et al.
Published: (2026)
ShaLa: Multimodal Shared Latent Space Modelling
by: Cui, Jiali, et al.
Published: (2025)
by: Cui, Jiali, et al.
Published: (2025)
Learning Multimodal Latent Generative Models with Energy-Based Prior
by: Yuan, Shiyu, et al.
Published: (2024)
by: Yuan, Shiyu, et al.
Published: (2024)
AudioStory: Generating Long-Form Narrative Audio with Large Language Models
by: Guo, Yuxin, et al.
Published: (2025)
by: Guo, Yuxin, et al.
Published: (2025)
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
by: Ge, Yuying, et al.
Published: (2024)
by: Ge, Yuying, et al.
Published: (2024)
GNFactor: Multi-Task Real Robot Learning with Generalizable Neural Feature Fields
by: Ze, Yanjie, et al.
Published: (2023)
by: Ze, Yanjie, et al.
Published: (2023)
ST-LLM: Large Language Models Are Effective Temporal Learners
by: Liu, Ruyang, et al.
Published: (2024)
by: Liu, Ruyang, et al.
Published: (2024)
Aligning Data Selection with Performance: Performance-driven Reinforcement Learning for Active Learning in Object Detection
by: Liang, Zhixuan, et al.
Published: (2023)
by: Liang, Zhixuan, et al.
Published: (2023)
Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models
by: He, Dailan, et al.
Published: (2025)
by: He, Dailan, et al.
Published: (2025)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
Flow Priors for Linear Inverse Problems via Iterative Corrupted Trajectory Matching
by: Zhang, Yasi, et al.
Published: (2024)
by: Zhang, Yasi, et al.
Published: (2024)
Latent Diffusion Inversion Requires Understanding the Latent Space
by: Rao, Mingxing, et al.
Published: (2025)
by: Rao, Mingxing, et al.
Published: (2025)
CASL: Concept-Aligned Sparse Latents for Interpreting Diffusion Models
by: He, Zhenghao, et al.
Published: (2026)
by: He, Zhenghao, et al.
Published: (2026)
PixelFlow: Pixel-Space Generative Models with Flow
by: Chen, Shoufa, et al.
Published: (2025)
by: Chen, Shoufa, et al.
Published: (2025)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
by: Liu, Ruyang, et al.
Published: (2023)
by: Liu, Ruyang, et al.
Published: (2023)
Align Your Flow: Scaling Continuous-Time Flow Map Distillation
by: Sabour, Amirmojtaba, et al.
Published: (2025)
by: Sabour, Amirmojtaba, et al.
Published: (2025)
Moving Object Proposals with Deep Learned Optical Flow for Video Object Segmentation
by: Shi, Ge, et al.
Published: (2024)
by: Shi, Ge, et al.
Published: (2024)
Aligning Few-Step Diffusion Models with Dense Reward Difference Learning
by: Zhang, Ziyi, et al.
Published: (2024)
by: Zhang, Ziyi, et al.
Published: (2024)
Flow-Factory: A Unified Framework for Reinforcement Learning in Flow-Matching Models
by: Ping, Bowen, et al.
Published: (2026)
by: Ping, Bowen, et al.
Published: (2026)
Learning Discrete Autoregressive Priors with Wasserstein Gradient Flow
by: Zheng, Bowen, et al.
Published: (2026)
by: Zheng, Bowen, et al.
Published: (2026)
Similar Items
-
DiCoDe: Diffusion-Compressed Deep Tokens for Autoregressive Video Generation with Language Models
by: Li, Yizhuo, et al.
Published: (2024) -
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
by: Ge, Yuying, et al.
Published: (2024) -
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
by: Chen, Yi, et al.
Published: (2024) -
AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation
by: Qiu, Lu, et al.
Published: (2025) -
DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA
by: Chen, Yi, et al.
Published: (2026)