Posterior Augmented Flow Matching
Fuente:
arXiv
Saved in:
| Main Authors: | Stoica, George, Paul, Sayak, Wallingford, Matthew, Ramanujan, Vivek, Nori, Abhay, Han, Winson, Farhadi, Ali, Krishna, Ranjay, Hoffman, Judy |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Contrastive Flow Matching
by: Stoica, George, et al.
Published: (2025)
by: Stoica, George, et al.
Published: (2025)
The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs Better
by: Geng, Scott, et al.
Published: (2024)
by: Geng, Scott, et al.
Published: (2024)
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
by: Yadav, Tanush, et al.
Published: (2026)
by: Yadav, Tanush, et al.
Published: (2026)
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
by: Wallingford, Matthew, et al.
Published: (2024)
by: Wallingford, Matthew, et al.
Published: (2024)
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
by: Ramanujan, Vivek, et al.
Published: (2024)
by: Ramanujan, Vivek, et al.
Published: (2024)
Beyond the Frame: Generating 360 Panoramic Videos from Perspective Videos
by: Luo, Rundong, et al.
Published: (2025)
by: Luo, Rundong, et al.
Published: (2025)
Matryoshka Representation Learning
by: Kusupati, Aditya, et al.
Published: (2022)
by: Kusupati, Aditya, et al.
Published: (2022)
Model merging with SVD to tie the Knots
by: Stoica, George, et al.
Published: (2024)
by: Stoica, George, et al.
Published: (2024)
WildDet3D: Scaling Promptable 3D Detection in the Wild
by: Huang, Weikai, et al.
Published: (2026)
by: Huang, Weikai, et al.
Published: (2026)
Visual Representations inside the Language Model
by: Liu, Benlin, et al.
Published: (2025)
by: Liu, Benlin, et al.
Published: (2025)
Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation
by: Fei, Yang, et al.
Published: (2025)
by: Fei, Yang, et al.
Published: (2025)
ZipIt! Merging Models from Different Tasks without Training
by: Stoica, George, et al.
Published: (2023)
by: Stoica, George, et al.
Published: (2023)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
by: Eftekhar, Ainaz, et al.
Published: (2023)
by: Eftekhar, Ainaz, et al.
Published: (2023)
OmniView: An All-Seeing Diffusion Model for 3D and 4D View Synthesis
by: Fan, Xiang, et al.
Published: (2025)
by: Fan, Xiang, et al.
Published: (2025)
Resolving Interference (RI): Disentangling Models for Improved Model Merging
by: Ramesh, Pratik, et al.
Published: (2026)
by: Ramesh, Pratik, et al.
Published: (2026)
Multilingual Diversity Improves Vision-Language Representations
by: Nguyen, Thao, et al.
Published: (2024)
by: Nguyen, Thao, et al.
Published: (2024)
Seeing Fast and Slow: Learning the Flow of Time in Videos
by: Wu, Yen-Siang, et al.
Published: (2026)
by: Wu, Yen-Siang, et al.
Published: (2026)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
Flow Matching Posterior Sampling: A Training-free Conditional Generation for Flow Matching
by: Song, Kaiyu, et al.
Published: (2024)
by: Song, Kaiyu, et al.
Published: (2024)
The One RING: a Robotic Indoor Navigation Generalist
by: Eftekhar, Ainaz, et al.
Published: (2024)
by: Eftekhar, Ainaz, et al.
Published: (2024)
Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
by: Huang, Weikai, et al.
Published: (2025)
by: Huang, Weikai, et al.
Published: (2025)
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
by: Gupta, Tanmay, et al.
Published: (2026)
by: Gupta, Tanmay, et al.
Published: (2026)
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
by: Clark, Christopher, et al.
Published: (2026)
by: Clark, Christopher, et al.
Published: (2026)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
by: Zhang, Tianyi, et al.
Published: (2026)
by: Zhang, Tianyi, et al.
Published: (2026)
We're Not Using Videos Effectively: An Updated Domain Adaptive Video Segmentation Baseline
by: Kareer, Simar, et al.
Published: (2024)
by: Kareer, Simar, et al.
Published: (2024)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation
by: Duggal, Shivam, et al.
Published: (2025)
by: Duggal, Shivam, et al.
Published: (2025)
AUGCAL: Improving Sim2Real Adaptation by Uncertainty Calibration on Augmented Synthetic Images
by: Chattopadhyay, Prithvijit, et al.
Published: (2023)
by: Chattopadhyay, Prithvijit, et al.
Published: (2023)
Iterated Learning Improves Compositionality in Large Vision-Language Models
by: Zheng, Chenhao, et al.
Published: (2024)
by: Zheng, Chenhao, et al.
Published: (2024)
Divergence is Uncertainty: A Closed-Form Posterior Covariance for Flow Matching
by: Xing, Jiarui, et al.
Published: (2026)
by: Xing, Jiarui, et al.
Published: (2026)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
by: Clark, Christopher, et al.
Published: (2026)
by: Clark, Christopher, et al.
Published: (2026)
Task Me Anything
by: Zhang, Jieyu, et al.
Published: (2024)
by: Zhang, Jieyu, et al.
Published: (2024)
Synthetic Visual Genome
by: Park, Jae Sung, et al.
Published: (2025)
by: Park, Jae Sung, et al.
Published: (2025)
Semi-Truths: A Large-Scale Dataset of AI-Augmented Images for Evaluating Robustness of AI-Generated Image detectors
by: Pal, Anisha, et al.
Published: (2024)
by: Pal, Anisha, et al.
Published: (2024)
Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
by: Gao, Ziqi, et al.
Published: (2026)
by: Gao, Ziqi, et al.
Published: (2026)
RAFM: Retrieval-Augmented Flow Matching for Unpaired CBCT-to-CT Translation
by: Zhou, Xianhao, et al.
Published: (2026)
by: Zhou, Xianhao, et al.
Published: (2026)
MPFlow: Multi-modal Posterior-Guided Flow Matching for Zero-Shot MRI Reconstruction
by: Kim, Seunghoi, et al.
Published: (2026)
by: Kim, Seunghoi, et al.
Published: (2026)
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
by: Fan, Xiang, et al.
Published: (2024)
by: Fan, Xiang, et al.
Published: (2024)
Unbiased Diffusion Variational Inversion via Principled Posterior Matching
by: Bai, Weimin, et al.
Published: (2026)
by: Bai, Weimin, et al.
Published: (2026)
Bytes Are All You Need: Transformers Operating Directly On File Bytes
by: Horton, Maxwell, et al.
Published: (2023)
by: Horton, Maxwell, et al.
Published: (2023)
Similar Items
-
Contrastive Flow Matching
by: Stoica, George, et al.
Published: (2025) -
The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs Better
by: Geng, Scott, et al.
Published: (2024) -
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
by: Yadav, Tanush, et al.
Published: (2026) -
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
by: Wallingford, Matthew, et al.
Published: (2024) -
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
by: Ramanujan, Vivek, et al.
Published: (2024)