VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | An, Zhaochong, Kupyn, Orest, Uscidda, Théo, Colaco, Andrea, Ahuja, Karan, Belongie, Serge, Gonzalez-Franco, Mar, Gazulla, Marta Tintore |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dataset Enhancement with Instance-Level Augmentations
von: Kupyn, Orest, et al.
Veröffentlicht: (2024)
von: Kupyn, Orest, et al.
Veröffentlicht: (2024)
S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
von: Kupyn, Orest, et al.
Veröffentlicht: (2025)
von: Kupyn, Orest, et al.
Veröffentlicht: (2025)
VGGHeads: 3D Multi Head Alignment with a Large-Scale Synthetic Dataset
von: Kupyn, Orest, et al.
Veröffentlicht: (2024)
von: Kupyn, Orest, et al.
Veröffentlicht: (2024)
Epipolar Geometry Improves Video Generation Models
von: Kupyn, Orest, et al.
Veröffentlicht: (2025)
von: Kupyn, Orest, et al.
Veröffentlicht: (2025)
SurfaceXR: Fusing Smartwatch IMUs and Egocentric Hand Pose for Seamless Surface Interactions
von: Xu, Vasco, et al.
Veröffentlicht: (2026)
von: Xu, Vasco, et al.
Veröffentlicht: (2026)
Geometry Fidelity for Spherical Images
von: Christensen, Anders, et al.
Veröffentlicht: (2024)
von: Christensen, Anders, et al.
Veröffentlicht: (2024)
Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model
von: An, Zhaochong, et al.
Veröffentlicht: (2025)
von: An, Zhaochong, et al.
Veröffentlicht: (2025)
MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
von: Segu, Mattia, et al.
Veröffentlicht: (2025)
von: Segu, Mattia, et al.
Veröffentlicht: (2025)
Augmented Object Intelligence with XR-Objects
von: Dogan, Mustafa Doga, et al.
Veröffentlicht: (2024)
von: Dogan, Mustafa Doga, et al.
Veröffentlicht: (2024)
DAD-3DHeads: A Large-scale Dense, Accurate and Diverse Dataset for 3D Head Alignment from a Single Image
von: Martyniuk, Tetiana, et al.
Veröffentlicht: (2022)
von: Martyniuk, Tetiana, et al.
Veröffentlicht: (2022)
PARSE-Ego4D: Personal Action Recommendation Suggestions for Egocentric Videos
von: Abreu, Steven, et al.
Veröffentlicht: (2024)
von: Abreu, Steven, et al.
Veröffentlicht: (2024)
Rethinking Few-shot 3D Point Cloud Semantic Segmentation
von: An, Zhaochong, et al.
Veröffentlicht: (2024)
von: An, Zhaochong, et al.
Veröffentlicht: (2024)
Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation
von: An, Zhaochong, et al.
Veröffentlicht: (2024)
von: An, Zhaochong, et al.
Veröffentlicht: (2024)
GeOT: A spatially explicit framework for evaluating spatio-temporal predictions
von: Wiedemann, Nina, et al.
Veröffentlicht: (2024)
von: Wiedemann, Nina, et al.
Veröffentlicht: (2024)
PoseDreamer: Scalable and Photorealistic Human Data Generation Pipeline with Diffusion Models
von: Prospero, Lorenza, et al.
Veröffentlicht: (2026)
von: Prospero, Lorenza, et al.
Veröffentlicht: (2026)
ChatMotion: A Multimodal Multi-Agent for Human Motion Analysis
von: Li, Lei, et al.
Veröffentlicht: (2025)
von: Li, Lei, et al.
Veröffentlicht: (2025)
Revisiting the Perception-Distortion Trade-off with Spatial-Semantic Guided Super-Resolution
von: Wang, Dan, et al.
Veröffentlicht: (2026)
von: Wang, Dan, et al.
Veröffentlicht: (2026)
Shifting the Breaking Point of Flow Matching for Multi-Instance Editing
von: Zaccagnino, Carmine, et al.
Veröffentlicht: (2026)
von: Zaccagnino, Carmine, et al.
Veröffentlicht: (2026)
Practical and Rich User Digitization
von: Ahuja, Karan
Veröffentlicht: (2024)
von: Ahuja, Karan
Veröffentlicht: (2024)
Symbiotic AI: Augmenting Human Cognition from PCs to Cars
von: Bovo, Riccardo, et al.
Veröffentlicht: (2025)
von: Bovo, Riccardo, et al.
Veröffentlicht: (2025)
The Latent Color Subspace: Emergent Order in High-Dimensional Chaos
von: Pach, Mateusz, et al.
Veröffentlicht: (2026)
von: Pach, Mateusz, et al.
Veröffentlicht: (2026)
Generalized Discrete Diffusion from Snapshots
von: Zekri, Oussama, et al.
Veröffentlicht: (2026)
von: Zekri, Oussama, et al.
Veröffentlicht: (2026)
GENOT: Entropic (Gromov) Wasserstein Flow Matching with Applications to Single-Cell Genomics
von: Klein, Dominik, et al.
Veröffentlicht: (2023)
von: Klein, Dominik, et al.
Veröffentlicht: (2023)
Noise-Coded Illumination for Forensic and Photometric Video Analysis
von: Michael, Peter F., et al.
Veröffentlicht: (2025)
von: Michael, Peter F., et al.
Veröffentlicht: (2025)
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion
von: Tian, Junjiao, et al.
Veröffentlicht: (2023)
von: Tian, Junjiao, et al.
Veröffentlicht: (2023)
Video Understanding: From Geometry and Semantics to Unified Models
von: An, Zhaochong, et al.
Veröffentlicht: (2026)
von: An, Zhaochong, et al.
Veröffentlicht: (2026)
EmBARDiment: an Embodied AI Agent for Productivity in XR
von: Bovo, Riccardo, et al.
Veröffentlicht: (2024)
von: Bovo, Riccardo, et al.
Veröffentlicht: (2024)
PhysConvex: Physics-Informed 3D Dynamic Convex Radiance Fields for Reconstruction and Simulation
von: Wang, Dan, et al.
Veröffentlicht: (2026)
von: Wang, Dan, et al.
Veröffentlicht: (2026)
Unlearning-based Neural Interpretations
von: Choi, Ching Lam, et al.
Veröffentlicht: (2024)
von: Choi, Ching Lam, et al.
Veröffentlicht: (2024)
Large Vision-Language Models for Knowledge-Grounded Data Annotation of Memes
von: Deng, Shiling, et al.
Veröffentlicht: (2025)
von: Deng, Shiling, et al.
Veröffentlicht: (2025)
Reward Guided Latent Consistency Distillation
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
von: Li, Jiachen, et al.
Veröffentlicht: (2024)
Real-time surface current data in the Ibiza Channel from January to October 2016
von: Tintore, Joaquín
Veröffentlicht: (2016)
von: Tintore, Joaquín
Veröffentlicht: (2016)
Panel on Synthesizing BPs .
von: Tintore, Joaquin
Veröffentlicht: (2019)
von: Tintore, Joaquin
Veröffentlicht: (2019)
From Videos to Conversations: Egocentric Instructions for Task Assistance
von: Aggarwal, Lavisha, et al.
Veröffentlicht: (2026)
von: Aggarwal, Lavisha, et al.
Veröffentlicht: (2026)
Stitched Value Model for Diffusion Alignment
von: Go, Hyojun, et al.
Veröffentlicht: (2026)
von: Go, Hyojun, et al.
Veröffentlicht: (2026)
OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
von: An, Zhaochong, et al.
Veröffentlicht: (2025)
von: An, Zhaochong, et al.
Veröffentlicht: (2025)
MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training
von: Gordon, Lucia, et al.
Veröffentlicht: (2026)
von: Gordon, Lucia, et al.
Veröffentlicht: (2026)
Text Entry for XR Trove (TEXT): Collecting and Analyzing Techniques for Text Input in XR
von: Bhatia, Arpit, et al.
Veröffentlicht: (2025)
von: Bhatia, Arpit, et al.
Veröffentlicht: (2025)
On the structural stability of a simple cosmological model in $R+αR^{2}$ theory of gravity
von: Hrycyna, Orest
Veröffentlicht: (2022)
von: Hrycyna, Orest
Veröffentlicht: (2022)
Ähnliche Einträge
-
Dataset Enhancement with Instance-Level Augmentations
von: Kupyn, Orest, et al.
Veröffentlicht: (2024) -
S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
von: Kupyn, Orest, et al.
Veröffentlicht: (2025) -
VGGHeads: 3D Multi Head Alignment with a Large-Scale Synthetic Dataset
von: Kupyn, Orest, et al.
Veröffentlicht: (2024) -
Epipolar Geometry Improves Video Generation Models
von: Kupyn, Orest, et al.
Veröffentlicht: (2025) -
SurfaceXR: Fusing Smartwatch IMUs and Egocentric Hand Pose for Seamless Surface Interactions
von: Xu, Vasco, et al.
Veröffentlicht: (2026)