Dense Semantic Matching with VGGT Prior
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Songlin, Wei, Tianyi, Lan, Yushi, Xiao, Zeqi, Rao, Anyi, Pan, Xingang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Textured 3D Regenerative Morphing with 3D Diffusion Prior
by: Yang, Songlin, et al.
Published: (2025)
by: Yang, Songlin, et al.
Published: (2025)
Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
WorldMem: Long-term Consistent World Simulation with Memory
by: Xiao, Zeqi, et al.
Published: (2025)
by: Xiao, Zeqi, et al.
Published: (2025)
MvDrag3D: Drag-based Creative 3D Editing via Multi-view Generation-Reconstruction Priors
by: Chen, Honghua, et al.
Published: (2024)
by: Chen, Honghua, et al.
Published: (2024)
Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
Video Diffusion Models are Training-free Motion Interpreter and Controller
by: Xiao, Zeqi, et al.
Published: (2024)
by: Xiao, Zeqi, et al.
Published: (2024)
Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models
by: Yang, Songlin, et al.
Published: (2026)
by: Yang, Songlin, et al.
Published: (2026)
VGGT-X: When VGGT Meets Dense Novel View Synthesis
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
PnP-U3D: Plug-and-Play 3D Framework Bridging Autoregression and Diffusion for Unified Understanding and Generation
by: Chen, Yongwei, et al.
Published: (2026)
by: Chen, Yongwei, et al.
Published: (2026)
ArtiLatent: Realistic Articulated 3D Object Generation via Structured Latents
by: Chen, Honghua, et al.
Published: (2025)
by: Chen, Honghua, et al.
Published: (2025)
FastMesh: Efficient Artistic Mesh Generation via Component Decoupling
by: Kim, Jeonghwan, et al.
Published: (2025)
by: Kim, Jeonghwan, et al.
Published: (2025)
SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAE
by: Chen, Yongwei, et al.
Published: (2024)
by: Chen, Yongwei, et al.
Published: (2024)
SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural Alignment
by: Zhao, Zhuoran, et al.
Published: (2026)
by: Zhao, Zhuoran, et al.
Published: (2026)
4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere
by: Luo, Yihang, et al.
Published: (2026)
by: Luo, Yihang, et al.
Published: (2026)
3DEnhancer: Consistent Multi-View Diffusion for 3D Enhancement
by: Luo, Yihang, et al.
Published: (2024)
by: Luo, Yihang, et al.
Published: (2024)
Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation
by: Xiao, Zeqi, et al.
Published: (2025)
by: Xiao, Zeqi, et al.
Published: (2025)
Trajectory Attention for Fine-grained Video Motion Control
by: Xiao, Zeqi, et al.
Published: (2024)
by: Xiao, Zeqi, et al.
Published: (2024)
VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
by: Cao, Yang, et al.
Published: (2026)
by: Cao, Yang, et al.
Published: (2026)
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
by: Peng, Haosong, et al.
Published: (2025)
by: Peng, Haosong, et al.
Published: (2025)
Boosting Monocular Metric Depth Estimation via Bokeh Rendering
by: Zhang, Hangwei, et al.
Published: (2025)
by: Zhang, Hangwei, et al.
Published: (2025)
FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing
by: Wei, Tianyi, et al.
Published: (2025)
by: Wei, Tianyi, et al.
Published: (2025)
Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation
by: Wei, Tianyi, et al.
Published: (2024)
by: Wei, Tianyi, et al.
Published: (2024)
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
by: Shu, Zhijian, et al.
Published: (2025)
by: Shu, Zhijian, et al.
Published: (2025)
VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
by: Maggio, Dominic, et al.
Published: (2025)
by: Maggio, Dominic, et al.
Published: (2025)
GaussianAnything: Interactive Point Cloud Flow Matching For 3D Object Generation
by: Lan, Yushi, et al.
Published: (2024)
by: Lan, Yushi, et al.
Published: (2024)
TokensGen: Harnessing Condensed Tokens for Long Video Generation
by: Ouyang, Wenqi, et al.
Published: (2025)
by: Ouyang, Wenqi, et al.
Published: (2025)
Bokeh Diffusion: Defocus Blur Control in Text-to-Image Diffusion Models
by: Fortes, Armando, et al.
Published: (2025)
by: Fortes, Armando, et al.
Published: (2025)
TableCenterNet: A one-stage network for table structure recognition
by: Xiao, Anyi, et al.
Published: (2025)
by: Xiao, Anyi, et al.
Published: (2025)
VGGT-$Ω$
by: Wang, Jianyuan, et al.
Published: (2026)
by: Wang, Jianyuan, et al.
Published: (2026)
MVIP-NeRF: Multi-view 3D Inpainting on NeRF Scenes via Diffusion Prior
by: Chen, Honghua, et al.
Published: (2024)
by: Chen, Honghua, et al.
Published: (2024)
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
by: Sun, Xiangyu, et al.
Published: (2026)
by: Sun, Xiangyu, et al.
Published: (2026)
From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space
by: Bu, Jiazi, et al.
Published: (2026)
by: Bu, Jiazi, et al.
Published: (2026)
VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
by: Deng, Kai, et al.
Published: (2025)
by: Deng, Kai, et al.
Published: (2025)
VGGT-SLAM 2.0: Real-time Dense Feed-forward Scene Reconstruction
by: Maggio, Dominic, et al.
Published: (2026)
by: Maggio, Dominic, et al.
Published: (2026)
VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments
by: Xu, Jingyi, et al.
Published: (2026)
by: Xu, Jingyi, et al.
Published: (2026)
FrameVGGT: Geometry-Aligned Frame-Level Memory for Bounded Streaming VGGT
by: Xu, Zhisong, et al.
Published: (2026)
by: Xu, Zhisong, et al.
Published: (2026)
VGGT-SLAM++
by: Mandal, Avilasha, et al.
Published: (2026)
by: Mandal, Avilasha, et al.
Published: (2026)
PI-Light: Physics-Inspired Diffusion for Full-Image Relighting
by: Liang, Zhexin, et al.
Published: (2026)
by: Liang, Zhexin, et al.
Published: (2026)
Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures
by: Wang, Yuxi, et al.
Published: (2026)
by: Wang, Yuxi, et al.
Published: (2026)
Beyond Inserting: Learning Identity Embedding for Semantic-Fidelity Personalized Diffusion Generation
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
Similar Items
-
Textured 3D Regenerative Morphing with 3D Diffusion Prior
by: Yang, Songlin, et al.
Published: (2025) -
Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
by: Zhou, Yifan, et al.
Published: (2025) -
WorldMem: Long-term Consistent World Simulation with Memory
by: Xiao, Zeqi, et al.
Published: (2025) -
MvDrag3D: Drag-based Creative 3D Editing via Multi-view Generation-Reconstruction Priors
by: Chen, Honghua, et al.
Published: (2024) -
Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space
by: Zhou, Yifan, et al.
Published: (2025)