PartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction Model
Fuente:
arXiv
Guardado en:
| Autores principales: | Gao, Mingju, Pan, Yike, Gao, Huan-ang, Zhang, Zongzheng, Li, Wenyi, Dong, Hao, Tang, Hao, Yi, Li, Zhao, Hao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Training-Free Model Merging for Multi-target Domain Adaptation
por: Li, Wenyi, et al.
Publicado: (2024)
por: Li, Wenyi, et al.
Publicado: (2024)
SCP-Diff: Spatial-Categorical Joint Prior for Diffusion Based Semantic Image Synthesis
por: Gao, Huan-ang, et al.
Publicado: (2024)
por: Gao, Huan-ang, et al.
Publicado: (2024)
Benchmarking PhD-Level Coding in 3D Geometric Computer Vision
por: Li, Wenyi, et al.
Publicado: (2026)
por: Li, Wenyi, et al.
Publicado: (2026)
FairDiff: Fair Segmentation with Point-Image Diffusion
por: Li, Wenyi, et al.
Publicado: (2024)
por: Li, Wenyi, et al.
Publicado: (2024)
FB-4D: Spatial-Temporal Coherent Dynamic 3D Content Generation with Feature Banks
por: Li, Jinwei, et al.
Publicado: (2025)
por: Li, Jinwei, et al.
Publicado: (2025)
PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation
por: Gao, Mingju, et al.
Publicado: (2026)
por: Gao, Mingju, et al.
Publicado: (2026)
PartRAG: Retrieval-Augmented Part-Level 3D Generation and Editing
por: Li, Peize, et al.
Publicado: (2026)
por: Li, Peize, et al.
Publicado: (2026)
PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and Environments
por: Ding, Kairui, et al.
Publicado: (2024)
por: Ding, Kairui, et al.
Publicado: (2024)
Challenger: Affordable Adversarial Driving Video Generation
por: Xu, Zhiyuan, et al.
Publicado: (2025)
por: Xu, Zhiyuan, et al.
Publicado: (2025)
Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward Modeling
por: Zhang, Guiyu, et al.
Publicado: (2024)
por: Zhang, Guiyu, et al.
Publicado: (2024)
Dual-frame Fluid Motion Estimation with Test-time Optimization and Zero-divergence Loss
por: Zhang, Yifei, et al.
Publicado: (2024)
por: Zhang, Yifei, et al.
Publicado: (2024)
Alias-free 4D Gaussian Splatting
por: Chen, Zilong, et al.
Publicado: (2025)
por: Chen, Zilong, et al.
Publicado: (2025)
TwinAligner: Visual-Dynamic Alignment Empowers Physics-aware Real2Sim2Real for Robotic Manipulation
por: Fan, Hongwei, et al.
Publicado: (2025)
por: Fan, Hongwei, et al.
Publicado: (2025)
RGM: Reconstructing High-fidelity 3D Car Assets with Relightable 3D-GS Generative Model from a Single Image
por: Chen, Xiaoxue, et al.
Publicado: (2024)
por: Chen, Xiaoxue, et al.
Publicado: (2024)
Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
por: Chi, Haohan, et al.
Publicado: (2025)
por: Chi, Haohan, et al.
Publicado: (2025)
Ultraman: Single Image 3D Human Reconstruction with Ultra Speed and Detail
por: Chen, Mingjin, et al.
Publicado: (2024)
por: Chen, Mingjin, et al.
Publicado: (2024)
Diffusion-based Visual Anagram as Multi-task Learning
por: Xu, Zhiyuan, et al.
Publicado: (2024)
por: Xu, Zhiyuan, et al.
Publicado: (2024)
Reusing Attention for One-stage Lane Topology Understanding
por: Li, Yang, et al.
Publicado: (2025)
por: Li, Yang, et al.
Publicado: (2025)
Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation
por: Zhu, Xiaomeng, et al.
Publicado: (2025)
por: Zhu, Xiaomeng, et al.
Publicado: (2025)
TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
por: Zhang, Zongzheng, et al.
Publicado: (2025)
por: Zhang, Zongzheng, et al.
Publicado: (2025)
SUM Parts: Benchmarking Part-Level Semantic Segmentation of Urban Meshes
por: Gao, Weixiao, et al.
Publicado: (2025)
por: Gao, Weixiao, et al.
Publicado: (2025)
GaussianArt: Unified Modeling of Geometry and Motion for Articulated Objects
por: Shen, Licheng, et al.
Publicado: (2025)
por: Shen, Licheng, et al.
Publicado: (2025)
Delving into Mapping Uncertainty for Mapless Trajectory Prediction
por: Zhang, Zongzheng, et al.
Publicado: (2025)
por: Zhang, Zongzheng, et al.
Publicado: (2025)
P-MapNet: Far-seeing Map Generator Enhanced by both SDMap and HDMap Priors
por: Jiang, Zhou, et al.
Publicado: (2024)
por: Jiang, Zhou, et al.
Publicado: (2024)
UniUncer: Unified Dynamic Static Uncertainty for End to End Driving
por: Gao, Yu, et al.
Publicado: (2026)
por: Gao, Yu, et al.
Publicado: (2026)
CRUISE: Cooperative Reconstruction and Editing in V2X Scenarios using Gaussian Splatting
por: Xu, Haoran, et al.
Publicado: (2025)
por: Xu, Haoran, et al.
Publicado: (2025)
Locate n' Rotate: Two-stage Openable Part Detection with Foundation Model Priors
por: Li, Siqi, et al.
Publicado: (2024)
por: Li, Siqi, et al.
Publicado: (2024)
AVD2: Accident Video Diffusion for Accident Video Description
por: Li, Cheng, et al.
Publicado: (2025)
por: Li, Cheng, et al.
Publicado: (2025)
FuRPE: Learning Full-body Reconstruction from Part Experts
por: Fan, Zhaoxin, et al.
Publicado: (2022)
por: Fan, Zhaoxin, et al.
Publicado: (2022)
Adaptive Articulated Object Manipulation On The Fly with Foundation Model Reasoning and Part Grounding
por: Zhang, Xiaojie, et al.
Publicado: (2025)
por: Zhang, Xiaojie, et al.
Publicado: (2025)
SA-GS: Scale-Adaptive Gaussian Splatting for Training-Free Anti-Aliasing
por: Song, Xiaowei, et al.
Publicado: (2024)
por: Song, Xiaowei, et al.
Publicado: (2024)
Dynamic Scene Reconstruction: Recent Advance in Real-time Rendering and Streaming
por: Zhu, Jiaxuan, et al.
Publicado: (2025)
por: Zhu, Jiaxuan, et al.
Publicado: (2025)
Part123: Part-aware 3D Reconstruction from a Single-view Image
por: Liu, Anran, et al.
Publicado: (2024)
por: Liu, Anran, et al.
Publicado: (2024)
TUNI: Real-time RGB-T Semantic Segmentation with Unified Multi-Modal Feature Extraction and Cross-Modal Feature Fusion
por: Guo, Xiaodong, et al.
Publicado: (2025)
por: Guo, Xiaodong, et al.
Publicado: (2025)
DOEPatch: Dynamically Optimized Ensemble Model for Adversarial Patches Generation
por: Tan, Wenyi, et al.
Publicado: (2023)
por: Tan, Wenyi, et al.
Publicado: (2023)
FoundIR: Unleashing Million-scale Training Data to Advance Foundation Models for Image Restoration
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding
por: Blume, Ansel, et al.
Publicado: (2025)
por: Blume, Ansel, et al.
Publicado: (2025)
Scalable Visual State Space Model with Fractal Scanning
por: Tang, Lv, et al.
Publicado: (2024)
por: Tang, Lv, et al.
Publicado: (2024)
Post-hoc Part-prototype Networks
por: Tan, Andong, et al.
Publicado: (2024)
por: Tan, Andong, et al.
Publicado: (2024)
PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth
por: Jin, Bu, et al.
Publicado: (2025)
por: Jin, Bu, et al.
Publicado: (2025)
Ejemplares similares
-
Training-Free Model Merging for Multi-target Domain Adaptation
por: Li, Wenyi, et al.
Publicado: (2024) -
SCP-Diff: Spatial-Categorical Joint Prior for Diffusion Based Semantic Image Synthesis
por: Gao, Huan-ang, et al.
Publicado: (2024) -
Benchmarking PhD-Level Coding in 3D Geometric Computer Vision
por: Li, Wenyi, et al.
Publicado: (2026) -
FairDiff: Fair Segmentation with Point-Image Diffusion
por: Li, Wenyi, et al.
Publicado: (2024) -
FB-4D: Spatial-Temporal Coherent Dynamic 3D Content Generation with Feature Banks
por: Li, Jinwei, et al.
Publicado: (2025)