E-M3RF: An Equivariant Multimodal 3D Re-assembly Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Islam, Adeela, Fiorini, Stefano, Lecha, Manuel, Tsesmelis, Theodore, James, Stuart, Morerio, Pietro, Del Bue, Alessio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909926625902592
author Islam, Adeela
Fiorini, Stefano
Lecha, Manuel
Tsesmelis, Theodore
James, Stuart
Morerio, Pietro
Del Bue, Alessio
author_facet Islam, Adeela
Fiorini, Stefano
Lecha, Manuel
Tsesmelis, Theodore
James, Stuart
Morerio, Pietro
Del Bue, Alessio
contents 3D reassembly is a fundamental geometric problem, and in recent years it has increasingly been challenged by deep learning methods rather than classical optimization. While learning approaches have shown promising results, most still rely primarily on geometric features to assemble a whole from its parts. As a result, methods struggle when geometry alone is insufficient or ambiguous, for example, for small, eroded, or symmetric fragments. Additionally, solutions do not impose physical constraints that explicitly prevent overlapping assemblies. To address these limitations, we introduce E-M3RF, an equivariant multimodal 3D reassembly framework that takes as input the point clouds, containing both point positions and colors of fractured fragments, and predicts the transformations required to reassemble them using SE(3) flow matching. Each fragment is represented by both geometric and color features: i) 3D point positions are encoded as rotationconsistent geometric features using a rotation-equivariant encoder, ii) the colors at each 3D point are encoded with a transformer. The two feature sets are then combined to form a multimodal representation. We experimented on four datasets: two synthetic datasets, Breaking Bad and Fantastic Breaks, and two real-world cultural heritage datasets, RePAIR and Presious, demonstrating that E-M3RF on the RePAIR dataset reduces rotation error by 23.1% and translation error by 13.2%, while Chamfer Distance decreases by 18.4% compared to competing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21422
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle E-M3RF: An Equivariant Multimodal 3D Re-assembly Framework
Islam, Adeela
Fiorini, Stefano
Lecha, Manuel
Tsesmelis, Theodore
James, Stuart
Morerio, Pietro
Del Bue, Alessio
Computer Vision and Pattern Recognition
3D reassembly is a fundamental geometric problem, and in recent years it has increasingly been challenged by deep learning methods rather than classical optimization. While learning approaches have shown promising results, most still rely primarily on geometric features to assemble a whole from its parts. As a result, methods struggle when geometry alone is insufficient or ambiguous, for example, for small, eroded, or symmetric fragments. Additionally, solutions do not impose physical constraints that explicitly prevent overlapping assemblies. To address these limitations, we introduce E-M3RF, an equivariant multimodal 3D reassembly framework that takes as input the point clouds, containing both point positions and colors of fractured fragments, and predicts the transformations required to reassemble them using SE(3) flow matching. Each fragment is represented by both geometric and color features: i) 3D point positions are encoded as rotationconsistent geometric features using a rotation-equivariant encoder, ii) the colors at each 3D point are encoded with a transformer. The two feature sets are then combined to form a multimodal representation. We experimented on four datasets: two synthetic datasets, Breaking Bad and Fantastic Breaks, and two real-world cultural heritage datasets, RePAIR and Presious, demonstrating that E-M3RF on the RePAIR dataset reduces rotation error by 23.1% and translation error by 13.2%, while Chamfer Distance decreases by 18.4% compared to competing methods.
title E-M3RF: An Equivariant Multimodal 3D Re-assembly Framework
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.21422