Φeat: Physically-Grounded Feature Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vecchio, Giuseppe, Kaiser, Adrien, Romain, Rouffet, Martin, Rosalie, Garces, Elena, Boubekeur, Tamy
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912709177507840
author Vecchio, Giuseppe
Kaiser, Adrien
Romain, Rouffet
Martin, Rosalie
Garces, Elena
Boubekeur, Tamy
author_facet Vecchio, Giuseppe
Kaiser, Adrien
Romain, Rouffet
Martin, Rosalie
Garces, Elena
Boubekeur, Tamy
contents Foundation models have emerged as effective backbones for many vision tasks. However, current self-supervised features entangle high-level semantics with low-level physical factors, such as geometry and illumination, hindering their use in tasks requiring explicit physical reasoning. In this paper, we introduce $Φ$eat, a novel physically-grounded visual backbone that encourages a representation sensitive to material identity, including reflectance cues and geometric mesostructure. Our key idea is to employ a pretraining strategy that contrasts spatial crops and physical augmentations of the same material under varying shapes and lighting conditions. While similar data have been used in high-end supervised tasks such as intrinsic decomposition or material estimation, we demonstrate that a pure self-supervised training strategy, without explicit labels, already provides a strong prior for tasks requiring robust features invariant to external physical factors. We evaluate the learned representations through feature similarity analysis and material selection, showing that $Φ$eat captures physically-grounded structure beyond semantic grouping. These findings highlight the promise of unsupervised physical feature learning as a foundation for physics-aware perception in vision and graphics. These findings highlight the promise of unsupervised physical feature learning as a foundation for physics-aware perception in vision and graphics.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11270
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Φeat: Physically-Grounded Feature Representation
Vecchio, Giuseppe
Kaiser, Adrien
Romain, Rouffet
Martin, Rosalie
Garces, Elena
Boubekeur, Tamy
Computer Vision and Pattern Recognition
Foundation models have emerged as effective backbones for many vision tasks. However, current self-supervised features entangle high-level semantics with low-level physical factors, such as geometry and illumination, hindering their use in tasks requiring explicit physical reasoning. In this paper, we introduce $Φ$eat, a novel physically-grounded visual backbone that encourages a representation sensitive to material identity, including reflectance cues and geometric mesostructure. Our key idea is to employ a pretraining strategy that contrasts spatial crops and physical augmentations of the same material under varying shapes and lighting conditions. While similar data have been used in high-end supervised tasks such as intrinsic decomposition or material estimation, we demonstrate that a pure self-supervised training strategy, without explicit labels, already provides a strong prior for tasks requiring robust features invariant to external physical factors. We evaluate the learned representations through feature similarity analysis and material selection, showing that $Φ$eat captures physically-grounded structure beyond semantic grouping. These findings highlight the promise of unsupervised physical feature learning as a foundation for physics-aware perception in vision and graphics. These findings highlight the promise of unsupervised physical feature learning as a foundation for physics-aware perception in vision and graphics.
title Φeat: Physically-Grounded Feature Representation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.11270