Saved in:
Bibliographic Details
Main Authors: Li, Haoyang, You, Yang, Su, Hao, Guibas, Leonidas
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.20323
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918478803369984
author Li, Haoyang
You, Yang
Su, Hao
Guibas, Leonidas
author_facet Li, Haoyang
You, Yang
Su, Hao
Guibas, Leonidas
contents Reliable object manipulation requires understanding physical properties that vary across objects and environments. Vision-language model (VLM) planners can reason about friction and stability in general terms; however, they often cannot predict how a specific ball will roll on a particular surface or which stone will provide a stable foundation without direct experience. We present PhysMem, a memory framework that enables VLM robot planners to learn physical principles from interaction at test time, without updating model parameters. The system records experiences, generates candidate hypotheses, and verifies them through targeted interaction before promoting validated knowledge to guide future decisions. A central design choice is verification before application: the system tests hypotheses against new observations rather than applying retrieved experience directly, reducing rigid reliance on prior experience when physical conditions change. We evaluate PhysMem on three real-world manipulation tasks and simulation benchmarks across four VLM backbones. On a controlled brick insertion task, principled abstraction achieves 76% success compared to 23% for direct experience retrieval, and real-world experiments show consistent improvement over 30-minute deployment sessions.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20323
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning
Li, Haoyang
You, Yang
Su, Hao
Guibas, Leonidas
Robotics
Artificial Intelligence
Reliable object manipulation requires understanding physical properties that vary across objects and environments. Vision-language model (VLM) planners can reason about friction and stability in general terms; however, they often cannot predict how a specific ball will roll on a particular surface or which stone will provide a stable foundation without direct experience. We present PhysMem, a memory framework that enables VLM robot planners to learn physical principles from interaction at test time, without updating model parameters. The system records experiences, generates candidate hypotheses, and verifies them through targeted interaction before promoting validated knowledge to guide future decisions. A central design choice is verification before application: the system tests hypotheses against new observations rather than applying retrieved experience directly, reducing rigid reliance on prior experience when physical conditions change. We evaluate PhysMem on three real-world manipulation tasks and simulation benchmarks across four VLM backbones. On a controlled brick insertion task, principled abstraction achieves 76% success compared to 23% for direct experience retrieval, and real-world experiments show consistent improvement over 30-minute deployment sessions.
title PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2602.20323