Beyond Static Perception: Integrating Temporal Context into VLMs for Cloth Folding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Barbany, Oriol, Colomé, Adrià, Torras, Carme |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BiFold: Bimanual Cloth Folding with Language Guidance
von: Barbany, Oriol, et al.
Veröffentlicht: (2025)
von: Barbany, Oriol, et al.
Veröffentlicht: (2025)
Benchmarking the Sim-to-Real Gap in Cloth Manipulation
von: Blanco-Mulero, David, et al.
Veröffentlicht: (2023)
von: Blanco-Mulero, David, et al.
Veröffentlicht: (2023)
Dataset for BiFold: Bimanual Cloth Folding with Language Guidance
von: Barbany, Oriol, et al.
Veröffentlicht: (2025)
von: Barbany, Oriol, et al.
Veröffentlicht: (2025)
CloSE: A Geometric Shape-Agnostic Cloth State Representation
von: Kamat, Jay, et al.
Veröffentlicht: (2025)
von: Kamat, Jay, et al.
Veröffentlicht: (2025)
Learning a General Model: Folding Clothing with Topological Dynamics
von: Liu, Yiming, et al.
Veröffentlicht: (2025)
von: Liu, Yiming, et al.
Veröffentlicht: (2025)
Cloth-Splatting: 3D Cloth State Estimation from RGB Supervision
von: Longhini, Alberta, et al.
Veröffentlicht: (2025)
von: Longhini, Alberta, et al.
Veröffentlicht: (2025)
From Steering to Pedalling: Do Autonomous Driving VLMs Generalize to Cyclist-Assistive Spatial Perception and Planning?
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2026)
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2026)
From Static to Dynamic: a Survey of Topology-Aware Perception in Autonomous Driving
von: Chen, Yixiao, et al.
Veröffentlicht: (2025)
von: Chen, Yixiao, et al.
Veröffentlicht: (2025)
Leveraging Large Language Models for Multimodal Search
von: Barbany, Oriol, et al.
Veröffentlicht: (2024)
von: Barbany, Oriol, et al.
Veröffentlicht: (2024)
Temporally Consistent Unsupervised Segmentation for Mobile Robot Perception
von: Ellis, Christian, et al.
Veröffentlicht: (2025)
von: Ellis, Christian, et al.
Veröffentlicht: (2025)
Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving
von: Tang, Zecong, et al.
Veröffentlicht: (2026)
von: Tang, Zecong, et al.
Veröffentlicht: (2026)
FoldNet: Learning Generalizable Closed-Loop Policy for Garment Folding via Keypoint-Driven Asset and Demonstration Synthesis
von: Chen, Yuxing, et al.
Veröffentlicht: (2025)
von: Chen, Yuxing, et al.
Veröffentlicht: (2025)
Flat'n'Fold: A Diverse Multi-Modal Dataset for Garment Perception and Manipulation
von: Zhuang, Lipeng, et al.
Veröffentlicht: (2024)
von: Zhuang, Lipeng, et al.
Veröffentlicht: (2024)
STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from Gaits
von: Bhattacharya, Uttaran, et al.
Veröffentlicht: (2019)
von: Bhattacharya, Uttaran, et al.
Veröffentlicht: (2019)
SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation
von: Munje, Michael J., et al.
Veröffentlicht: (2025)
von: Munje, Michael J., et al.
Veröffentlicht: (2025)
ST-Booster: An Iterative SpatioTemporal Perception Booster for Vision-and-Language Navigation in Continuous Environments
von: Yue, Lu, et al.
Veröffentlicht: (2025)
von: Yue, Lu, et al.
Veröffentlicht: (2025)
Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning
von: Li, Yue, et al.
Veröffentlicht: (2025)
von: Li, Yue, et al.
Veröffentlicht: (2025)
AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives
von: Xie, Shaoyuan, et al.
Veröffentlicht: (2025)
von: Xie, Shaoyuan, et al.
Veröffentlicht: (2025)
CRASH: Crash Recognition and Anticipation System Harnessing with Context-Aware and Temporal Focus Attentions
von: Liao, Haicheng, et al.
Veröffentlicht: (2024)
von: Liao, Haicheng, et al.
Veröffentlicht: (2024)
ClothPPO: A Proximal Policy Optimization Enhancing Framework for Robotic Cloth Manipulation with Observation-Aligned Action Spaces
von: Yang, Libing, et al.
Veröffentlicht: (2024)
von: Yang, Libing, et al.
Veröffentlicht: (2024)
TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers
von: Yu, Bin, et al.
Veröffentlicht: (2026)
von: Yu, Bin, et al.
Veröffentlicht: (2026)
DME-Driver: Integrating Human Decision Logic and 3D Scene Perception in Autonomous Driving
von: Han, Wencheng, et al.
Veröffentlicht: (2024)
von: Han, Wencheng, et al.
Veröffentlicht: (2024)
DISC: Dense Integrated Semantic Context for Large-Scale Open-Set Semantic Mapping
von: Igelbrink, Felix, et al.
Veröffentlicht: (2026)
von: Igelbrink, Felix, et al.
Veröffentlicht: (2026)
Perception Helps Planning: Facilitating Multi-Stage Lane-Level Integration via Double-Edge Structures
von: You, Guoliang, et al.
Veröffentlicht: (2024)
von: You, Guoliang, et al.
Veröffentlicht: (2024)
Perpetua: Multi-Hypothesis Persistence Modeling for Semi-Static Environments
von: Saavedra-Ruiz, Miguel, et al.
Veröffentlicht: (2025)
von: Saavedra-Ruiz, Miguel, et al.
Veröffentlicht: (2025)
Relative Pose for Nonrigid Multi-Perspective Cameras: The Static Case
von: Li, Min, et al.
Veröffentlicht: (2024)
von: Li, Min, et al.
Veröffentlicht: (2024)
Prof. Robot: Differentiable Robot Rendering Without Static and Self-Collisions
von: Ruan, Quanyuan, et al.
Veröffentlicht: (2025)
von: Ruan, Quanyuan, et al.
Veröffentlicht: (2025)
UniUncer: Unified Dynamic Static Uncertainty for End to End Driving
von: Gao, Yu, et al.
Veröffentlicht: (2026)
von: Gao, Yu, et al.
Veröffentlicht: (2026)
On-Road Object Importance Estimation: A New Dataset and A Model with Multi-Fold Top-Down Guidance
von: Nan, Zhixiong, et al.
Veröffentlicht: (2024)
von: Nan, Zhixiong, et al.
Veröffentlicht: (2024)
Diffusion Dynamics Models with Generative State Estimation for Cloth Manipulation
von: Tian, Tongxuan, et al.
Veröffentlicht: (2025)
von: Tian, Tongxuan, et al.
Veröffentlicht: (2025)
Efficient Long-Horizon Vision-Language-Action Models via Static-Dynamic Disentanglement
von: Qiu, Weikang, et al.
Veröffentlicht: (2026)
von: Qiu, Weikang, et al.
Veröffentlicht: (2026)
Structured Analysis and Comparison of Alphabets in Historical Handwritten Ciphers
von: Méndez, Martín, et al.
Veröffentlicht: (2024)
von: Méndez, Martín, et al.
Veröffentlicht: (2024)
Exploring the Use of VLMs for Navigation Assistance for People with Blindness and Low Vision
von: Li, Yu, et al.
Veröffentlicht: (2026)
von: Li, Yu, et al.
Veröffentlicht: (2026)
Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
von: Han, Jianhua, et al.
Veröffentlicht: (2025)
von: Han, Jianhua, et al.
Veröffentlicht: (2025)
ERASOR++: Height Coding Plus Egocentric Ratio Based Dynamic Object Removal for Static Point Cloud Mapping
von: Zhang, Jiabao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiabao, et al.
Veröffentlicht: (2024)
Learning Underwater Active Perception in Simulation
von: Cardaillac, Alexandre, et al.
Veröffentlicht: (2025)
von: Cardaillac, Alexandre, et al.
Veröffentlicht: (2025)
Adaptive Illumination Control for Robot Perception
von: Turkar, Yash, et al.
Veröffentlicht: (2026)
von: Turkar, Yash, et al.
Veröffentlicht: (2026)
Masked Depth Modeling for Spatial Perception
von: Tan, Bin, et al.
Veröffentlicht: (2026)
von: Tan, Bin, et al.
Veröffentlicht: (2026)
PERSEUS: Perception with Semantic Endoscopic Understanding and SLAM
von: Acar, Ayberk, et al.
Veröffentlicht: (2025)
von: Acar, Ayberk, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BiFold: Bimanual Cloth Folding with Language Guidance
von: Barbany, Oriol, et al.
Veröffentlicht: (2025) -
Benchmarking the Sim-to-Real Gap in Cloth Manipulation
von: Blanco-Mulero, David, et al.
Veröffentlicht: (2023) -
Dataset for BiFold: Bimanual Cloth Folding with Language Guidance
von: Barbany, Oriol, et al.
Veröffentlicht: (2025) -
CloSE: A Geometric Shape-Agnostic Cloth State Representation
von: Kamat, Jay, et al.
Veröffentlicht: (2025) -
Learning a General Model: Folding Clothing with Topological Dynamics
von: Liu, Yiming, et al.
Veröffentlicht: (2025)