ViTA-Seg: Vision Transformer for Amodal Segmentation in Robotics
Fuente:
arXiv
Saved in:
| Main Authors: | Caramia, Donato, Pokorny, Florian T., Triggiani, Giuseppe, Ruffino, Denis, Naso, David, Massenio, Paolo Roberto |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
F-ViTA: Foundation Model Guided Visible to Thermal Translation
by: Paranjape, Jay N., et al.
Published: (2025)
by: Paranjape, Jay N., et al.
Published: (2025)
How Physics and Background Attributes Impact Video Transformers in Robotic Manipulation: A Case Study on Planar Pushing
by: Jin, Shutong, et al.
Published: (2023)
by: Jin, Shutong, et al.
Published: (2023)
RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
by: Mei, Haiyang, et al.
Published: (2025)
by: Mei, Haiyang, et al.
Published: (2025)
ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition
by: Park, Minjeong, et al.
Published: (2025)
by: Park, Minjeong, et al.
Published: (2025)
VAT: Vision Action Transformer by Unlocking Full Representation of ViT
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Amodal Optical Flow
by: Luz, Maximilian, et al.
Published: (2023)
by: Luz, Maximilian, et al.
Published: (2023)
ViT-VS: On the Applicability of Pretrained Vision Transformer Features for Generalizable Visual Servoing
by: Scherl, Alessandro, et al.
Published: (2025)
by: Scherl, Alessandro, et al.
Published: (2025)
LiteViLNet: Lightweight Vision-LiDAR Fusion Network for Efficient Road Segmentation
by: Peng, Daojie, et al.
Published: (2026)
by: Peng, Daojie, et al.
Published: (2026)
Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation
by: Ding, Hongyu, et al.
Published: (2026)
by: Ding, Hongyu, et al.
Published: (2026)
PACA: Perspective-Aware Cross-Attention Representation for Zero-Shot Scene Rearrangement
by: Jin, Shutong, et al.
Published: (2024)
by: Jin, Shutong, et al.
Published: (2024)
TransForSeg: A Multitask Stereo ViT for Joint Stereo Segmentation and 3D Force Estimation in Catheterization
by: Fekri, Pedram, et al.
Published: (2025)
by: Fekri, Pedram, et al.
Published: (2025)
Robot Manipulation in Salient Vision through Referring Image Segmentation and Geometric Constraints
by: Jiang, Chen, et al.
Published: (2024)
by: Jiang, Chen, et al.
Published: (2024)
AISFormer: Amodal Instance Segmentation with Transformer
by: Tran, Minh, et al.
Published: (2022)
by: Tran, Minh, et al.
Published: (2022)
ReViP: Mitigating False Completion in Vision-Language-Action Models with Vision-Proprioception Rebalance
by: Li, Zhuohao, et al.
Published: (2026)
by: Li, Zhuohao, et al.
Published: (2026)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
by: Chen, Zerui, et al.
Published: (2024)
by: Chen, Zerui, et al.
Published: (2024)
Low Resolution Next Best View for Robot Packing
by: Preziosa, Giuseppe Fabio, et al.
Published: (2025)
by: Preziosa, Giuseppe Fabio, et al.
Published: (2025)
ScenarioControl: Vision-Language Controllable Vectorized Latent Scenario Generation
by: Gao, Lili, et al.
Published: (2026)
by: Gao, Lili, et al.
Published: (2026)
ZISVFM: Zero-Shot Object Instance Segmentation in Indoor Robotic Environments with Vision Foundation Models
by: Zhang, Ying, et al.
Published: (2025)
by: Zhang, Ying, et al.
Published: (2025)
LuSeg: Efficient Negative and Positive Obstacles Segmentation via Contrast-Driven Multi-Modal Feature Fusion on the Lunar
by: Jiao, Shuaifeng, et al.
Published: (2025)
by: Jiao, Shuaifeng, et al.
Published: (2025)
SegSTRONG-C: Segmenting Surgical Tools Robustly On Non-adversarial Generated Corruptions -- An EndoVis'24 Challenge
by: Ding, Hao, et al.
Published: (2024)
by: Ding, Hao, et al.
Published: (2024)
Clutt3R-Seg: Sparse-view 3D Instance Segmentation for Language-grounded Grasping in Cluttered Scenes
by: Noh, Jeongho, et al.
Published: (2026)
by: Noh, Jeongho, et al.
Published: (2026)
ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance Segmentation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
SegVec3D: A Method for Vector Embedding of 3D Objects Oriented Towards Robot manipulation
by: Kang, Zhihan, et al.
Published: (2025)
by: Kang, Zhihan, et al.
Published: (2025)
MultiGraspNet: A Multitask 3D Vision Model for Multi-gripper Robotic Grasping
by: Ortuno-Chanelo, Stephany, et al.
Published: (2026)
by: Ortuno-Chanelo, Stephany, et al.
Published: (2026)
SEMNAV: Enhancing Visual Semantic Navigation in Robotics through Semantic Segmentation
by: Flor-Rodríguez, Rafael, et al.
Published: (2025)
by: Flor-Rodríguez, Rafael, et al.
Published: (2025)
ViVa-SAFELAND: a New Freeware for Safe Validation of Vision-based Navigation in Aerial Vehicles
by: Soriano-García, Miguel S., et al.
Published: (2025)
by: Soriano-García, Miguel S., et al.
Published: (2025)
vAccSOL: Efficient and Transparent AI Vision Offloading for Mobile Robots
by: Zahir, Adam, et al.
Published: (2026)
by: Zahir, Adam, et al.
Published: (2026)
Garbage Segmentation and Attribute Analysis by Robotic Dogs
by: Xu, Nuo, et al.
Published: (2024)
by: Xu, Nuo, et al.
Published: (2024)
CaveSeg: Deep Semantic Segmentation and Scene Parsing for Autonomous Underwater Cave Exploration
by: Abdullah, A., et al.
Published: (2023)
by: Abdullah, A., et al.
Published: (2023)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
by: Han, Xiaofeng, et al.
Published: (2025)
by: Han, Xiaofeng, et al.
Published: (2025)
Learn Fast, Segment Well: Fast Object Segmentation Learning on the iCub Robot
by: Ceola, Federico, et al.
Published: (2022)
by: Ceola, Federico, et al.
Published: (2022)
RAPiD-Seg: Range-Aware Pointwise Distance Distribution Networks for 3D LiDAR Segmentation
by: Li, Li, et al.
Published: (2024)
by: Li, Li, et al.
Published: (2024)
Robot Instance Segmentation with Few Annotations for Grasping
by: Kimhi, Moshe, et al.
Published: (2024)
by: Kimhi, Moshe, et al.
Published: (2024)
Towards an Accurate and Effective Robot Vision (The Problem of Topological Localization for Mobile Robots)
by: Boros, Emanuela
Published: (2025)
by: Boros, Emanuela
Published: (2025)
Temporally Consistent Unsupervised Segmentation for Mobile Robot Perception
by: Ellis, Christian, et al.
Published: (2025)
by: Ellis, Christian, et al.
Published: (2025)
SegXAL: Explainable Active Learning for Semantic Segmentation in Driving Scene Scenarios
by: Mandalika, Sriram, et al.
Published: (2024)
by: Mandalika, Sriram, et al.
Published: (2024)
More than Segmentation: Benchmarking SAM 3 for Segmentation, 3D Perception, and Reconstruction in Robotic Surgery
by: Dong, Wenzhen, et al.
Published: (2025)
by: Dong, Wenzhen, et al.
Published: (2025)
LiPS: Lightweight Panoptic Segmentation for Resource-Constrained Robotics
by: Galagain, Calvin, et al.
Published: (2026)
by: Galagain, Calvin, et al.
Published: (2026)
Nano-U: Efficient Terrain Segmentation for Tiny Robot Navigation
by: Pizzolato, Federico, et al.
Published: (2026)
by: Pizzolato, Federico, et al.
Published: (2026)
Learning Priors of Human Motion With Vision Transformers
by: Falqueto, Placido, et al.
Published: (2025)
by: Falqueto, Placido, et al.
Published: (2025)
Similar Items
-
F-ViTA: Foundation Model Guided Visible to Thermal Translation
by: Paranjape, Jay N., et al.
Published: (2025) -
How Physics and Background Attributes Impact Video Transformers in Robotic Manipulation: A Case Study on Planar Pushing
by: Jin, Shutong, et al.
Published: (2023) -
RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
by: Mei, Haiyang, et al.
Published: (2025) -
ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition
by: Park, Minjeong, et al.
Published: (2025) -
VAT: Vision Action Transformer by Unlocking Full Representation of ViT
by: Li, Wenhao, et al.
Published: (2025)