Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions
Fuente:
arXiv
Saved in:
| Main Authors: | Galliena, Tommaso, Apicella, Tommaso, Rosa, Stefano, Morerio, Pietro, Del Bue, Alessio, Natale, Lorenzo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning
by: Galliena, Tommaso, et al.
Published: (2026)
by: Galliena, Tommaso, et al.
Published: (2026)
Learning to Evaluate Autonomous Behaviour in Human-Robot Interaction
by: Tiezzi, Matteo, et al.
Published: (2025)
by: Tiezzi, Matteo, et al.
Published: (2025)
Look Around and Learn: Self-Training Object Detection by Exploration
by: Scarpellini, Gianluca, et al.
Published: (2023)
by: Scarpellini, Gianluca, et al.
Published: (2023)
Visual Affordance Prediction: Survey and Reproducibility
by: Apicella, Tommaso, et al.
Published: (2025)
by: Apicella, Tommaso, et al.
Published: (2025)
Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment
by: Yu, Fanqi, et al.
Published: (2026)
by: Yu, Fanqi, et al.
Published: (2026)
CloseUpAvatar: High-Fidelity Animatable Full-Body Avatars with Mixture of Multi-Scale Textures
by: Svitov, David, et al.
Published: (2025)
by: Svitov, David, et al.
Published: (2025)
BillBoard Splatting (BBSplat): Learnable Textured Primitives for Novel View Synthesis
by: Svitov, David, et al.
Published: (2024)
by: Svitov, David, et al.
Published: (2024)
HAHA: Highly Articulated Gaussian Human Avatars with Textured Mesh Prior
by: Svitov, David, et al.
Published: (2024)
by: Svitov, David, et al.
Published: (2024)
DiffAssemble: A Unified Graph-Diffusion Model for 2D and 3D Reassembly
by: Scarpellini, Gianluca, et al.
Published: (2024)
by: Scarpellini, Gianluca, et al.
Published: (2024)
ReassembleNet: Learnable Keypoints and Diffusion for 2D Fresco Reconstruction
by: Islam, Adeela, et al.
Published: (2025)
by: Islam, Adeela, et al.
Published: (2025)
Pre-trained Multiple Latent Variable Generative Models are good defenders against Adversarial Attacks
by: Serez, Dario, et al.
Published: (2024)
by: Serez, Dario, et al.
Published: (2024)
Anticipating Next Active Objects for Egocentric Videos
by: Thakur, Sanket, et al.
Published: (2023)
by: Thakur, Sanket, et al.
Published: (2023)
A Mutual Information Perspective on Multiple Latent Variable Generative Models for Positive View Generation
by: Serez, Dario, et al.
Published: (2025)
by: Serez, Dario, et al.
Published: (2025)
PersONAL: Towards a Comprehensive Benchmark for Personalized Embodied Agents
by: Ziliotto, Filippo, et al.
Published: (2025)
by: Ziliotto, Filippo, et al.
Published: (2025)
Container Localisation and Mass Estimation with an RGB-D Camera
by: Apicella, Tommaso, et al.
Published: (2022)
by: Apicella, Tommaso, et al.
Published: (2022)
SelfGeo: Self-supervised and Geodesic-consistent Estimation of Keypoints on Deformable Shapes
by: Zohaib, Mohammad, et al.
Published: (2024)
by: Zohaib, Mohammad, et al.
Published: (2024)
E-M3RF: An Equivariant Multimodal 3D Re-assembly Framework
by: Islam, Adeela, et al.
Published: (2025)
by: Islam, Adeela, et al.
Published: (2025)
IFFNeRF: Initialisation Free and Fast 6DoF pose estimation from a single image and a NeRF model
by: Bortolon, Matteo, et al.
Published: (2024)
by: Bortolon, Matteo, et al.
Published: (2024)
Embodied Agents for Efficient Exploration and Smart Scene Description
by: Bigazzi, Roberto, et al.
Published: (2023)
by: Bigazzi, Roberto, et al.
Published: (2023)
The impact of Compositionality in Zero-shot Multi-label action recognition for Object-based tasks
by: Calabrese, Carmela, et al.
Published: (2024)
by: Calabrese, Carmela, et al.
Published: (2024)
Segmenting Object Affordances: Reproducibility and Sensitivity to Scale
by: Apicella, Tommaso, et al.
Published: (2024)
by: Apicella, Tommaso, et al.
Published: (2024)
Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
by: Liu, Dayong, et al.
Published: (2025)
by: Liu, Dayong, et al.
Published: (2025)
Image Quality Assessment for Embodied AI
by: Li, Chunyi, et al.
Published: (2025)
by: Li, Chunyi, et al.
Published: (2025)
GRASPLAT: Enabling dexterous grasping through novel view synthesis
by: Bortolon, Matteo, et al.
Published: (2025)
by: Bortolon, Matteo, et al.
Published: (2025)
Learn Fast, Segment Well: Fast Object Segmentation Learning on the iCub Robot
by: Ceola, Federico, et al.
Published: (2022)
by: Ceola, Federico, et al.
Published: (2022)
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Maps from Motion (MfM): Generating 2D Semantic Maps from Sparse Multi-view Images
by: Toso, Matteo, et al.
Published: (2024)
by: Toso, Matteo, et al.
Published: (2024)
Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch
by: Caddeo, Gabriele Mario, et al.
Published: (2026)
by: Caddeo, Gabriele Mario, et al.
Published: (2026)
Uncertainty-guided Open-Set Source-Free Unsupervised Domain Adaptation with Target-private Class Segregation
by: Litrico, Mattia, et al.
Published: (2024)
by: Litrico, Mattia, et al.
Published: (2024)
Embodied Navigation with Auxiliary Task of Action Description Prediction
by: Kondoh, Haru, et al.
Published: (2025)
by: Kondoh, Haru, et al.
Published: (2025)
MOPA: Modular Object Navigation with PointGoal Agents
by: Raychaudhuri, Sonia, et al.
Published: (2023)
by: Raychaudhuri, Sonia, et al.
Published: (2023)
Ego to World: Collaborative Spatial Reasoning in Embodied Systems via Reinforcement Learning
by: Zhou, Heng, et al.
Published: (2026)
by: Zhou, Heng, et al.
Published: (2026)
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models
by: Zhang, Jiyao, et al.
Published: (2026)
by: Zhang, Jiyao, et al.
Published: (2026)
Following the Human Thread in Social Navigation
by: Scofano, Luca, et al.
Published: (2024)
by: Scofano, Luca, et al.
Published: (2024)
Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
XBG: End-to-end Imitation Learning for Autonomous Behaviour in Human-Robot Interaction and Collaboration
by: Cardenas-Perez, Carlos, et al.
Published: (2024)
by: Cardenas-Perez, Carlos, et al.
Published: (2024)
The Safety Challenge of World Models for Embodied AI Agents: A Review
by: Baraldi, Lorenzo, et al.
Published: (2025)
by: Baraldi, Lorenzo, et al.
Published: (2025)
Spot the Difference: A Novel Task for Embodied Agents in Changing Environments
by: Landi, Federico, et al.
Published: (2022)
by: Landi, Federico, et al.
Published: (2022)
SpatialPoint: Spatial-aware Point Prediction for Embodied Localization
by: Zhu, Qiming, et al.
Published: (2026)
by: Zhu, Qiming, et al.
Published: (2026)
Robo-Cortex: A Self-Evolving Embodied Agent via Dual-Grain Cognitive Memory and Autonomous Knowledge Induction
by: Chan, Nga Teng, et al.
Published: (2026)
by: Chan, Nga Teng, et al.
Published: (2026)
Similar Items
-
Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning
by: Galliena, Tommaso, et al.
Published: (2026) -
Learning to Evaluate Autonomous Behaviour in Human-Robot Interaction
by: Tiezzi, Matteo, et al.
Published: (2025) -
Look Around and Learn: Self-Training Object Detection by Exploration
by: Scarpellini, Gianluca, et al.
Published: (2023) -
Visual Affordance Prediction: Survey and Reproducibility
by: Apicella, Tommaso, et al.
Published: (2025) -
Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment
by: Yu, Fanqi, et al.
Published: (2026)