Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Galliena, Tommaso, Rosa, Stefano, Apicella, Tommaso, Morerio, Pietro, Del Bue, Alessio, Natale, Lorenzo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions
by: Galliena, Tommaso, et al.
Published: (2025)
by: Galliena, Tommaso, et al.
Published: (2025)
Look Around and Learn: Self-Training Object Detection by Exploration
by: Scarpellini, Gianluca, et al.
Published: (2023)
by: Scarpellini, Gianluca, et al.
Published: (2023)
Learning to Evaluate Autonomous Behaviour in Human-Robot Interaction
by: Tiezzi, Matteo, et al.
Published: (2025)
by: Tiezzi, Matteo, et al.
Published: (2025)
Anticipating Next Active Objects for Egocentric Videos
by: Thakur, Sanket, et al.
Published: (2023)
by: Thakur, Sanket, et al.
Published: (2023)
CloseUpAvatar: High-Fidelity Animatable Full-Body Avatars with Mixture of Multi-Scale Textures
by: Svitov, David, et al.
Published: (2025)
by: Svitov, David, et al.
Published: (2025)
BillBoard Splatting (BBSplat): Learnable Textured Primitives for Novel View Synthesis
by: Svitov, David, et al.
Published: (2024)
by: Svitov, David, et al.
Published: (2024)
HAHA: Highly Articulated Gaussian Human Avatars with Textured Mesh Prior
by: Svitov, David, et al.
Published: (2024)
by: Svitov, David, et al.
Published: (2024)
DiffAssemble: A Unified Graph-Diffusion Model for 2D and 3D Reassembly
by: Scarpellini, Gianluca, et al.
Published: (2024)
by: Scarpellini, Gianluca, et al.
Published: (2024)
ReassembleNet: Learnable Keypoints and Diffusion for 2D Fresco Reconstruction
by: Islam, Adeela, et al.
Published: (2025)
by: Islam, Adeela, et al.
Published: (2025)
Segmenting Object Affordances: Reproducibility and Sensitivity to Scale
by: Apicella, Tommaso, et al.
Published: (2024)
by: Apicella, Tommaso, et al.
Published: (2024)
Pre-trained Multiple Latent Variable Generative Models are good defenders against Adversarial Attacks
by: Serez, Dario, et al.
Published: (2024)
by: Serez, Dario, et al.
Published: (2024)
A Mutual Information Perspective on Multiple Latent Variable Generative Models for Positive View Generation
by: Serez, Dario, et al.
Published: (2025)
by: Serez, Dario, et al.
Published: (2025)
Visual Affordance Prediction: Survey and Reproducibility
by: Apicella, Tommaso, et al.
Published: (2025)
by: Apicella, Tommaso, et al.
Published: (2025)
E-M3RF: An Equivariant Multimodal 3D Re-assembly Framework
by: Islam, Adeela, et al.
Published: (2025)
by: Islam, Adeela, et al.
Published: (2025)
Maps from Motion (MfM): Generating 2D Semantic Maps from Sparse Multi-view Images
by: Toso, Matteo, et al.
Published: (2024)
by: Toso, Matteo, et al.
Published: (2024)
Uncertainty-guided Open-Set Source-Free Unsupervised Domain Adaptation with Target-private Class Segregation
by: Litrico, Mattia, et al.
Published: (2024)
by: Litrico, Mattia, et al.
Published: (2024)
Affordance segmentation of hand-occluded containers from exocentric images
by: Apicella, Tommaso, et al.
Published: (2023)
by: Apicella, Tommaso, et al.
Published: (2023)
Model Debiasing by Learnable Data Augmentation
by: Morerio, Pietro, et al.
Published: (2024)
by: Morerio, Pietro, et al.
Published: (2024)
PRAGO: Differentiable Multi-View Pose Optimization From Objectness Detections
by: Taiana, Matteo, et al.
Published: (2024)
by: Taiana, Matteo, et al.
Published: (2024)
Gaussian Heritage: 3D Digitization of Cultural Heritage with Integrated Object Segmentation
by: Dahaghin, Mahtab, et al.
Published: (2024)
by: Dahaghin, Mahtab, et al.
Published: (2024)
SelfGeo: Self-supervised and Geodesic-consistent Estimation of Keypoints on Deformable Shapes
by: Zohaib, Mohammad, et al.
Published: (2024)
by: Zohaib, Mohammad, et al.
Published: (2024)
Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment
by: Yu, Fanqi, et al.
Published: (2026)
by: Yu, Fanqi, et al.
Published: (2026)
The impact of Compositionality in Zero-shot Multi-label action recognition for Object-based tasks
by: Calabrese, Carmela, et al.
Published: (2024)
by: Calabrese, Carmela, et al.
Published: (2024)
Training-Free Semantic Multi-Object Tracking with Vision-Language Models
by: Bonat, Laurence, et al.
Published: (2026)
by: Bonat, Laurence, et al.
Published: (2026)
A Baseline Study and Benchmark for Few-Shot Open-Set Action Recognition with Feature Residual Discrimination
by: Berti, Stefano, et al.
Published: (2026)
by: Berti, Stefano, et al.
Published: (2026)
Container Localisation and Mass Estimation with an RGB-D Camera
by: Apicella, Tommaso, et al.
Published: (2022)
by: Apicella, Tommaso, et al.
Published: (2022)
SplatFill: 3D Scene Inpainting via Depth-Guided Gaussian Splatting
by: Dahaghin, Mahtab, et al.
Published: (2025)
by: Dahaghin, Mahtab, et al.
Published: (2025)
Effectively Enhancing Vision Language Large Models by Prompt Augmentation and Caption Utilization
by: Zhao, Minyi, et al.
Published: (2024)
by: Zhao, Minyi, et al.
Published: (2024)
MeaCap: Memory-Augmented Zero-shot Image Captioning
by: Zeng, Zequn, et al.
Published: (2024)
by: Zeng, Zequn, et al.
Published: (2024)
ViTOC: Vision Transformer and Object-aware Captioner
by: Huang, Feiyang
Published: (2024)
by: Huang, Feiyang
Published: (2024)
6DGS: 6D Pose Estimation from a Single Image and a 3D Gaussian Splatting Model
by: Bortolon, Matteo, et al.
Published: (2024)
by: Bortolon, Matteo, et al.
Published: (2024)
MOPA: Modular Object Navigation with PointGoal Agents
by: Raychaudhuri, Sonia, et al.
Published: (2023)
by: Raychaudhuri, Sonia, et al.
Published: (2023)
6-DoF Object Tracking with Event-based Optical Flow and Frames
by: Li, Zhichao, et al.
Published: (2025)
by: Li, Zhichao, et al.
Published: (2025)
Event-based Motion & Appearance Fusion for 6D Object Pose Tracking
by: Li, Zhichao, et al.
Published: (2026)
by: Li, Zhichao, et al.
Published: (2026)
Sim2Real Bilevel Adaptation for Object Surface Classification using Vision-Based Tactile Sensors
by: Caddeo, Gabriele M., et al.
Published: (2023)
by: Caddeo, Gabriele M., et al.
Published: (2023)
Diffusion Based Augmentation for Captioning and Retrieval in Cultural Heritage
by: Cioni, Dario, et al.
Published: (2023)
by: Cioni, Dario, et al.
Published: (2023)
Gaussian-Augmented Physics Simulation and System Identification with Complex Colliders
by: Vasile, Federico, et al.
Published: (2025)
by: Vasile, Federico, et al.
Published: (2025)
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
by: Li, Qiming, et al.
Published: (2025)
by: Li, Qiming, et al.
Published: (2025)
IFFNeRF: Initialisation Free and Fast 6DoF pose estimation from a single image and a NeRF model
by: Bortolon, Matteo, et al.
Published: (2024)
by: Bortolon, Matteo, et al.
Published: (2024)
Contrastive Gaussian Clustering: Weakly Supervised 3D Scene Segmentation
by: Silva, Myrna C., et al.
Published: (2024)
by: Silva, Myrna C., et al.
Published: (2024)
Similar Items
-
Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions
by: Galliena, Tommaso, et al.
Published: (2025) -
Look Around and Learn: Self-Training Object Detection by Exploration
by: Scarpellini, Gianluca, et al.
Published: (2023) -
Learning to Evaluate Autonomous Behaviour in Human-Robot Interaction
by: Tiezzi, Matteo, et al.
Published: (2025) -
Anticipating Next Active Objects for Egocentric Videos
by: Thakur, Sanket, et al.
Published: (2023) -
CloseUpAvatar: High-Fidelity Animatable Full-Body Avatars with Mixture of Multi-Scale Textures
by: Svitov, David, et al.
Published: (2025)