Growing Perspectives: Modelling Embodied Perspective Taking and Inner Narrative Development Using Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Patania, Sabrina, Annese, Luca, Lambiase, Anna, Pellegrini, Anita, Foulsham, Tom, Ruggeri, Azzurra, Rossi, Silvia, Serino, Silvia, Ognibene, Dimitri |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PerspAct: Enhancing LLM Situated Collaboration Skills through Perspective Taking and Active Vision
por: Patania, Sabrina, et al.
Publicado: (2025)
por: Patania, Sabrina, et al.
Publicado: (2025)
Who Sees What? Structured Thought-Action Sequences for Epistemic Reasoning in LLMs
por: Annese, Luca, et al.
Publicado: (2025)
por: Annese, Luca, et al.
Publicado: (2025)
AI Pedagogy: Dialogic Social Learning for Artificial Agents
por: Patania, Sabrina, et al.
Publicado: (2025)
por: Patania, Sabrina, et al.
Publicado: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
por: Raoufi, Behnam, et al.
Publicado: (2025)
por: Raoufi, Behnam, et al.
Publicado: (2025)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
por: Lim, Shoon Kit, et al.
Publicado: (2025)
por: Lim, Shoon Kit, et al.
Publicado: (2025)
EmbodiedLGR: Integrating Lightweight Graph Representation and Retrieval for Semantic-Spatial Memory in Robotic Agents
por: Riva, Paolo, et al.
Publicado: (2026)
por: Riva, Paolo, et al.
Publicado: (2026)
Motion Perceiver: Real-Time Occupancy Forecasting for Embedded Systems
por: Ferenczi, Bryce, et al.
Publicado: (2023)
por: Ferenczi, Bryce, et al.
Publicado: (2023)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
por: Romero, Angel, et al.
Publicado: (2025)
por: Romero, Angel, et al.
Publicado: (2025)
EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique
por: Zhu, Chenglin, et al.
Publicado: (2025)
por: Zhu, Chenglin, et al.
Publicado: (2025)
Closed-Loop Neural Activation Control in Vision-Language-Action Models
por: Babu, Abhijith, et al.
Publicado: (2026)
por: Babu, Abhijith, et al.
Publicado: (2026)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026)
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
por: Li, Danyang, et al.
Publicado: (2025)
por: Li, Danyang, et al.
Publicado: (2025)
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
por: Gasparino, Mateus Valverde, et al.
Publicado: (2024)
por: Gasparino, Mateus Valverde, et al.
Publicado: (2024)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
Deep Probabilistic Traversability with Test-time Adaptation for Uncertainty-aware Planetary Rover Navigation
por: Endo, Masafumi, et al.
Publicado: (2024)
por: Endo, Masafumi, et al.
Publicado: (2024)
CoMoCAVs: Cohesive Decision-Guided Motion Planning for Connected and Autonomous Vehicles with Multi-Policy Reinforcement Learning
por: Hu, Pan
Publicado: (2025)
por: Hu, Pan
Publicado: (2025)
Cortex 2.0: Grounding World Models in Real-World Industrial Deployment
por: Aida, Adriana, et al.
Publicado: (2026)
por: Aida, Adriana, et al.
Publicado: (2026)
CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion
por: Römer, Ralf, et al.
Publicado: (2026)
por: Römer, Ralf, et al.
Publicado: (2026)
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
por: Syed, Shahram Najam, et al.
Publicado: (2025)
por: Syed, Shahram Najam, et al.
Publicado: (2025)
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
por: Zhang, Sinin, et al.
Publicado: (2026)
por: Zhang, Sinin, et al.
Publicado: (2026)
Task Singular Vectors: Reducing Task Interference in Model Merging
por: Gargiulo, Antonio Andrea, et al.
Publicado: (2024)
por: Gargiulo, Antonio Andrea, et al.
Publicado: (2024)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
por: Chen, Yuangong, et al.
Publicado: (2026)
por: Chen, Yuangong, et al.
Publicado: (2026)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
por: Li, Huibin, et al.
Publicado: (2025)
por: Li, Huibin, et al.
Publicado: (2025)
SUN Team's Contribution to ABAW 2024 Competition: Audio-visual Valence-Arousal Estimation and Expression Recognition
por: Dresvyanskiy, Denis, et al.
Publicado: (2024)
por: Dresvyanskiy, Denis, et al.
Publicado: (2024)
Learning the meanings of function words from grounded language using a visual question answering model
por: Portelance, Eva, et al.
Publicado: (2023)
por: Portelance, Eva, et al.
Publicado: (2023)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
por: Yang, Shan
Publicado: (2026)
por: Yang, Shan
Publicado: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
A Segmented Robot Grasping Perception Neural Network for Edge AI
por: Bröcheler, Casper, et al.
Publicado: (2025)
por: Bröcheler, Casper, et al.
Publicado: (2025)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
por: Gupta, Sunny, et al.
Publicado: (2024)
por: Gupta, Sunny, et al.
Publicado: (2024)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
por: Kashyap, Pankhi, et al.
Publicado: (2024)
por: Kashyap, Pankhi, et al.
Publicado: (2024)
Convolutional Model Trees
por: Armstrong, William Ward, et al.
Publicado: (2025)
por: Armstrong, William Ward, et al.
Publicado: (2025)
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
por: Tourani, Ali, et al.
Publicado: (2023)
por: Tourani, Ali, et al.
Publicado: (2023)
UAV-assisted Visual SLAM Generating Reconstructed 3D Scene Graphs in GPS-denied Environments
por: Radwan, Ahmed, et al.
Publicado: (2024)
por: Radwan, Ahmed, et al.
Publicado: (2024)
From Demonstrations to Safe Deployment: Path-Consistent Safety Filtering for Diffusion Policies
por: Römer, Ralf, et al.
Publicado: (2025)
por: Römer, Ralf, et al.
Publicado: (2025)
Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systems
por: Albiero, Daniel, et al.
Publicado: (2026)
por: Albiero, Daniel, et al.
Publicado: (2026)
Few-Shot Learning of a Graph-Based Neural Network Model Without Backpropagation
por: Lapin, Mykyta, et al.
Publicado: (2025)
por: Lapin, Mykyta, et al.
Publicado: (2025)
RoboPack: Learning Tactile-Informed Dynamics Models for Dense Packing
por: Ai, Bo, et al.
Publicado: (2024)
por: Ai, Bo, et al.
Publicado: (2024)
OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis
por: Chen, Junting, et al.
Publicado: (2025)
por: Chen, Junting, et al.
Publicado: (2025)
EvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experience
por: Xue, Taofeng, et al.
Publicado: (2026)
por: Xue, Taofeng, et al.
Publicado: (2026)
CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment
por: Kang, Li, et al.
Publicado: (2026)
por: Kang, Li, et al.
Publicado: (2026)
Ejemplares similares
-
PerspAct: Enhancing LLM Situated Collaboration Skills through Perspective Taking and Active Vision
por: Patania, Sabrina, et al.
Publicado: (2025) -
Who Sees What? Structured Thought-Action Sequences for Epistemic Reasoning in LLMs
por: Annese, Luca, et al.
Publicado: (2025) -
AI Pedagogy: Dialogic Social Learning for Artificial Agents
por: Patania, Sabrina, et al.
Publicado: (2025) -
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
por: Raoufi, Behnam, et al.
Publicado: (2025) -
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
por: Lim, Shoon Kit, et al.
Publicado: (2025)