PhotoBot: Reference-Guided Interactive Photography via Natural Language
Fuente:
arXiv
Saved in:
| Main Authors: | Limoyo, Oliver, Li, Jimmy, Rivkin, Dmitriy, Kelly, Jonathan, Dudek, Gregory |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Working Backwards: Learning to Place by Picking
by: Limoyo, Oliver, et al.
Published: (2023)
by: Limoyo, Oliver, et al.
Published: (2023)
Stable Multi-Drone GNSS Tracking System for Marine Robots
by: Wen, Shuo, et al.
Published: (2025)
by: Wen, Shuo, et al.
Published: (2025)
iFlyBot-VLA Technical Report
by: Zhang, Yuan, et al.
Published: (2025)
by: Zhang, Yuan, et al.
Published: (2025)
TidyBot++: An Open-Source Holonomic Mobile Manipulator for Robot Learning
by: Wu, Jimmy, et al.
Published: (2024)
by: Wu, Jimmy, et al.
Published: (2024)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
MANSION: Multi-floor lANguage-to-3D Scene generatIOn for loNg-horizon tasks
by: Che, Lirong, et al.
Published: (2026)
by: Che, Lirong, et al.
Published: (2026)
REALM: An RGB and Event Aligned Latent Manifold for Cross-Modal Perception
by: Polizzi, Vincenzo, et al.
Published: (2026)
by: Polizzi, Vincenzo, et al.
Published: (2026)
FlowBot3D: Learning 3D Articulation Flow to Manipulate Articulated Objects
by: Eisner, Ben, et al.
Published: (2022)
by: Eisner, Ben, et al.
Published: (2022)
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
Vega: Learning to Drive with Natural Language Instructions
by: Zuo, Sicheng, et al.
Published: (2026)
by: Zuo, Sicheng, et al.
Published: (2026)
PhotoAgent: A Robotic Photographer with Spatial and Aesthetic Understanding
by: Che, Lirong, et al.
Published: (2026)
by: Che, Lirong, et al.
Published: (2026)
GSWorld: Closed-Loop Photo-Realistic Simulation Suite for Robotic Manipulation
by: Jiang, Guangqi, et al.
Published: (2025)
by: Jiang, Guangqi, et al.
Published: (2025)
UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI
by: Zhong, Fangwei, et al.
Published: (2024)
by: Zhong, Fangwei, et al.
Published: (2024)
Object-Centric World Model for Language-Guided Manipulation
by: Jeong, Youngjoon, et al.
Published: (2025)
by: Jeong, Youngjoon, et al.
Published: (2025)
Multimodal and Force-Matched Imitation Learning with a See-Through Visuotactile Sensor
by: Ablett, Trevor, et al.
Published: (2023)
by: Ablett, Trevor, et al.
Published: (2023)
Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models
by: Wu, Qi, et al.
Published: (2024)
by: Wu, Qi, et al.
Published: (2024)
Knolling Bot: Teaching Robots the Human Notion of Tidiness
by: Hu, Yuhang, et al.
Published: (2023)
by: Hu, Yuhang, et al.
Published: (2023)
Toward Reliable AR-Guided Surgical Navigation: Interactive Deformation Modeling with Data-Driven Biomechanics and Prompts
by: Han, Zheng, et al.
Published: (2025)
by: Han, Zheng, et al.
Published: (2025)
DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving
by: Jiang, Anqing, et al.
Published: (2025)
by: Jiang, Anqing, et al.
Published: (2025)
A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring
by: Wang, Wenze, et al.
Published: (2026)
by: Wang, Wenze, et al.
Published: (2026)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
PlaceIt3D: Language-Guided Object Placement in Real 3D Scenes
by: Abdelreheem, Ahmed, et al.
Published: (2025)
by: Abdelreheem, Ahmed, et al.
Published: (2025)
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation
by: Chen, Jiaqi, et al.
Published: (2024)
by: Chen, Jiaqi, et al.
Published: (2024)
Dream to Recall: Imagination-Guided Experience Retrieval for Memory-Persistent Vision-and-Language Navigation
by: Xu, Yunzhe, et al.
Published: (2025)
by: Xu, Yunzhe, et al.
Published: (2025)
Enhancing Vision-Language Models with Scene Graphs for Traffic Accident Understanding
by: Lohner, Aaron, et al.
Published: (2024)
by: Lohner, Aaron, et al.
Published: (2024)
Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions
by: Li, Heng, et al.
Published: (2024)
by: Li, Heng, et al.
Published: (2024)
FunGraph: Functionality Aware 3D Scene Graphs for Language-Prompted Scene Interaction
by: Rotondi, Dennis, et al.
Published: (2025)
by: Rotondi, Dennis, et al.
Published: (2025)
VITAL: Interactive Few-Shot Imitation Learning via Visual Human-in-the-Loop Corrections
by: Kasaei, Hamidreza, et al.
Published: (2024)
by: Kasaei, Hamidreza, et al.
Published: (2024)
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
by: Chen, Xinyi, et al.
Published: (2025)
by: Chen, Xinyi, et al.
Published: (2025)
SynHLMA:Synthesizing Hand Language Manipulation for Articulated Object with Discrete Human Object Interaction Representation
by: zhi, Wang, et al.
Published: (2025)
by: zhi, Wang, et al.
Published: (2025)
SSL-Interactions: Pretext Tasks for Interactive Trajectory Prediction
by: Bhattacharyya, Prarthana, et al.
Published: (2024)
by: Bhattacharyya, Prarthana, et al.
Published: (2024)
PhotoFlow: Agentic 3D Virtual Photography Missions
by: Guo, Jiarui, et al.
Published: (2026)
by: Guo, Jiarui, et al.
Published: (2026)
Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition
by: Wen, Yuhang, et al.
Published: (2023)
by: Wen, Yuhang, et al.
Published: (2023)
Going Places: Place Recognition in Artificial and Natural Systems
by: Milford, Michael, et al.
Published: (2025)
by: Milford, Michael, et al.
Published: (2025)
J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception
by: Atuhurra, Jesse, et al.
Published: (2025)
by: Atuhurra, Jesse, et al.
Published: (2025)
IRL-VLA: Training an Vision-Language-Action Policy via Reward World Model
by: Jiang, Anqing, et al.
Published: (2025)
by: Jiang, Anqing, et al.
Published: (2025)
Autonomous Vision-Guided Resection of Central Airway Obstruction
by: Smith, M. E., et al.
Published: (2025)
by: Smith, M. E., et al.
Published: (2025)
ActiveVLN: Towards Active Exploration via Multi-Turn RL in Vision-and-Language Navigation
by: Zhang, Zekai, et al.
Published: (2025)
by: Zhang, Zekai, et al.
Published: (2025)
Hand-Object Interaction Pretraining from Videos
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
PhotoReg: Photometrically Registering 3D Gaussian Splatting Models
by: Yuan, Ziwen, et al.
Published: (2024)
by: Yuan, Ziwen, et al.
Published: (2024)
Similar Items
-
Working Backwards: Learning to Place by Picking
by: Limoyo, Oliver, et al.
Published: (2023) -
Stable Multi-Drone GNSS Tracking System for Marine Robots
by: Wen, Shuo, et al.
Published: (2025) -
iFlyBot-VLA Technical Report
by: Zhang, Yuan, et al.
Published: (2025) -
TidyBot++: An Open-Source Holonomic Mobile Manipulator for Robot Learning
by: Wu, Jimmy, et al.
Published: (2024) -
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)