J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Atuhurra, Jesse, Kamigaito, Hidetaka, Watanabe, Taro, Yoshino, Koichiro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
by: Atuhurra, Jesse, et al.
Published: (2025)
by: Atuhurra, Jesse, et al.
Published: (2025)
Constructing Multilingual Visual-Text Datasets Revealing Visual Multilingual Ability of Vision Language Models
by: Atuhurra, Jesse, et al.
Published: (2024)
by: Atuhurra, Jesse, et al.
Published: (2024)
NERsocial: Efficient Named Entity Recognition Dataset Construction for Human-Robot Interaction Utilizing RapidNER
by: Atuhurra, Jesse, et al.
Published: (2024)
by: Atuhurra, Jesse, et al.
Published: (2024)
Diagnosing Vision Language Models' Perception by Leveraging Human Methods for Color Vision Deficiencies
by: Hayashi, Kazuki, et al.
Published: (2025)
by: Hayashi, Kazuki, et al.
Published: (2025)
Revealing Trends in Datasets from the 2022 ACL and EMNLP Conferences
by: Atuhurra, Jesse, et al.
Published: (2024)
by: Atuhurra, Jesse, et al.
Published: (2024)
Towards Temporal Change Explanations from Bi-Temporal Satellite Images
by: Tsujimoto, Ryo, et al.
Published: (2024)
by: Tsujimoto, Ryo, et al.
Published: (2024)
Towards Artwork Explanation in Large-scale Vision Language Models
by: Hayashi, Kazuki, et al.
Published: (2024)
by: Hayashi, Kazuki, et al.
Published: (2024)
Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation
by: Li, Yuyang, et al.
Published: (2025)
by: Li, Yuyang, et al.
Published: (2025)
ClearDepth: Enhanced Stereo Perception of Transparent Objects for Robotic Manipulation
by: Bai, Kaixin, et al.
Published: (2024)
by: Bai, Kaixin, et al.
Published: (2024)
HOH: Markerless Multimodal Human-Object-Human Handover Dataset with Large Object Count
by: Wiederhold, Noah, et al.
Published: (2023)
by: Wiederhold, Noah, et al.
Published: (2023)
Towards Comprehensive Multimodal Perception: Introducing the Touch-Language-Vision Dataset
by: Cheng, Ning, et al.
Published: (2024)
by: Cheng, Ning, et al.
Published: (2024)
Object-Centric Action-Enhanced Representations for Robot Visuo-Motor Policy Learning
by: Giannakakis, Nikos, et al.
Published: (2025)
by: Giannakakis, Nikos, et al.
Published: (2025)
Eye, Robot: Learning to Look to Act with a BC-RL Perception-Action Loop
by: Kerr, Justin, et al.
Published: (2025)
by: Kerr, Justin, et al.
Published: (2025)
TiROD: Tiny Robotics Dataset and Benchmark for Continual Object Detection
by: Pasti, Francesco, et al.
Published: (2024)
by: Pasti, Francesco, et al.
Published: (2024)
TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation
by: Patel, Manthan, et al.
Published: (2025)
by: Patel, Manthan, et al.
Published: (2025)
S3E: A Multi-Robot Multimodal Dataset for Collaborative SLAM
by: Feng, Dapeng, et al.
Published: (2022)
by: Feng, Dapeng, et al.
Published: (2022)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
by: Liu, Mengzhen, et al.
Published: (2026)
by: Liu, Mengzhen, et al.
Published: (2026)
PokeFlex: A Real-World Dataset of Volumetric Deformable Objects for Robotics
by: Obrist, Jan, et al.
Published: (2024)
by: Obrist, Jan, et al.
Published: (2024)
SeePerSea: Multi-modal Perception Dataset of In-water Objects for Autonomous Surface Vehicles
by: Jeong, Mingi, et al.
Published: (2024)
by: Jeong, Mingi, et al.
Published: (2024)
MMCIG: Multimodal Cover Image Generation for Text-only Documents and Its Dataset Construction via Pseudo-labeling
by: Kim, Hyeyeon, et al.
Published: (2025)
by: Kim, Hyeyeon, et al.
Published: (2025)
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
by: Wang, Kuanning, et al.
Published: (2026)
by: Wang, Kuanning, et al.
Published: (2026)
VizFlyt: Perception-centric Pedagogical Framework For Autonomous Aerial Robots
by: Srivastava, Kushagra, et al.
Published: (2025)
by: Srivastava, Kushagra, et al.
Published: (2025)
OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
by: Guo, Heyu, et al.
Published: (2025)
by: Guo, Heyu, et al.
Published: (2025)
Handling Object Symmetries in CNN-based Pose Estimation
by: Richter-Klug, Jesse, et al.
Published: (2020)
by: Richter-Klug, Jesse, et al.
Published: (2020)
OceanSim: A GPU-Accelerated Underwater Robot Perception Simulation Framework
by: Song, Jingyu, et al.
Published: (2025)
by: Song, Jingyu, et al.
Published: (2025)
DIO: Dataset of 3D Mesh Models of Indoor Objects for Robotics and Computer Vision Applications
by: Nimal, Nillan, et al.
Published: (2024)
by: Nimal, Nillan, et al.
Published: (2024)
WD-DETR: Wavelet Denoising-Enhanced Real-Time Object Detection Transformer for Robot Perception with Event Cameras
by: Cui, Yangjie, et al.
Published: (2025)
by: Cui, Yangjie, et al.
Published: (2025)
Decision-Driven Semantic Object Exploration for Legged Robots via Confidence-Calibrated Perception and Topological Subgoal Selection
by: Zhao, Guoyang, et al.
Published: (2025)
by: Zhao, Guoyang, et al.
Published: (2025)
MBE-ARI: A Multimodal Dataset Mapping Bi-directional Engagement in Animal-Robot Interaction
by: Noronha, Ian, et al.
Published: (2025)
by: Noronha, Ian, et al.
Published: (2025)
Introducing Syllable Tokenization for Low-resource Languages: A Case Study with Swahili
by: Atuhurra, Jesse, et al.
Published: (2024)
by: Atuhurra, Jesse, et al.
Published: (2024)
Domain Adaptation in Intent Classification Systems: A Review
by: Atuhurra, Jesse, et al.
Published: (2024)
by: Atuhurra, Jesse, et al.
Published: (2024)
MATT-GS: Masked Attention-based 3DGS for Robot Perception and Object Detection
by: Lee, Jee Won, et al.
Published: (2025)
by: Lee, Jee Won, et al.
Published: (2025)
CoPeD-Advancing Multi-Robot Collaborative Perception: A Comprehensive Dataset in Real-World Environments
by: Zhou, Yang, et al.
Published: (2024)
by: Zhou, Yang, et al.
Published: (2024)
A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions
by: Inadumi, Shun, et al.
Published: (2024)
by: Inadumi, Shun, et al.
Published: (2024)
V2X-ReaLO: An Open Online Framework and Dataset for Cooperative Perception in Reality
by: Xiang, Hao, et al.
Published: (2025)
by: Xiang, Hao, et al.
Published: (2025)
From Detection to Action Recognition: An Edge-Based Pipeline for Robot Human Perception
by: Toupas, Petros, et al.
Published: (2023)
by: Toupas, Petros, et al.
Published: (2023)
Rodrigues Network for Learning Robot Actions
by: Zhang, Jialiang, et al.
Published: (2025)
by: Zhang, Jialiang, et al.
Published: (2025)
Adaptive Illumination Control for Robot Perception
by: Turkar, Yash, et al.
Published: (2026)
by: Turkar, Yash, et al.
Published: (2026)
Survey on Datasets for Perception in Unstructured Outdoor Environments
by: Mortimer, Peter, et al.
Published: (2024)
by: Mortimer, Peter, et al.
Published: (2024)
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
by: Nguyen, Nghia, et al.
Published: (2024)
by: Nguyen, Nghia, et al.
Published: (2024)
Similar Items
-
VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
by: Atuhurra, Jesse, et al.
Published: (2025) -
Constructing Multilingual Visual-Text Datasets Revealing Visual Multilingual Ability of Vision Language Models
by: Atuhurra, Jesse, et al.
Published: (2024) -
NERsocial: Efficient Named Entity Recognition Dataset Construction for Human-Robot Interaction Utilizing RapidNER
by: Atuhurra, Jesse, et al.
Published: (2024) -
Diagnosing Vision Language Models' Perception by Leveraging Human Methods for Color Vision Deficiencies
by: Hayashi, Kazuki, et al.
Published: (2025) -
Revealing Trends in Datasets from the 2022 ACL and EMNLP Conferences
by: Atuhurra, Jesse, et al.
Published: (2024)