JRDB-Social: A Multifaceted Robotic Dataset for Understanding of Context and Dynamics of Human Interactions Within Social Groups
Fuente:
arXiv
Saved in:
| Main Authors: | Jahangard, Simindokht, Cai, Zhixi, Wen, Shiki, Rezatofighi, Hamid |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
by: Jahangard, Simindokht, et al.
Published: (2025)
by: Jahangard, Simindokht, et al.
Published: (2025)
HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning
by: Ke, Fucai, et al.
Published: (2024)
by: Ke, Fucai, et al.
Published: (2024)
A Multi-Modal Neuro-Symbolic Approach for Spatial Reasoning-Based Visual Grounding in Robotics
by: Jahangard, Simindokht, et al.
Published: (2025)
by: Jahangard, Simindokht, et al.
Published: (2025)
JRDB-Pose3D: A Multi-person 3D Human Pose and Shape Estimation Dataset for Robotics
by: Biswas, Sandika, et al.
Published: (2026)
by: Biswas, Sandika, et al.
Published: (2026)
JRDB-PanoTrack: An Open-world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments
by: Le, Duy-Tho, et al.
Published: (2024)
by: Le, Duy-Tho, et al.
Published: (2024)
NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning
by: Cai, Zhixi, et al.
Published: (2025)
by: Cai, Zhixi, et al.
Published: (2025)
Improving Visual Perception of a Social Robot for Controlled and In-the-wild Human-robot Interaction
by: Zhong, Wangjie, et al.
Published: (2024)
by: Zhong, Wangjie, et al.
Published: (2024)
Social-MAE: Social Masked Autoencoder for Multi-person Motion Representation Learning
by: Ehsanpour, Mahsa, et al.
Published: (2024)
by: Ehsanpour, Mahsa, et al.
Published: (2024)
Asset-Driven Sematic Reconstruction of Dynamic Scene with Multi-Human-Object Interactions
by: Biswas, Sandika, et al.
Published: (2025)
by: Biswas, Sandika, et al.
Published: (2025)
Marginalized Generalized IoU (MGIoU): A Unified Objective Function for Optimizing Any Convex Parametric Shapes
by: Le, Duy-Tho, et al.
Published: (2025)
by: Le, Duy-Tho, et al.
Published: (2025)
TFS-NeRF: Template-Free NeRF for Semantic 3D Reconstruction of Dynamic Scene
by: Biswas, Sandika, et al.
Published: (2024)
by: Biswas, Sandika, et al.
Published: (2024)
DrVideo: Document Retrieval Based Long Video Understanding
by: Ma, Ziyu, et al.
Published: (2024)
by: Ma, Ziyu, et al.
Published: (2024)
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations
by: Ke, Fucai, et al.
Published: (2026)
by: Ke, Fucai, et al.
Published: (2026)
DifFUSER: Diffusion Model for Robust Multi-Sensor Fusion in 3D Object Detection and BEV Segmentation
by: Le, Duy-Tho, et al.
Published: (2024)
by: Le, Duy-Tho, et al.
Published: (2024)
MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual Reasoning
by: Cai, Zhixi, et al.
Published: (2026)
by: Cai, Zhixi, et al.
Published: (2026)
ASAP-Textured Gaussians: Enhancing Textured Gaussians with Adaptive Sampling and Anisotropic Parameterization
by: Wei, Meng, et al.
Published: (2025)
by: Wei, Meng, et al.
Published: (2025)
Normal-GS: 3D Gaussian Splatting with Normal-Involved Rendering
by: Wei, Meng, et al.
Published: (2024)
by: Wei, Meng, et al.
Published: (2024)
DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning
by: Ke, Fucai, et al.
Published: (2025)
by: Ke, Fucai, et al.
Published: (2025)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
by: Shaker, Abdelrahman, et al.
Published: (2025)
by: Shaker, Abdelrahman, et al.
Published: (2025)
dinov3.seg: Open-Vocabulary Semantic Segmentation with DINOv3
by: Dutta, Saikat, et al.
Published: (2026)
by: Dutta, Saikat, et al.
Published: (2026)
Within the Dynamic Context: Inertia-aware 3D Human Modeling with Pose Sequence
by: Chen, Yutong, et al.
Published: (2024)
by: Chen, Yutong, et al.
Published: (2024)
Robot Interaction Behavior Generation based on Social Motion Forecasting for Human-Robot Interaction
by: Mascaro, Esteve Valls, et al.
Published: (2024)
by: Mascaro, Esteve Valls, et al.
Published: (2024)
SocialGen: Modeling Multi-Human Social Interaction with Language Models
by: Yu, Heng, et al.
Published: (2025)
by: Yu, Heng, et al.
Published: (2025)
SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition
by: Xie, Kunyuan, et al.
Published: (2026)
by: Xie, Kunyuan, et al.
Published: (2026)
Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots
by: Bartoli, Ermanno, et al.
Published: (2025)
by: Bartoli, Ermanno, et al.
Published: (2025)
Towards Online Multi-Modal Social Interaction Understanding
by: Li, Xinpeng, et al.
Published: (2025)
by: Li, Xinpeng, et al.
Published: (2025)
How Well Can Vision Language Models See Image Details?
by: Gou, Chenhui, et al.
Published: (2024)
by: Gou, Chenhui, et al.
Published: (2024)
SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation
by: Munje, Michael J., et al.
Published: (2025)
by: Munje, Michael J., et al.
Published: (2025)
Omni-MMSI: Toward Identity-attributed Social Interaction Understanding
by: Li, Xinpeng, et al.
Published: (2026)
by: Li, Xinpeng, et al.
Published: (2026)
RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives
by: Parikh, Chirag, et al.
Published: (2025)
by: Parikh, Chirag, et al.
Published: (2025)
HUMOF: Human Motion Forecasting in Interactive Social Scenes
by: Sun, Caiyi, et al.
Published: (2025)
by: Sun, Caiyi, et al.
Published: (2025)
CalibAnyView: Beyond Single-View Camera Calibration in the Wild
by: Li, Boying, et al.
Published: (2026)
by: Li, Boying, et al.
Published: (2026)
AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset
by: Cai, Zhixi, et al.
Published: (2023)
by: Cai, Zhixi, et al.
Published: (2023)
Explain Before You Answer: A Survey on Compositional Visual Reasoning
by: Ke, Fucai, et al.
Published: (2025)
by: Ke, Fucai, et al.
Published: (2025)
Do Blind Spots Matter for Word-Referent Mapping? A Computational Study with Infant Egocentric Video
by: Shi, Zekai, et al.
Published: (2025)
by: Shi, Zekai, et al.
Published: (2025)
Heterogeneous Network Based Contrastive Learning Method for PolSAR Land Cover Classification
by: Cai, Jianfeng, et al.
Published: (2024)
by: Cai, Jianfeng, et al.
Published: (2024)
AerOSeg: Harnessing SAM for Open-Vocabulary Segmentation in Remote Sensing Images
by: Dutta, Saikat, et al.
Published: (2025)
by: Dutta, Saikat, et al.
Published: (2025)
Learning Human-Object Interaction as Groups
by: Hong, Jiajun, et al.
Published: (2025)
by: Hong, Jiajun, et al.
Published: (2025)
An Empirical Study on How Video-LLMs Answer Video Questions
by: Gou, Chenhui, et al.
Published: (2025)
by: Gou, Chenhui, et al.
Published: (2025)
Who Walks With You Matters: Perceiving Social Interactions with Groups for Pedestrian Trajectory Prediction
by: Zou, Ziqian, et al.
Published: (2024)
by: Zou, Ziqian, et al.
Published: (2024)
Similar Items
-
JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
by: Jahangard, Simindokht, et al.
Published: (2025) -
HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning
by: Ke, Fucai, et al.
Published: (2024) -
A Multi-Modal Neuro-Symbolic Approach for Spatial Reasoning-Based Visual Grounding in Robotics
by: Jahangard, Simindokht, et al.
Published: (2025) -
JRDB-Pose3D: A Multi-person 3D Human Pose and Shape Estimation Dataset for Robotics
by: Biswas, Sandika, et al.
Published: (2026) -
JRDB-PanoTrack: An Open-world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments
by: Le, Duy-Tho, et al.
Published: (2024)