Clink! Chop! Thud! -- Learning Object Sounds from Real-World Interactions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Mengyu, Chen, Yiming, Pei, Haozheng, Agarwal, Siddhant, Vasudevan, Arun Balajee, Hays, James |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
von: Bansal, Siddhant, et al.
Veröffentlicht: (2024)
von: Bansal, Siddhant, et al.
Veröffentlicht: (2024)
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
ICTPolarReal: A Polarized Reflection and Material Dataset of Real World Objects
von: Yang, Jing, et al.
Veröffentlicht: (2026)
von: Yang, Jing, et al.
Veröffentlicht: (2026)
Live Interactive Training for Video Segmentation
von: Yang, Xinyu, et al.
Veröffentlicht: (2026)
von: Yang, Xinyu, et al.
Veröffentlicht: (2026)
RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos
von: Xia, Hongchi, et al.
Veröffentlicht: (2024)
von: Xia, Hongchi, et al.
Veröffentlicht: (2024)
WorldAct: Activating Monolithic 3D Worlds into Interactive-Ready Object-Centric Scenes
von: Hu, Jichen, et al.
Veröffentlicht: (2026)
von: Hu, Jichen, et al.
Veröffentlicht: (2026)
Self-Supervised Learning for Real-World Object Detection: a Survey
von: Ciocarlan, Alina, et al.
Veröffentlicht: (2024)
von: Ciocarlan, Alina, et al.
Veröffentlicht: (2024)
Learning Object-Centric Representations Based on Slots in Real World Scenarios
von: Akan, Adil Kaan
Veröffentlicht: (2025)
von: Akan, Adil Kaan
Veröffentlicht: (2025)
Streamlining the Development of Active Learning Methods in Real-World Object Detection
von: Sbeyti, Moussa Kassem, et al.
Veröffentlicht: (2025)
von: Sbeyti, Moussa Kassem, et al.
Veröffentlicht: (2025)
Physics-Informed Learning of Characteristic Trajectories for Smoke Reconstruction
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
von: Wang, Yiming, et al.
Veröffentlicht: (2024)
OAT: Object-Level Attention Transformer for Gaze Scanpath Prediction
von: Fang, Yini, et al.
Veröffentlicht: (2024)
von: Fang, Yini, et al.
Veröffentlicht: (2024)
Chimera: Compositional Image Generation using Part-based Concepting
von: Singh, Shivam, et al.
Veröffentlicht: (2025)
von: Singh, Shivam, et al.
Veröffentlicht: (2025)
Open World Object Detection: A Survey
von: Li, Yiming, et al.
Veröffentlicht: (2024)
von: Li, Yiming, et al.
Veröffentlicht: (2024)
Hearing Hands: Generating Sounds from Physical Interactions in 3D Scenes
von: Dou, Yiming, et al.
Veröffentlicht: (2025)
von: Dou, Yiming, et al.
Veröffentlicht: (2025)
Benchmarking Object Detectors under Real-World Distribution Shifts in Satellite Imagery
von: Al-Emadi, Sara, et al.
Veröffentlicht: (2025)
von: Al-Emadi, Sara, et al.
Veröffentlicht: (2025)
WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models
von: Gu, Bohai, et al.
Veröffentlicht: (2026)
von: Gu, Bohai, et al.
Veröffentlicht: (2026)
Open-World Human-Object Interaction Detection via Multi-modal Prompts
von: Yang, Jie, et al.
Veröffentlicht: (2024)
von: Yang, Jie, et al.
Veröffentlicht: (2024)
TrackDeform3D: Markerless and Autonomous 3D Keypoint Tracking and Dataset Collection for Deformable Objects
von: Zong, Yeheng, et al.
Veröffentlicht: (2026)
von: Zong, Yeheng, et al.
Veröffentlicht: (2026)
IntrinsicReal: Adapting IntrinsicAnything from Synthetic to Real Objects
von: Wei, Xiaokang, et al.
Veröffentlicht: (2025)
von: Wei, Xiaokang, et al.
Veröffentlicht: (2025)
YOLO-World: Real-Time Open-Vocabulary Object Detection
von: Cheng, Tianheng, et al.
Veröffentlicht: (2024)
von: Cheng, Tianheng, et al.
Veröffentlicht: (2024)
SDVPT: Semantic-Driven Visual Prompt Tuning for Open-World Object Counting
von: Zhao, Yiming, et al.
Veröffentlicht: (2025)
von: Zhao, Yiming, et al.
Veröffentlicht: (2025)
LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model
von: Gao, Quankai, et al.
Veröffentlicht: (2026)
von: Gao, Quankai, et al.
Veröffentlicht: (2026)
CIS-BA: Continuous Interaction Space Based Backdoor Attack for Object Detection in the Real-World
von: Zhao, Shuxin, et al.
Veröffentlicht: (2025)
von: Zhao, Shuxin, et al.
Veröffentlicht: (2025)
XWOD: A Real-World Benchmark for Object Detection under Extreme Weather Conditions
von: Chen, Chih-Hsin, et al.
Veröffentlicht: (2026)
von: Chen, Chih-Hsin, et al.
Veröffentlicht: (2026)
PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing
von: Xu, Ruihang, et al.
Veröffentlicht: (2026)
von: Xu, Ruihang, et al.
Veröffentlicht: (2026)
TUMTraf EMOT: Event-Based Multi-Object Tracking Dataset and Baseline for Traffic Scenarios
von: Li, Mengyu, et al.
Veröffentlicht: (2025)
von: Li, Mengyu, et al.
Veröffentlicht: (2025)
Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration
von: Sun, Haoze, et al.
Veröffentlicht: (2024)
von: Sun, Haoze, et al.
Veröffentlicht: (2024)
Towards Generalizable Deepfake Detection with Spatial-Frequency Collaborative Learning and Hierarchical Cross-Modal Fusion
von: Qiao, Mengyu, et al.
Veröffentlicht: (2025)
von: Qiao, Mengyu, et al.
Veröffentlicht: (2025)
POV: Prompt-Oriented View-Agnostic Learning for Egocentric Hand-Object Interaction in the Multi-View World
von: Xu, Boshen, et al.
Veröffentlicht: (2024)
von: Xu, Boshen, et al.
Veröffentlicht: (2024)
OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model
von: Zhang, Zhenhao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenhao, et al.
Veröffentlicht: (2025)
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
von: Huang, Yifei, et al.
Veröffentlicht: (2024)
von: Huang, Yifei, et al.
Veröffentlicht: (2024)
Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization
von: Bharadwaj, Siddhant, et al.
Veröffentlicht: (2026)
von: Bharadwaj, Siddhant, et al.
Veröffentlicht: (2026)
InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing
von: Zhang, Jinlu, et al.
Veröffentlicht: (2025)
von: Zhang, Jinlu, et al.
Veröffentlicht: (2025)
ClickTrack: Towards Real-time Interactive Single Object Tracking
von: Wang, Kuiran, et al.
Veröffentlicht: (2024)
von: Wang, Kuiran, et al.
Veröffentlicht: (2024)
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
von: Wang, Zile, et al.
Veröffentlicht: (2026)
von: Wang, Zile, et al.
Veröffentlicht: (2026)
Objects With Lighting: A Real-World Dataset for Evaluating Reconstruction and Rendering for Object Relighting
von: Ummenhofer, Benjamin, et al.
Veröffentlicht: (2024)
von: Ummenhofer, Benjamin, et al.
Veröffentlicht: (2024)
Learning Human-Object Interaction as Groups
von: Hong, Jiajun, et al.
Veröffentlicht: (2025)
von: Hong, Jiajun, et al.
Veröffentlicht: (2025)
Single-Slice-to-3D Reconstruction in Medical Imaging and Natural Objects: A Comparative Benchmark with SAM 3D
von: Luo, Yan, et al.
Veröffentlicht: (2026)
von: Luo, Yan, et al.
Veröffentlicht: (2026)
Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities
von: Zadaianchuk, Andrii, et al.
Veröffentlicht: (2023)
von: Zadaianchuk, Andrii, et al.
Veröffentlicht: (2023)
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
von: Goswami, Raktim Gautam, et al.
Veröffentlicht: (2025)
von: Goswami, Raktim Gautam, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
von: Bansal, Siddhant, et al.
Veröffentlicht: (2024) -
Reconstructing Objects along Hand Interaction Timelines in Egocentric Video
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025) -
ICTPolarReal: A Polarized Reflection and Material Dataset of Real World Objects
von: Yang, Jing, et al.
Veröffentlicht: (2026) -
Live Interactive Training for Video Segmentation
von: Yang, Xinyu, et al.
Veröffentlicht: (2026) -
RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos
von: Xia, Hongchi, et al.
Veröffentlicht: (2024)