AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Danrui, Zhang, Jiahao, Egger, Bernhard, Chatterjee, Moitreya, Lohit, Suhas, Marks, Tim K., Cherian, Anoop |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FreBIS: Frequency-Based Stratification for Neural Implicit Surface Representations
von: Sawada, Naoko, et al.
Veröffentlicht: (2025)
von: Sawada, Naoko, et al.
Veröffentlicht: (2025)
TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models
von: Ni, Haomiao, et al.
Veröffentlicht: (2024)
von: Ni, Haomiao, et al.
Veröffentlicht: (2024)
Programmatic Video Prediction Using Large Language Models
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
von: Cherian, Anoop, et al.
Veröffentlicht: (2025)
von: Cherian, Anoop, et al.
Veröffentlicht: (2025)
Improving Open-World Object Localization by Discovering Background
von: Singh, Ashish, et al.
Veröffentlicht: (2025)
von: Singh, Ashish, et al.
Veröffentlicht: (2025)
Multimodal Diffusion Bridge with Attention-Based SAR Fusion for Satellite Image Cloud Removal
von: Hu, Yuyang, et al.
Veröffentlicht: (2025)
von: Hu, Yuyang, et al.
Veröffentlicht: (2025)
Manual-PA: Learning 3D Part Assembly from Instruction Diagrams
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
von: Cherian, Anoop, et al.
Veröffentlicht: (2024)
von: Cherian, Anoop, et al.
Veröffentlicht: (2024)
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
Gear-NeRF: Free-Viewpoint Rendering and Tracking with Motion-aware Spatio-Temporal Sampling
von: Liu, Xinhang, et al.
Veröffentlicht: (2024)
von: Liu, Xinhang, et al.
Veröffentlicht: (2024)
MMHOI: Modeling Complex 3D Multi-Human Multi-Object Interactions
von: Kogashi, Kaen, et al.
Veröffentlicht: (2025)
von: Kogashi, Kaen, et al.
Veröffentlicht: (2025)
Time-Series U-Net with Recurrence for Noise-Robust Imaging Photoplethysmography
von: Shenoy, Vineet R., et al.
Veröffentlicht: (2025)
von: Shenoy, Vineet R., et al.
Veröffentlicht: (2025)
Recovering Pulse Waves from Video Using Deep Unrolling and Deep Equilibrium Models
von: Shenoy, Vineet R, et al.
Veröffentlicht: (2025)
von: Shenoy, Vineet R, et al.
Veröffentlicht: (2025)
GBOT: Graph-Based 3D Object Tracking for Augmented Reality-Assisted Assembly Guidance
von: Li, Shiyu, et al.
Veröffentlicht: (2024)
von: Li, Shiyu, et al.
Veröffentlicht: (2024)
LLMPhy: Parameter-Identifiable Physical Reasoning Combining Large Language Models and Physics Engines
von: Cherian, Anoop, et al.
Veröffentlicht: (2024)
von: Cherian, Anoop, et al.
Veröffentlicht: (2024)
Auto-Vocabulary 3D Object Detection
von: Zhang, Haomeng, et al.
Veröffentlicht: (2025)
von: Zhang, Haomeng, et al.
Veröffentlicht: (2025)
LLM-Guided Agentic Object Detection for Open-World Understanding
von: Mumcu, Furkan, et al.
Veröffentlicht: (2025)
von: Mumcu, Furkan, et al.
Veröffentlicht: (2025)
Towards Open-Vocabulary Multimodal 3D Object Detection with Attributes
von: Xiang, Xinhao, et al.
Veröffentlicht: (2025)
von: Xiang, Xinhao, et al.
Veröffentlicht: (2025)
ComplexVAD: Detecting Interaction Anomalies in Video
von: Mumcu, Furkan, et al.
Veröffentlicht: (2025)
von: Mumcu, Furkan, et al.
Veröffentlicht: (2025)
Multimodal 3D Object Detection on Unseen Domains
von: Hegde, Deepti, et al.
Veröffentlicht: (2024)
von: Hegde, Deepti, et al.
Veröffentlicht: (2024)
Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
von: An, Wenbin, et al.
Veröffentlicht: (2024)
von: An, Wenbin, et al.
Veröffentlicht: (2024)
Equivariant Spatio-Temporal Self-Supervision for LiDAR Object Detection
von: Hegde, Deepti, et al.
Veröffentlicht: (2024)
von: Hegde, Deepti, et al.
Veröffentlicht: (2024)
IMPACT: A Dataset for Multi-Granularity Human Procedural Action Understanding in Industrial Assembly
von: Wen, Di, et al.
Veröffentlicht: (2026)
von: Wen, Di, et al.
Veröffentlicht: (2026)
Find the Assembly Mistakes: Error Segmentation for Industrial Applications
von: Lehman, Dan, et al.
Veröffentlicht: (2024)
von: Lehman, Dan, et al.
Veröffentlicht: (2024)
Combinative Matching for Geometric Shape Assembly
von: Lee, Nahyuk, et al.
Veröffentlicht: (2025)
von: Lee, Nahyuk, et al.
Veröffentlicht: (2025)
Template-based Object Detection Using a Foundation Model
von: Braeutigam, Valentin, et al.
Veröffentlicht: (2026)
von: Braeutigam, Valentin, et al.
Veröffentlicht: (2026)
Aligning Step-by-Step Instructional Diagrams to Video Demonstrations
von: Zhang, Jiahao, et al.
Veröffentlicht: (2023)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2023)
Noise Consistency Regularization for Improved Subject-Driven Image Synthesis
von: Ni, Yao, et al.
Veröffentlicht: (2025)
von: Ni, Yao, et al.
Veröffentlicht: (2025)
Equivariant Flow Matching for Point Cloud Assembly
von: Wang, Ziming, et al.
Veröffentlicht: (2025)
von: Wang, Ziming, et al.
Veröffentlicht: (2025)
InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention
von: Xiang, Qiang, et al.
Veröffentlicht: (2025)
von: Xiang, Qiang, et al.
Veröffentlicht: (2025)
Temporally Grounding Instructional Diagrams in Unconstrained Videos
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly
von: Chetan, Aditya, et al.
Veröffentlicht: (2026)
von: Chetan, Aditya, et al.
Veröffentlicht: (2026)
A Probability-guided Sampler for Neural Implicit Surface Rendering
von: Pais, Gonçalo Dias, et al.
Veröffentlicht: (2025)
von: Pais, Gonçalo Dias, et al.
Veröffentlicht: (2025)
Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection
von: Mumcu, Furkan, et al.
Veröffentlicht: (2025)
von: Mumcu, Furkan, et al.
Veröffentlicht: (2025)
Is Video Anomaly Detection Misframed? Evidence from LLM-Based and Multi-Scene Models
von: Mumcu, Furkan, et al.
Veröffentlicht: (2026)
von: Mumcu, Furkan, et al.
Veröffentlicht: (2026)
Exploring Multi-modal Neural Scene Representations With Applications on Thermal Imaging
von: Özer, Mert, et al.
Veröffentlicht: (2024)
von: Özer, Mert, et al.
Veröffentlicht: (2024)
GeoGen: Geometry-Aware Generative Modeling via Signed Distance Functions
von: Esposito, Salvatore, et al.
Veröffentlicht: (2024)
von: Esposito, Salvatore, et al.
Veröffentlicht: (2024)
LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction
von: Ding, Tianye, et al.
Veröffentlicht: (2025)
von: Ding, Tianye, et al.
Veröffentlicht: (2025)
Understanding Dynamic Compute Allocation in Recurrent Transformers
von: Moosa, Ibraheem Muhammad, et al.
Veröffentlicht: (2026)
von: Moosa, Ibraheem Muhammad, et al.
Veröffentlicht: (2026)
ESA: Energy-Based Shot Assembly Optimization for Automatic Video Editing
von: Chen, Yaosen, et al.
Veröffentlicht: (2025)
von: Chen, Yaosen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FreBIS: Frequency-Based Stratification for Neural Implicit Surface Representations
von: Sawada, Naoko, et al.
Veröffentlicht: (2025) -
TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models
von: Ni, Haomiao, et al.
Veröffentlicht: (2024) -
Programmatic Video Prediction Using Large Language Models
von: Tang, Hao, et al.
Veröffentlicht: (2025) -
WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
von: Cherian, Anoop, et al.
Veröffentlicht: (2025) -
Improving Open-World Object Localization by Discovering Background
von: Singh, Ashish, et al.
Veröffentlicht: (2025)