AI-driven visual monitoring of industrial assembly tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Nardon, Mattia, Messelodi, Stefano, Granata, Antonio, Poiesi, Fabio, Danese, Alberto, Boscaini, Davide |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accurate and efficient zero-shot 6D pose estimation with frozen foundation models
by: Caraffa, Andrea, et al.
Published: (2025)
by: Caraffa, Andrea, et al.
Published: (2025)
Exploring Fine-grained Retail Product Discrimination with Zero-shot Object Classification Using Vision-Language Models
by: Tur, Anil Osman, et al.
Published: (2024)
by: Tur, Anil Osman, et al.
Published: (2024)
Leveraging Confident Image Regions for Source-Free Domain-Adaptive Object Detection
by: Mekhalfi, Mohamed Lamine, et al.
Published: (2025)
by: Mekhalfi, Mohamed Lamine, et al.
Published: (2025)
Distilling 3D distinctive local descriptors for 6D pose estimation
by: Hamza, Amir, et al.
Published: (2025)
by: Hamza, Amir, et al.
Published: (2025)
FreeZe: Training-free zero-shot 6D pose estimation with geometric and vision foundation models
by: Caraffa, Andrea, et al.
Published: (2023)
by: Caraffa, Andrea, et al.
Published: (2023)
CHIP: A multi-sensor dataset for 6D pose estimation of chairs in industrial settings
by: Nardon, Mattia, et al.
Published: (2025)
by: Nardon, Mattia, et al.
Published: (2025)
Open-vocabulary object 6D pose estimation
by: Corsetti, Jaime, et al.
Published: (2023)
by: Corsetti, Jaime, et al.
Published: (2023)
Functionality understanding and segmentation in 3D scenes
by: Corsetti, Jaime, et al.
Published: (2024)
by: Corsetti, Jaime, et al.
Published: (2024)
Generative 6D Pose Estimation via Conditional Flow Matching
by: Hamza, Amir, et al.
Published: (2026)
by: Hamza, Amir, et al.
Published: (2026)
3D Part Segmentation via Geometric Aggregation of 2D Visual Features
by: Garosi, Marco, et al.
Published: (2024)
by: Garosi, Marco, et al.
Published: (2024)
High-resolution open-vocabulary object 6D pose estimation
by: Corsetti, Jaime, et al.
Published: (2024)
by: Corsetti, Jaime, et al.
Published: (2024)
Action-guided generation of 3D functionality segmentation data
by: Corsetti, Jaime, et al.
Published: (2025)
by: Corsetti, Jaime, et al.
Published: (2025)
Geometrically-driven Aggregation for Zero-shot 3D Point Cloud Understanding
by: Mei, Guofeng, et al.
Published: (2023)
by: Mei, Guofeng, et al.
Published: (2023)
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
by: Li, Jinlong, et al.
Published: (2025)
by: Li, Jinlong, et al.
Published: (2025)
Event-based dataset for the detection and classification of manufacturing assembly tasks
by: Duarte, Laura, et al.
Published: (2024)
by: Duarte, Laura, et al.
Published: (2024)
An analysis of vision-language models for fabric retrieval
by: Giuliari, Francesco, et al.
Published: (2025)
by: Giuliari, Francesco, et al.
Published: (2025)
OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance Fields
by: Weijler, Lisa, et al.
Published: (2025)
by: Weijler, Lisa, et al.
Published: (2025)
Novel class discovery meets foundation models for 3D semantic segmentation
by: Riz, Luigi, et al.
Published: (2023)
by: Riz, Luigi, et al.
Published: (2023)
Vocabulary-Free 3D Instance Segmentation with Vision and Language Assistant
by: Mei, Guofeng, et al.
Published: (2024)
by: Mei, Guofeng, et al.
Published: (2024)
6DGS: 6D Pose Estimation from a Single Image and a 3D Gaussian Splatting Model
by: Bortolon, Matteo, et al.
Published: (2024)
by: Bortolon, Matteo, et al.
Published: (2024)
Wild Berry image dataset collected in Finnish forests and peatlands using drones
by: Riz, Luigi, et al.
Published: (2024)
by: Riz, Luigi, et al.
Published: (2024)
Fully-Geometric Cross-Attention for Point Cloud Registration
by: Wang, Weijie, et al.
Published: (2025)
by: Wang, Weijie, et al.
Published: (2025)
IFFNeRF: Initialisation Free and Fast 6DoF pose estimation from a single image and a NeRF model
by: Bortolon, Matteo, et al.
Published: (2024)
by: Bortolon, Matteo, et al.
Published: (2024)
The RaspGrade Dataset: Towards Automatic Raspberry Ripeness Grading with Deep Learning
by: Mekhalfi, Mohamed Lamine, et al.
Published: (2025)
by: Mekhalfi, Mohamed Lamine, et al.
Published: (2025)
Multimodal Fusion SLAM with Fourier Attention
by: Zhou, Youjie, et al.
Published: (2025)
by: Zhou, Youjie, et al.
Published: (2025)
Toward task-driven satellite image super-resolution
by: Ziaja, Maciej, et al.
Published: (2025)
by: Ziaja, Maciej, et al.
Published: (2025)
Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
by: Ren, Bin, et al.
Published: (2025)
by: Ren, Bin, et al.
Published: (2025)
ANTHROPOS-V: benchmarking the novel task of Crowd Volume Estimation
by: Collorone, Luca, et al.
Published: (2025)
by: Collorone, Luca, et al.
Published: (2025)
Vision encoders should be image size agnostic and task driven
by: Prisadnikov, Nedyalko, et al.
Published: (2025)
by: Prisadnikov, Nedyalko, et al.
Published: (2025)
A vision-based framework for human behavior understanding in industrial assembly lines
by: Papoutsakis, Konstantinos, et al.
Published: (2024)
by: Papoutsakis, Konstantinos, et al.
Published: (2024)
For a semiotic AI: Bridging computer vision and visual semiotics for computational observation of large scale facial image archives
by: Morra, Lia, et al.
Published: (2024)
by: Morra, Lia, et al.
Published: (2024)
ZeroReg: Zero-Shot Point Cloud Registration with Foundation Models
by: Wang, Weijie, et al.
Published: (2023)
by: Wang, Weijie, et al.
Published: (2023)
Flame quality monitoring of flare stack based on deep visual features
by: Mu, Xing
Published: (2024)
by: Mu, Xing
Published: (2024)
Efficient Encoder-Free Fourier-based 3D Large Multimodal Model
by: Mei, Guofeng, et al.
Published: (2026)
by: Mei, Guofeng, et al.
Published: (2026)
Learning reusable concepts across different egocentric video understanding tasks
by: Peirone, Simone Alberto, et al.
Published: (2025)
by: Peirone, Simone Alberto, et al.
Published: (2025)
GLASS: Graph and Vision-Language Assisted Semantic Shape Correspondence
by: Xiao, Qinfeng, et al.
Published: (2026)
by: Xiao, Qinfeng, et al.
Published: (2026)
GRASPLAT: Enabling dexterous grasping through novel view synthesis
by: Bortolon, Matteo, et al.
Published: (2025)
by: Bortolon, Matteo, et al.
Published: (2025)
PerLA: Perceptive 3D Language Assistant
by: Mei, Guofeng, et al.
Published: (2024)
by: Mei, Guofeng, et al.
Published: (2024)
FlowLet: Conditional 3D Brain MRI Synthesis using Wavelet Flow Matching
by: Danese, Danilo, et al.
Published: (2026)
by: Danese, Danilo, et al.
Published: (2026)
MoDiPO: text-to-motion alignment via AI-feedback-driven Direct Preference Optimization
by: Pappa, Massimiliano, et al.
Published: (2024)
by: Pappa, Massimiliano, et al.
Published: (2024)
Similar Items
-
Accurate and efficient zero-shot 6D pose estimation with frozen foundation models
by: Caraffa, Andrea, et al.
Published: (2025) -
Exploring Fine-grained Retail Product Discrimination with Zero-shot Object Classification Using Vision-Language Models
by: Tur, Anil Osman, et al.
Published: (2024) -
Leveraging Confident Image Regions for Source-Free Domain-Adaptive Object Detection
by: Mekhalfi, Mohamed Lamine, et al.
Published: (2025) -
Distilling 3D distinctive local descriptors for 6D pose estimation
by: Hamza, Amir, et al.
Published: (2025) -
FreeZe: Training-free zero-shot 6D pose estimation with geometric and vision foundation models
by: Caraffa, Andrea, et al.
Published: (2023)