ErgoChat: a Visual Query System for the Ergonomic Risk Assessment of Construction Workers
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Chao, Mei, Qipei, Wang, Xiaonan, Li, Xinming |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TrajGATFormer: A Graph-Based Transformer Approach for Worker and Obstacle Trajectory Prediction in Off-site Construction Environments
by: Alduais, Mohammed, et al.
Published: (2025)
by: Alduais, Mohammed, et al.
Published: (2025)
ErgoExplorer: Interactive Ergonomic Risk Assessment from Video Collections
by: Fernández, Manlio Massiris, et al.
Published: (2022)
by: Fernández, Manlio Massiris, et al.
Published: (2022)
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
by: Li, Xu, et al.
Published: (2025)
by: Li, Xu, et al.
Published: (2025)
TRACE: Transformer-based Risk Assessment for Clinical Evaluation
by: Christopoulos, Dionysis, et al.
Published: (2024)
by: Christopoulos, Dionysis, et al.
Published: (2024)
CC-GRMAS: A Multi-Agent Graph Neural System for Spatiotemporal Landslide Risk Assessment in High Mountain Asia
by: Panchal, Mihir, et al.
Published: (2025)
by: Panchal, Mihir, et al.
Published: (2025)
Vision-Language Models for Ergonomic Assessment of Manual Lifting Tasks: Estimating Horizontal and Vertical Hand Distances from RGB Video
by: Rajabi, Mohammad Sadra, et al.
Published: (2026)
by: Rajabi, Mohammad Sadra, et al.
Published: (2026)
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
by: Li, Jialuo, et al.
Published: (2025)
by: Li, Jialuo, et al.
Published: (2025)
SparVAR: Exploring Sparsity in Visual AutoRegressive Modeling for Training-Free Acceleration
by: Li, Zekun, et al.
Published: (2026)
by: Li, Zekun, et al.
Published: (2026)
ScaleKD: Strong Vision Transformers Could Be Excellent Teachers
by: Fan, Jiawei, et al.
Published: (2024)
by: Fan, Jiawei, et al.
Published: (2024)
MCAT: Visual Query-Based Localization of Standard Anatomical Clips in Fetal Ultrasound Videos Using Multi-Tier Class-Aware Token Transformer
by: Mishra, Divyanshu, et al.
Published: (2025)
by: Mishra, Divyanshu, et al.
Published: (2025)
Explainable Cross-Disease Reasoning for Cardiovascular Risk Assessment from Low-Dose Computed Tomography
by: Zhang, Yifei, et al.
Published: (2025)
by: Zhang, Yifei, et al.
Published: (2025)
V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
by: Fan, Chenrui, et al.
Published: (2025)
by: Fan, Chenrui, et al.
Published: (2025)
Lessons and Insights from a Unifying Study of Parameter-Efficient Fine-Tuning (PEFT) in Visual Recognition
by: Mai, Zheda, et al.
Published: (2024)
by: Mai, Zheda, et al.
Published: (2024)
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
by: Wang, Yiheng, et al.
Published: (2026)
by: Wang, Yiheng, et al.
Published: (2026)
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
by: Zeer, Ahmed, et al.
Published: (2024)
by: Zeer, Ahmed, et al.
Published: (2024)
Nacala-Roof-Material: Drone Imagery for Roof Detection, Classification, and Segmentation to Support Mosquito-borne Disease Risk Assessment
by: Guthula, Venkanna Babu, et al.
Published: (2024)
by: Guthula, Venkanna Babu, et al.
Published: (2024)
Masked Multi-Query Slot Attention for Unsupervised Object Discovery
by: Pramanik, Rishav, et al.
Published: (2024)
by: Pramanik, Rishav, et al.
Published: (2024)
Learning Semantic Segmentation with Query Points Supervision on Aerial Images
by: Rivier, Santiago, et al.
Published: (2023)
by: Rivier, Santiago, et al.
Published: (2023)
Effectiveness Assessment of Recent Large Vision-Language Models
by: Jiang, Yao, et al.
Published: (2024)
by: Jiang, Yao, et al.
Published: (2024)
Reversible Unfolding Network for Concealed Visual Perception with Generative Refinement
by: He, Chunming, et al.
Published: (2025)
by: He, Chunming, et al.
Published: (2025)
A Comprehensive Review of Machine Learning Advances on Data Change: A Cross-Field Perspective
by: Li, Jeng-Lin, et al.
Published: (2024)
by: Li, Jeng-Lin, et al.
Published: (2024)
Align Your Query: Representation Alignment for Multimodality Medical Object Detection
by: Seo, Ara, et al.
Published: (2025)
by: Seo, Ara, et al.
Published: (2025)
Hard-label based Small Query Black-box Adversarial Attack
by: Park, Jeonghwan, et al.
Published: (2024)
by: Park, Jeonghwan, et al.
Published: (2024)
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
by: Mai, Zheda, et al.
Published: (2025)
by: Mai, Zheda, et al.
Published: (2025)
Learning Concept-Based Causal Transition and Symbolic Reasoning for Visual Planning
by: Qian, Yilue, et al.
Published: (2023)
by: Qian, Yilue, et al.
Published: (2023)
We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
by: Qiao, Runqi, et al.
Published: (2025)
by: Qiao, Runqi, et al.
Published: (2025)
Seeing Unseen: Discover Novel Biomedical Concepts via Geometry-Constrained Probabilistic Modeling
by: Fan, Jianan, et al.
Published: (2024)
by: Fan, Jianan, et al.
Published: (2024)
Can ChatGPT Learn My Life From a Week of First-Person Video?
by: Harris, Keegan
Published: (2025)
by: Harris, Keegan
Published: (2025)
KernelWarehouse: Rethinking the Design of Dynamic Convolution
by: Li, Chao, et al.
Published: (2024)
by: Li, Chao, et al.
Published: (2024)
TextDestroyer: A Training- and Annotation-Free Diffusion Method for Destroying Anomal Text from Images
by: Li, Mengcheng, et al.
Published: (2024)
by: Li, Mengcheng, et al.
Published: (2024)
Optimizing Contrail Detection: A Deep Learning Approach with EfficientNet-b4 Encoding
by: Lin, Qunwei, et al.
Published: (2024)
by: Lin, Qunwei, et al.
Published: (2024)
Moving Healthcare AI-Support Systems for Visually Detectable Diseases onto Constrained Devices
by: Watt, Tess, et al.
Published: (2024)
by: Watt, Tess, et al.
Published: (2024)
Relational Visual Similarity
by: Nguyen, Thao, et al.
Published: (2025)
by: Nguyen, Thao, et al.
Published: (2025)
Token Activation Map to Visually Explain Multimodal LLMs
by: Li, Yi, et al.
Published: (2025)
by: Li, Yi, et al.
Published: (2025)
Uncertainty-DTW for Sequences and Visual Tokens
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
Text-to-CadQuery: A New Paradigm for CAD Generation with Scalable Large Model Capabilities
by: Xie, Haoyang, et al.
Published: (2025)
by: Xie, Haoyang, et al.
Published: (2025)
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
by: Du, Zilin, et al.
Published: (2024)
by: Du, Zilin, et al.
Published: (2024)
StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
by: Li, Bingyu, et al.
Published: (2024)
by: Li, Bingyu, et al.
Published: (2024)
Jailbreaking Vision-Language Models Through the Visual Modality
by: Azulay, Aharon, et al.
Published: (2026)
by: Azulay, Aharon, et al.
Published: (2026)
A Large-scale Medical Visual Task Adaptation Benchmark
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
Similar Items
-
TrajGATFormer: A Graph-Based Transformer Approach for Worker and Obstacle Trajectory Prediction in Off-site Construction Environments
by: Alduais, Mohammed, et al.
Published: (2025) -
ErgoExplorer: Interactive Ergonomic Risk Assessment from Video Collections
by: Fernández, Manlio Massiris, et al.
Published: (2022) -
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
by: Li, Xu, et al.
Published: (2025) -
TRACE: Transformer-based Risk Assessment for Clinical Evaluation
by: Christopoulos, Dionysis, et al.
Published: (2024) -
CC-GRMAS: A Multi-Agent Graph Neural System for Spatiotemporal Landslide Risk Assessment in High Mountain Asia
by: Panchal, Mihir, et al.
Published: (2025)