ErgoChat: a Visual Query System for the Ergonomic Risk Assessment of Construction Workers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fan, Chao, Mei, Qipei, Wang, Xiaonan, Li, Xinming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TrajGATFormer: A Graph-Based Transformer Approach for Worker and Obstacle Trajectory Prediction in Off-site Construction Environments
von: Alduais, Mohammed, et al.
Veröffentlicht: (2025)
von: Alduais, Mohammed, et al.
Veröffentlicht: (2025)
ErgoExplorer: Interactive Ergonomic Risk Assessment from Video Collections
von: Fernández, Manlio Massiris, et al.
Veröffentlicht: (2022)
von: Fernández, Manlio Massiris, et al.
Veröffentlicht: (2022)
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
von: Li, Xu, et al.
Veröffentlicht: (2025)
von: Li, Xu, et al.
Veröffentlicht: (2025)
TRACE: Transformer-based Risk Assessment for Clinical Evaluation
von: Christopoulos, Dionysis, et al.
Veröffentlicht: (2024)
von: Christopoulos, Dionysis, et al.
Veröffentlicht: (2024)
CC-GRMAS: A Multi-Agent Graph Neural System for Spatiotemporal Landslide Risk Assessment in High Mountain Asia
von: Panchal, Mihir, et al.
Veröffentlicht: (2025)
von: Panchal, Mihir, et al.
Veröffentlicht: (2025)
Vision-Language Models for Ergonomic Assessment of Manual Lifting Tasks: Estimating Horizontal and Vertical Hand Distances from RGB Video
von: Rajabi, Mohammad Sadra, et al.
Veröffentlicht: (2026)
von: Rajabi, Mohammad Sadra, et al.
Veröffentlicht: (2026)
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
von: Li, Jialuo, et al.
Veröffentlicht: (2025)
von: Li, Jialuo, et al.
Veröffentlicht: (2025)
SparVAR: Exploring Sparsity in Visual AutoRegressive Modeling for Training-Free Acceleration
von: Li, Zekun, et al.
Veröffentlicht: (2026)
von: Li, Zekun, et al.
Veröffentlicht: (2026)
ScaleKD: Strong Vision Transformers Could Be Excellent Teachers
von: Fan, Jiawei, et al.
Veröffentlicht: (2024)
von: Fan, Jiawei, et al.
Veröffentlicht: (2024)
MCAT: Visual Query-Based Localization of Standard Anatomical Clips in Fetal Ultrasound Videos Using Multi-Tier Class-Aware Token Transformer
von: Mishra, Divyanshu, et al.
Veröffentlicht: (2025)
von: Mishra, Divyanshu, et al.
Veröffentlicht: (2025)
Explainable Cross-Disease Reasoning for Cardiovascular Risk Assessment from Low-Dose Computed Tomography
von: Zhang, Yifei, et al.
Veröffentlicht: (2025)
von: Zhang, Yifei, et al.
Veröffentlicht: (2025)
V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
von: Fan, Chenrui, et al.
Veröffentlicht: (2025)
Lessons and Insights from a Unifying Study of Parameter-Efficient Fine-Tuning (PEFT) in Visual Recognition
von: Mai, Zheda, et al.
Veröffentlicht: (2024)
von: Mai, Zheda, et al.
Veröffentlicht: (2024)
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
von: Wang, Yiheng, et al.
Veröffentlicht: (2026)
von: Wang, Yiheng, et al.
Veröffentlicht: (2026)
Cosmos-LLaVA: Chatting with the Visual Cosmos-LLaVA: Görselle Sohbet Etmek
von: Zeer, Ahmed, et al.
Veröffentlicht: (2024)
von: Zeer, Ahmed, et al.
Veröffentlicht: (2024)
Nacala-Roof-Material: Drone Imagery for Roof Detection, Classification, and Segmentation to Support Mosquito-borne Disease Risk Assessment
von: Guthula, Venkanna Babu, et al.
Veröffentlicht: (2024)
von: Guthula, Venkanna Babu, et al.
Veröffentlicht: (2024)
Masked Multi-Query Slot Attention for Unsupervised Object Discovery
von: Pramanik, Rishav, et al.
Veröffentlicht: (2024)
von: Pramanik, Rishav, et al.
Veröffentlicht: (2024)
Learning Semantic Segmentation with Query Points Supervision on Aerial Images
von: Rivier, Santiago, et al.
Veröffentlicht: (2023)
von: Rivier, Santiago, et al.
Veröffentlicht: (2023)
Effectiveness Assessment of Recent Large Vision-Language Models
von: Jiang, Yao, et al.
Veröffentlicht: (2024)
von: Jiang, Yao, et al.
Veröffentlicht: (2024)
Reversible Unfolding Network for Concealed Visual Perception with Generative Refinement
von: He, Chunming, et al.
Veröffentlicht: (2025)
von: He, Chunming, et al.
Veröffentlicht: (2025)
A Comprehensive Review of Machine Learning Advances on Data Change: A Cross-Field Perspective
von: Li, Jeng-Lin, et al.
Veröffentlicht: (2024)
von: Li, Jeng-Lin, et al.
Veröffentlicht: (2024)
Align Your Query: Representation Alignment for Multimodality Medical Object Detection
von: Seo, Ara, et al.
Veröffentlicht: (2025)
von: Seo, Ara, et al.
Veröffentlicht: (2025)
Hard-label based Small Query Black-box Adversarial Attack
von: Park, Jeonghwan, et al.
Veröffentlicht: (2024)
von: Park, Jeonghwan, et al.
Veröffentlicht: (2024)
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
von: Mai, Zheda, et al.
Veröffentlicht: (2025)
von: Mai, Zheda, et al.
Veröffentlicht: (2025)
Learning Concept-Based Causal Transition and Symbolic Reasoning for Visual Planning
von: Qian, Yilue, et al.
Veröffentlicht: (2023)
von: Qian, Yilue, et al.
Veröffentlicht: (2023)
We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
von: Qiao, Runqi, et al.
Veröffentlicht: (2025)
von: Qiao, Runqi, et al.
Veröffentlicht: (2025)
Seeing Unseen: Discover Novel Biomedical Concepts via Geometry-Constrained Probabilistic Modeling
von: Fan, Jianan, et al.
Veröffentlicht: (2024)
von: Fan, Jianan, et al.
Veröffentlicht: (2024)
Can ChatGPT Learn My Life From a Week of First-Person Video?
von: Harris, Keegan
Veröffentlicht: (2025)
von: Harris, Keegan
Veröffentlicht: (2025)
KernelWarehouse: Rethinking the Design of Dynamic Convolution
von: Li, Chao, et al.
Veröffentlicht: (2024)
von: Li, Chao, et al.
Veröffentlicht: (2024)
TextDestroyer: A Training- and Annotation-Free Diffusion Method for Destroying Anomal Text from Images
von: Li, Mengcheng, et al.
Veröffentlicht: (2024)
von: Li, Mengcheng, et al.
Veröffentlicht: (2024)
Optimizing Contrail Detection: A Deep Learning Approach with EfficientNet-b4 Encoding
von: Lin, Qunwei, et al.
Veröffentlicht: (2024)
von: Lin, Qunwei, et al.
Veröffentlicht: (2024)
Moving Healthcare AI-Support Systems for Visually Detectable Diseases onto Constrained Devices
von: Watt, Tess, et al.
Veröffentlicht: (2024)
von: Watt, Tess, et al.
Veröffentlicht: (2024)
Relational Visual Similarity
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
Token Activation Map to Visually Explain Multimodal LLMs
von: Li, Yi, et al.
Veröffentlicht: (2025)
von: Li, Yi, et al.
Veröffentlicht: (2025)
Uncertainty-DTW for Sequences and Visual Tokens
von: Wang, Lei, et al.
Veröffentlicht: (2026)
von: Wang, Lei, et al.
Veröffentlicht: (2026)
Text-to-CadQuery: A New Paradigm for CAD Generation with Scalable Large Model Capabilities
von: Xie, Haoyang, et al.
Veröffentlicht: (2025)
von: Xie, Haoyang, et al.
Veröffentlicht: (2025)
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
von: Du, Zilin, et al.
Veröffentlicht: (2024)
von: Du, Zilin, et al.
Veröffentlicht: (2024)
StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
von: Li, Bingyu, et al.
Veröffentlicht: (2024)
von: Li, Bingyu, et al.
Veröffentlicht: (2024)
Jailbreaking Vision-Language Models Through the Visual Modality
von: Azulay, Aharon, et al.
Veröffentlicht: (2026)
von: Azulay, Aharon, et al.
Veröffentlicht: (2026)
A Large-scale Medical Visual Task Adaptation Benchmark
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TrajGATFormer: A Graph-Based Transformer Approach for Worker and Obstacle Trajectory Prediction in Off-site Construction Environments
von: Alduais, Mohammed, et al.
Veröffentlicht: (2025) -
ErgoExplorer: Interactive Ergonomic Risk Assessment from Video Collections
von: Fernández, Manlio Massiris, et al.
Veröffentlicht: (2022) -
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
von: Li, Xu, et al.
Veröffentlicht: (2025) -
TRACE: Transformer-based Risk Assessment for Clinical Evaluation
von: Christopoulos, Dionysis, et al.
Veröffentlicht: (2024) -
CC-GRMAS: A Multi-Agent Graph Neural System for Spatiotemporal Landslide Risk Assessment in High Mountain Asia
von: Panchal, Mihir, et al.
Veröffentlicht: (2025)