Explaining the Unseen: Multimodal Vision-Language Reasoning for Situational Awareness in Underground Mining Disasters
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jewel, Mizanur Rahman, Elmahallawy, Mohamed, Madria, Sanjay, Frimpong, Samuel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DIS-Mine: Instance Segmentation for Disaster-Awareness in Poor-Light Condition in Underground Mines
von: Jewel, Mizanur Rahman, et al.
Veröffentlicht: (2024)
von: Jewel, Mizanur Rahman, et al.
Veröffentlicht: (2024)
Secure and Privacy-Preserving Federated Learning for Next-Generation Underground Mine Safety
von: Elmahallawy, Mohamed, et al.
Veröffentlicht: (2025)
von: Elmahallawy, Mohamed, et al.
Veröffentlicht: (2025)
Detecting Untargeted Attacks and Mitigating Unreliable Updates in Federated Learning for Underground Mining Operations
von: Rahman, Md Sazedur, et al.
Veröffentlicht: (2025)
von: Rahman, Md Sazedur, et al.
Veröffentlicht: (2025)
Prototype Fusion: A Training-Free Multi-Layer Approach to OOD Detection
von: Gul, Shreen, et al.
Veröffentlicht: (2026)
von: Gul, Shreen, et al.
Veröffentlicht: (2026)
FisherMask: Enhancing Neural Network Labeling Efficiency in Image Classification Using Fisher Information
von: Gul, Shreen, et al.
Veröffentlicht: (2024)
von: Gul, Shreen, et al.
Veröffentlicht: (2024)
LPLgrad: Optimizing Active Learning Through Gradient Norm Sample Selection and Auxiliary Model Training
von: Gul, Shreen, et al.
Veröffentlicht: (2024)
von: Gul, Shreen, et al.
Veröffentlicht: (2024)
Future Mining: Learning for Safety and Security
von: Rahman, Md Sazedur, et al.
Veröffentlicht: (2026)
von: Rahman, Md Sazedur, et al.
Veröffentlicht: (2026)
Landmark-based Localization using Stereo Vision and Deep Learning in GPS-Denied Battlefield Environment
von: Sapkota, Ganesh, et al.
Veröffentlicht: (2024)
von: Sapkota, Ganesh, et al.
Veröffentlicht: (2024)
Landmark Stereo Dataset for Landmark Recognition and Moving Node Localization in a Non-GPS Battlefield Environment
von: Sapkota, Ganesh, et al.
Veröffentlicht: (2024)
von: Sapkota, Ganesh, et al.
Veröffentlicht: (2024)
Secure Navigation using Landmark-based Localization in a GPS-denied Environment
von: Sapkota, Ganesh, et al.
Veröffentlicht: (2024)
von: Sapkota, Ganesh, et al.
Veröffentlicht: (2024)
CAV-AD: A Robust Framework for Detection of Anomalous Data and Malicious Sensors in CAV Networks
von: Rahman, Md Sazedur, et al.
Veröffentlicht: (2024)
von: Rahman, Md Sazedur, et al.
Veröffentlicht: (2024)
Enhancing Vision Language Models with Logic Reasoning for Situational Awareness
von: Pradeep, Pavana, et al.
Veröffentlicht: (2026)
von: Pradeep, Pavana, et al.
Veröffentlicht: (2026)
Toward Generalized Detection of Synthetic Media: Limitations, Challenges, and the Path to Multimodal Solutions
von: Hussain, Redwan, et al.
Veröffentlicht: (2025)
von: Hussain, Redwan, et al.
Veröffentlicht: (2025)
Situational Awareness Matters in 3D Vision Language Reasoning
von: Man, Yunze, et al.
Veröffentlicht: (2024)
von: Man, Yunze, et al.
Veröffentlicht: (2024)
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
DisasterInsight: A Multimodal Benchmark for Function-Aware and Grounded Disaster Assessment
von: Tehrani, Sara, et al.
Veröffentlicht: (2026)
von: Tehrani, Sara, et al.
Veröffentlicht: (2026)
Unlocking Neural Transparency: Jacobian Maps for Explainable AI in Alzheimer's Detection
von: Mustafa, Yasmine, et al.
Veröffentlicht: (2025)
von: Mustafa, Yasmine, et al.
Veröffentlicht: (2025)
Question Aware Vision Transformer for Multimodal Reasoning
von: Ganz, Roy, et al.
Veröffentlicht: (2024)
von: Ganz, Roy, et al.
Veröffentlicht: (2024)
Edge-Optimized Vision-Language Models for Underground Infrastructure Assessment
von: Lopez, Johny J., et al.
Veröffentlicht: (2026)
von: Lopez, Johny J., et al.
Veröffentlicht: (2026)
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
von: Kim, Geewook, et al.
Veröffentlicht: (2024)
von: Kim, Geewook, et al.
Veröffentlicht: (2024)
Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
von: Zhou, Yuchen, et al.
Veröffentlicht: (2025)
von: Zhou, Yuchen, et al.
Veröffentlicht: (2025)
Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model
von: Liu, Ruiping, et al.
Veröffentlicht: (2025)
von: Liu, Ruiping, et al.
Veröffentlicht: (2025)
MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models
von: Xia, Yinan, et al.
Veröffentlicht: (2025)
von: Xia, Yinan, et al.
Veröffentlicht: (2025)
Empowering Large Language Models with 3D Situation Awareness
von: Yuan, Zhihao, et al.
Veröffentlicht: (2025)
von: Yuan, Zhihao, et al.
Veröffentlicht: (2025)
Efficient Brain Imaging Analysis for Alzheimer's and Dementia Detection Using Convolution-Derivative Operations
von: Mustafa, Yasmine, et al.
Veröffentlicht: (2024)
von: Mustafa, Yasmine, et al.
Veröffentlicht: (2024)
PoseGAM: Robust Unseen Object Pose Estimation via Geometry-Aware Multi-View Reasoning
von: Chen, Jianqi, et al.
Veröffentlicht: (2025)
von: Chen, Jianqi, et al.
Veröffentlicht: (2025)
MPCAR: Multi-Perspective Contextual Augmentation for Enhanced Visual Reasoning in Large Vision-Language Models
von: Rahman, Amirul, et al.
Veröffentlicht: (2025)
von: Rahman, Amirul, et al.
Veröffentlicht: (2025)
Vision-Based Localization and LLM-based Navigation for Indoor Environments
von: Rahimi, Keyan, et al.
Veröffentlicht: (2025)
von: Rahimi, Keyan, et al.
Veröffentlicht: (2025)
MedDChest: A Content-Aware Multimodal Foundational Vision Model for Thoracic Imaging
von: Soliman, Mahmoud, et al.
Veröffentlicht: (2025)
von: Soliman, Mahmoud, et al.
Veröffentlicht: (2025)
Evolution of ReID: From Early Methods to LLM Integration
von: Bhuiyan, Amran, et al.
Veröffentlicht: (2025)
von: Bhuiyan, Amran, et al.
Veröffentlicht: (2025)
Multimodal 3D Object Detection on Unseen Domains
von: Hegde, Deepti, et al.
Veröffentlicht: (2024)
von: Hegde, Deepti, et al.
Veröffentlicht: (2024)
BanglaMM-Disaster: A Multimodal Transformer-Based Deep Learning Framework for Multiclass Disaster Classification in Bangla
von: Islam, Ariful, et al.
Veröffentlicht: (2025)
von: Islam, Ariful, et al.
Veröffentlicht: (2025)
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
von: Zhong, Qihuang, et al.
Veröffentlicht: (2026)
von: Zhong, Qihuang, et al.
Veröffentlicht: (2026)
Hallucination-Aware Multimodal Benchmark for Gastrointestinal Image Analysis with Large Vision-Language Models
von: Khanal, Bidur, et al.
Veröffentlicht: (2025)
von: Khanal, Bidur, et al.
Veröffentlicht: (2025)
Recov-Vision: Linking Street View Imagery and Vision-Language Models for Post-Disaster Recovery
von: Xiao, Yiming, et al.
Veröffentlicht: (2025)
von: Xiao, Yiming, et al.
Veröffentlicht: (2025)
ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining
von: Huang, Zhenyang, et al.
Veröffentlicht: (2025)
von: Huang, Zhenyang, et al.
Veröffentlicht: (2025)
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
von: Khan, Mohammed Safi Ur Rahman, et al.
Veröffentlicht: (2026)
von: Khan, Mohammed Safi Ur Rahman, et al.
Veröffentlicht: (2026)
SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation
von: Du, Hao, et al.
Veröffentlicht: (2025)
von: Du, Hao, et al.
Veröffentlicht: (2025)
Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DIS-Mine: Instance Segmentation for Disaster-Awareness in Poor-Light Condition in Underground Mines
von: Jewel, Mizanur Rahman, et al.
Veröffentlicht: (2024) -
Secure and Privacy-Preserving Federated Learning for Next-Generation Underground Mine Safety
von: Elmahallawy, Mohamed, et al.
Veröffentlicht: (2025) -
Detecting Untargeted Attacks and Mitigating Unreliable Updates in Federated Learning for Underground Mining Operations
von: Rahman, Md Sazedur, et al.
Veröffentlicht: (2025) -
Prototype Fusion: A Training-Free Multi-Layer Approach to OOD Detection
von: Gul, Shreen, et al.
Veröffentlicht: (2026) -
FisherMask: Enhancing Neural Network Labeling Efficiency in Image Classification Using Fisher Information
von: Gul, Shreen, et al.
Veröffentlicht: (2024)