IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial Scenarios
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yifan, Chen, Yuhang, Dao, Anh, Li, Lichi, Cai, Zhongyi, Tan, Zhen, Chen, Tianlong, Kong, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
EQA-RM: A Generative Embodied Reward Model with Test-time Scaling
by: Chen, Yuhang, et al.
Published: (2025)
by: Chen, Yuhang, et al.
Published: (2025)
EfficientEQA: An Efficient Approach to Open-Vocabulary Embodied Question Answering
by: Cheng, Kai, et al.
Published: (2024)
by: Cheng, Kai, et al.
Published: (2024)
NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
Facial Affective Behavior Analysis with Instruction Tuning
by: Li, Yifan, et al.
Published: (2024)
by: Li, Yifan, et al.
Published: (2024)
StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
by: Wang, Yifei, et al.
Published: (2025)
by: Wang, Yifei, et al.
Published: (2025)
GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering
by: Saxena, Saumya, et al.
Published: (2024)
by: Saxena, Saumya, et al.
Published: (2024)
IndustryAssetEQA: A Neurosymbolic Operational Intelligence System for Embodied Question Answering in Industrial Asset Maintenance
by: Shyalika, Chathurangi, et al.
Published: (2026)
by: Shyalika, Chathurangi, et al.
Published: (2026)
Vision to Geometry: 3D Spatial Memory for Sequential Embodied MLLM Reasoning and Exploration
by: Cai, Zhongyi, et al.
Published: (2025)
by: Cai, Zhongyi, et al.
Published: (2025)
CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space
by: Zhao, Yong, et al.
Published: (2025)
by: Zhao, Yong, et al.
Published: (2025)
Improving Data Augmentation for Robust Visual Question Answering with Effective Curriculum Learning
by: Zheng, Yuhang, et al.
Published: (2024)
by: Zheng, Yuhang, et al.
Published: (2024)
Prune-Then-Plan: Step-Level Calibration for Stable Frontier Exploration in Embodied Question Answering
by: Frahm, Noah, et al.
Published: (2025)
by: Frahm, Noah, et al.
Published: (2025)
FAST-EQA: Efficient Embodied Question Answering with Global and Local Region Relevancy
by: Zhang, Haochen, et al.
Published: (2026)
by: Zhang, Haochen, et al.
Published: (2026)
Window Token Concatenation for Efficient Visual Large Language Models
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering
by: Jiang, Kaixuan, et al.
Published: (2025)
by: Jiang, Kaixuan, et al.
Published: (2025)
BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections
by: Varghese, Subin, et al.
Published: (2025)
by: Varghese, Subin, et al.
Published: (2025)
3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
by: Xu, Rongtao, et al.
Published: (2025)
by: Xu, Rongtao, et al.
Published: (2025)
Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
by: Li, Zechuan, et al.
Published: (2025)
by: Li, Zechuan, et al.
Published: (2025)
Research on Vision-Language Question Answering Models for Industrial Robots
by: Li, Ping, et al.
Published: (2026)
by: Li, Ping, et al.
Published: (2026)
VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
by: Li, Yuyi, et al.
Published: (2025)
by: Li, Yuyi, et al.
Published: (2025)
Embodied Understanding of Driving Scenarios
by: Zhou, Yunsong, et al.
Published: (2024)
by: Zhou, Yunsong, et al.
Published: (2024)
Open Set Face Forgery Detection via Dual-Level Evidence Collection
by: Cai, Zhongyi, et al.
Published: (2025)
by: Cai, Zhongyi, et al.
Published: (2025)
NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
by: Qian, Tianwen, et al.
Published: (2023)
by: Qian, Tianwen, et al.
Published: (2023)
Path-RAG: Knowledge-Guided Key Region Retrieval for Open-ended Pathology Visual Question Answering
by: Naeem, Awais, et al.
Published: (2024)
by: Naeem, Awais, et al.
Published: (2024)
Visual Environment-Interactive Planning for Embodied Complex-Question Answering
by: Lan, Ning, et al.
Published: (2025)
by: Lan, Ning, et al.
Published: (2025)
DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
by: Luo, Jingzhou, et al.
Published: (2025)
by: Luo, Jingzhou, et al.
Published: (2025)
Visual Large Language Models for Generalized and Specialized Applications
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Multi-Sourced Compositional Generalization in Visual Question Answering
by: Li, Chuanhao, et al.
Published: (2025)
by: Li, Chuanhao, et al.
Published: (2025)
Map-based Modular Approach for Zero-shot Embodied Question Answering
by: Sakamoto, Koya, et al.
Published: (2024)
by: Sakamoto, Koya, et al.
Published: (2024)
Memory-Guided View Refinement for Dynamic Human-in-the-loop EQA
by: Lu, Xin, et al.
Published: (2026)
by: Lu, Xin, et al.
Published: (2026)
MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
by: Nguyen, Hai-Dang, et al.
Published: (2025)
by: Nguyen, Hai-Dang, et al.
Published: (2025)
Explore until Confident: Efficient Exploration for Embodied Question Answering
by: Ren, Allen Z., et al.
Published: (2024)
by: Ren, Allen Z., et al.
Published: (2024)
Task-Aware Resolution Optimization for Visual Large Language Models
by: Luo, Weiqing, et al.
Published: (2025)
by: Luo, Weiqing, et al.
Published: (2025)
Pushing the Limits of Sparsity: A Bag of Tricks for Extreme Pruning
by: Li, Andy, et al.
Published: (2024)
by: Li, Andy, et al.
Published: (2024)
MerRec: A Large-scale Multipurpose Mercari Dataset for Consumer-to-Consumer Recommendation Systems
by: Li, Lichi, et al.
Published: (2024)
by: Li, Lichi, et al.
Published: (2024)
ConEQsA: Concurrent and Asynchronous Embodied Questions Scheduling and Answering
by: Wang, Haisheng, et al.
Published: (2025)
by: Wang, Haisheng, et al.
Published: (2025)
VietMEAgent: Culturally-Aware Few-Shot Multimodal Explanation for Vietnamese Visual Question Answering
by: Nguyen, Hai-Dang, et al.
Published: (2025)
by: Nguyen, Hai-Dang, et al.
Published: (2025)
EMOv2: Pushing 5M Vision Model Frontier
by: Zhang, Jiangning, et al.
Published: (2024)
by: Zhang, Jiangning, et al.
Published: (2024)
FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering
by: Zemskova, Tatiana, et al.
Published: (2026)
by: Zemskova, Tatiana, et al.
Published: (2026)
Mitigating Easy Option Bias in Multiple-Choice Question Answering
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
Similar Items
-
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
by: Li, Yifan, et al.
Published: (2025) -
EQA-RM: A Generative Embodied Reward Model with Test-time Scaling
by: Chen, Yuhang, et al.
Published: (2025) -
EfficientEQA: An Efficient Approach to Open-Vocabulary Embodied Question Answering
by: Cheng, Kai, et al.
Published: (2024) -
NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries
by: Wu, Tao, et al.
Published: (2024) -
Facial Affective Behavior Analysis with Instruction Tuning
by: Li, Yifan, et al.
Published: (2024)