Saved in:
| Main Authors: | Zhang, Sen, Li, Runmei, Deng, Shizhuang, Zheng, Zhichao, Zhang, Yuhe, Li, Jiani, Zhang, Kailun, Zhang, Tao, Wu, Wenjun, Wang, Qunbo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.27112 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reflecting with Two Voices: A Co-Adaptive Dual-Strategy Framework for LLM-Based Agent Decision Making
by: Zhang, Wentao, et al.
Published: (2025)
by: Zhang, Wentao, et al.
Published: (2025)
CogRail: Benchmarking VLMs in Cognitive Intrusion Perception for Intelligent Railway Transportation Systems
by: Tian, Yonglin, et al.
Published: (2026)
by: Tian, Yonglin, et al.
Published: (2026)
Hot Deformation Behavior and Processing Map of Eutectoid Pearlite Rail Steel
by: Haibo Feng, et al.
Published: (2025)
by: Haibo Feng, et al.
Published: (2025)
Knowledge Condensation and Reasoning for Knowledge-based VQA
by: Hao, Dongze, et al.
Published: (2024)
by: Hao, Dongze, et al.
Published: (2024)
An LLM-based Framework for Human-Swarm Teaming Cognition in Disaster Search and Rescue
by: Ji, Kailun, et al.
Published: (2025)
by: Ji, Kailun, et al.
Published: (2025)
Marginal Debiased Network for Fair Visual Recognition
by: Wang, Mei, et al.
Published: (2024)
by: Wang, Mei, et al.
Published: (2024)
Prompt-tuning for Clickbait Detection via Text Summarization
by: Deng, Haoxiang, et al.
Published: (2024)
by: Deng, Haoxiang, et al.
Published: (2024)
QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA
by: Li, Shuai, et al.
Published: (2025)
by: Li, Shuai, et al.
Published: (2025)
DPC-VQA: Decoupling Quality Perception and Residual Calibration for Video Quality Assessment
by: Li, Xinyue, et al.
Published: (2026)
by: Li, Xinyue, et al.
Published: (2026)
Multi-Scale Implicit Transformer with Re-parameterize for Arbitrary-Scale Super-Resolution
by: Zhu, Jinchen, et al.
Published: (2024)
by: Zhu, Jinchen, et al.
Published: (2024)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
by: Zhang, Chengyi, et al.
Published: (2026)
by: Zhang, Chengyi, et al.
Published: (2026)
Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge
by: Xu, Yinsong, et al.
Published: (2026)
by: Xu, Yinsong, et al.
Published: (2026)
MaS-VQA: A Mask-and-Select Framework for Knowledge-Based Visual Question Answering
by: Mao, Xianwei, et al.
Published: (2026)
by: Mao, Xianwei, et al.
Published: (2026)
Analysis of Rail Transit Operation and Maintenance Fault Recognition Considering Bayesian Knowledge Recognition Algorithm
by: Yanyan Zhang
Published: (2025)
by: Yanyan Zhang
Published: (2025)
Nonlinear predictive control of a quadrotor based on differential flatness
by: Runmei Zhang, et al.
Published: (2025)
by: Runmei Zhang, et al.
Published: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
by: Zhang, Xiaoman, et al.
Published: (2023)
by: Zhang, Xiaoman, et al.
Published: (2023)
Mechanical Mode Analysis of Centrifugal Pump Impellers Based on Numerical Simulations
by: Jiani Zhang
Published: (2025)
by: Jiani Zhang
Published: (2025)
Visual Robustness Benchmark for Visual Question Answering (VQA)
by: Ishmam, Md Farhan, et al.
Published: (2024)
by: Ishmam, Md Farhan, et al.
Published: (2024)
SpecVQA: A Benchmark for Spectral Understanding and Visual Question Answering in Scientific Images
by: Shen, Jialu, et al.
Published: (2026)
by: Shen, Jialu, et al.
Published: (2026)
StoryTailor:A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual Narratives
by: Hu, Jinghao, et al.
Published: (2026)
by: Hu, Jinghao, et al.
Published: (2026)
KNVQA: A Benchmark for evaluation knowledge-based VQA
by: Cheng, Sirui, et al.
Published: (2023)
by: Cheng, Sirui, et al.
Published: (2023)
Study on Elevated Interval Evacuation of Metro Train considering the Influence of Fire, Railings, and Track Bed
by: Jianyao Tu, et al.
Published: (2024)
by: Jianyao Tu, et al.
Published: (2024)
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
by: Chen, Pingyi, et al.
Published: (2024)
by: Chen, Pingyi, et al.
Published: (2024)
NEUROLOGIC: From Neural Representations to Interpretable Logic Rules
by: Geng, Chuqin, et al.
Published: (2025)
by: Geng, Chuqin, et al.
Published: (2025)
SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory
by: Alam, Samiul, et al.
Published: (2026)
by: Alam, Samiul, et al.
Published: (2026)
SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent
by: Li, Yangning, et al.
Published: (2024)
by: Li, Yangning, et al.
Published: (2024)
PVPUFormer: Probabilistic Visual Prompt Unified Transformer for Interactive Image Segmentation
by: Zhang, Xu, et al.
Published: (2023)
by: Zhang, Xu, et al.
Published: (2023)
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
by: Zhang, Nonghai, et al.
Published: (2025)
by: Zhang, Nonghai, et al.
Published: (2025)
OccSTeP: Benchmarking 4D Occupancy Spatio-Temporal Persistence
by: Zheng, Yu, et al.
Published: (2025)
by: Zheng, Yu, et al.
Published: (2025)
DWAFM: Dynamic Weighted Graph Structure Embedding Integrated with Attention and Frequency-Domain MLPs for Traffic Forecasting
by: Shi, Sen, et al.
Published: (2026)
by: Shi, Sen, et al.
Published: (2026)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
by: Wang, Siting, et al.
Published: (2025)
by: Wang, Siting, et al.
Published: (2025)
Tetris: A Compilation Framework for VQA Applications in Quantum Computing
by: Jin, Yuwei, et al.
Published: (2023)
by: Jin, Yuwei, et al.
Published: (2023)
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
by: Ma, Dongsheng, et al.
Published: (2026)
by: Ma, Dongsheng, et al.
Published: (2026)
Diffusion Signals Reveal Hidden Connections: A Physics-Inspired Framework for Link Prediction via Personalized PageRank Signals
by: Deng, Huilin Wang Wenjun Zhang Weibing
Published: (2025)
by: Deng, Huilin Wang Wenjun Zhang Weibing
Published: (2025)
VQA$^2$: Visual Question Answering for Video Quality Assessment
by: Jia, Ziheng, et al.
Published: (2024)
by: Jia, Ziheng, et al.
Published: (2024)
Revisiting Plasticity in Visual Reinforcement Learning: Data, Modules and Training Stages
by: Ma, Guozheng, et al.
Published: (2023)
by: Ma, Guozheng, et al.
Published: (2023)
Audio Outperforms Text for Visual Decoding
by: Zhang, Zhengdi, et al.
Published: (2026)
by: Zhang, Zhengdi, et al.
Published: (2026)
Seeing Beyond Words: MatVQA for Challenging Visual-Scientific Reasoning in Materials Science
by: Wu, Sifan, et al.
Published: (2025)
by: Wu, Sifan, et al.
Published: (2025)
RankDVQA: Deep VQA based on Ranking-inspired Hybrid Training
by: Feng, Chen, et al.
Published: (2022)
by: Feng, Chen, et al.
Published: (2022)
Similar Items
-
Reflecting with Two Voices: A Co-Adaptive Dual-Strategy Framework for LLM-Based Agent Decision Making
by: Zhang, Wentao, et al.
Published: (2025) -
CogRail: Benchmarking VLMs in Cognitive Intrusion Perception for Intelligent Railway Transportation Systems
by: Tian, Yonglin, et al.
Published: (2026) -
Hot Deformation Behavior and Processing Map of Eutectoid Pearlite Rail Steel
by: Haibo Feng, et al.
Published: (2025) -
Knowledge Condensation and Reasoning for Knowledge-based VQA
by: Hao, Dongze, et al.
Published: (2024) -
An LLM-based Framework for Human-Swarm Teaming Cognition in Disaster Search and Rescue
by: Ji, Kailun, et al.
Published: (2025)