Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Yin, Jianghao, Li, Qingbin, Sun, Kun, Ding, Cheng, Wang, Jie, Chen, Qin, Zhou, Jie, Wang, Nan, Li, Changqing, Wu, Pei, Xu, Jian, Yang, Zheming, He, Liang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
di: Liang, Zichen, et al.
Pubblicazione: (2025)
di: Liang, Zichen, et al.
Pubblicazione: (2025)
From Atoms to Chains: Divergence-Guided Reasoning Curriculum for Unlabeled LLM Domain Adaptation
di: Wang, Yongqi, et al.
Pubblicazione: (2026)
di: Wang, Yongqi, et al.
Pubblicazione: (2026)
Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
di: Lu, Haolang, et al.
Pubblicazione: (2025)
di: Lu, Haolang, et al.
Pubblicazione: (2025)
Comment on “Association of Gastrointestinal Symptoms With Severity and Progression of Cognitive Impairment in Parkinson's Disease: A Systematic Review and Meta‐Analysis”
di: Sheng‐Nan Li, et al.
Pubblicazione: (2026)
di: Sheng‐Nan Li, et al.
Pubblicazione: (2026)
EasyMimic: A Low-Cost Framework for Robot Imitation Learning from Human Videos
di: Zhang, Tao, et al.
Pubblicazione: (2026)
di: Zhang, Tao, et al.
Pubblicazione: (2026)
An Area-Efficient 20-100-GHz Phase-Invariant Switch-Type Attenuator Achieving 0.1-dB Tuning Step in 65-nm CMOS
di: Li, Qingbin, et al.
Pubblicazione: (2025)
di: Li, Qingbin, et al.
Pubblicazione: (2025)
Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer
di: Xing, Bohao, et al.
Pubblicazione: (2026)
di: Xing, Bohao, et al.
Pubblicazione: (2026)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
di: Yin, Jianghao, et al.
Pubblicazione: (2026)
di: Yin, Jianghao, et al.
Pubblicazione: (2026)
Boosting Conversational Question Answering with Fine-Grained Retrieval-Augmentation and Self-Check
di: Ye, Linhao, et al.
Pubblicazione: (2024)
di: Ye, Linhao, et al.
Pubblicazione: (2024)
How Likely Do LLMs with CoT Mimic Human Reasoning?
di: Bao, Guangsheng, et al.
Pubblicazione: (2024)
di: Bao, Guangsheng, et al.
Pubblicazione: (2024)
Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
di: Zhou, Guanghao, et al.
Pubblicazione: (2025)
di: Zhou, Guanghao, et al.
Pubblicazione: (2025)
Model Privacy: A Unified Framework for Understanding Model Stealing Attacks and Defenses
di: Wang, Ganghua, et al.
Pubblicazione: (2025)
di: Wang, Ganghua, et al.
Pubblicazione: (2025)
CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention
di: Li, Qingbin, et al.
Pubblicazione: (2025)
di: Li, Qingbin, et al.
Pubblicazione: (2025)
Design, Synthesis and Activity Evaluation of Methylene‐H4MPT Mimics
di: Yutian Wang, et al.
Pubblicazione: (2025)
di: Yutian Wang, et al.
Pubblicazione: (2025)
MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance Detection
di: Wang, Bingbing, et al.
Pubblicazione: (2025)
di: Wang, Bingbing, et al.
Pubblicazione: (2025)
A Stepwise Distillation Learning Strategy for Non-differentiable Visual Programming Frameworks on Visual Reasoning Tasks
di: Wan, Wentao, et al.
Pubblicazione: (2023)
di: Wan, Wentao, et al.
Pubblicazione: (2023)
MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training
di: Li, Haoyun, et al.
Pubblicazione: (2025)
di: Li, Haoyun, et al.
Pubblicazione: (2025)
Foundation Model for Skeleton-Based Human Action Understanding
di: Wang, Hongsong, et al.
Pubblicazione: (2025)
di: Wang, Hongsong, et al.
Pubblicazione: (2025)
Autonomous Imagination: Closed-Loop Decomposition of Visual-to-Textual Conversion in Visual Reasoning for Multimodal Large Language Models
di: Liu, Jingming, et al.
Pubblicazione: (2024)
di: Liu, Jingming, et al.
Pubblicazione: (2024)
From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning
di: Wang, Shaojie, et al.
Pubblicazione: (2026)
di: Wang, Shaojie, et al.
Pubblicazione: (2026)
Region-aware Image-based Human Action Retrieval with Transformers
di: Wang, Hongsong, et al.
Pubblicazione: (2024)
di: Wang, Hongsong, et al.
Pubblicazione: (2024)
MECO: A Multimodal Dataset for Emotion and Cognitive Understanding in Older Adults
di: Chen, Hongbin, et al.
Pubblicazione: (2026)
di: Chen, Hongbin, et al.
Pubblicazione: (2026)
VisTR: Visualizations as Representations for Time-series Table Reasoning
di: Hao, Jianing, et al.
Pubblicazione: (2024)
di: Hao, Jianing, et al.
Pubblicazione: (2024)
Unifying Language-Action Understanding and Generation for Autonomous Driving
di: Wang, Xinyang, et al.
Pubblicazione: (2026)
di: Wang, Xinyang, et al.
Pubblicazione: (2026)
LSTC-MDA: A Unified Framework for Long-Short Term Temporal Convolution and Mixed Data Augmentation in Skeleton-Based Action Recognition
di: Ding, Feng, et al.
Pubblicazione: (2025)
di: Ding, Feng, et al.
Pubblicazione: (2025)
AdaCodec: A Predictive Visual Code for Video MLLMs
di: Hou, Haowen, et al.
Pubblicazione: (2026)
di: Hou, Haowen, et al.
Pubblicazione: (2026)
ChartAgent: A Chart Understanding Framework with Tool Integrated Reasoning
di: Wang, Boran, et al.
Pubblicazione: (2025)
di: Wang, Boran, et al.
Pubblicazione: (2025)
AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance
di: He, Lixuan, et al.
Pubblicazione: (2025)
di: He, Lixuan, et al.
Pubblicazione: (2025)
Evaluating Accounting Reasoning Capabilities of Large Language Models
di: Zhou, Jie, et al.
Pubblicazione: (2026)
di: Zhou, Jie, et al.
Pubblicazione: (2026)
CAMS: A CityGPT-Powered Agentic Framework for Urban Human Mobility Simulation
di: Du, Yuwei, et al.
Pubblicazione: (2025)
di: Du, Yuwei, et al.
Pubblicazione: (2025)
SCALE: Semantic- and Confidence-Aware Conditional Variational Autoencoder for Zero-shot Skeleton-based Action Recognition
di: Oraki, Soroush, et al.
Pubblicazione: (2026)
di: Oraki, Soroush, et al.
Pubblicazione: (2026)
Global Meta‐Analysis of Individual and Combined Nitrogen Inhibitors: Enhancing Plant Productivity and Reducing Environmental Losses
di: Wenyu Wang, et al.
Pubblicazione: (2024)
di: Wenyu Wang, et al.
Pubblicazione: (2024)
Brainformer: Mimic Human Visual Brain Functions to Machine Vision Models via fMRI
di: Nguyen, Xuan-Bac, et al.
Pubblicazione: (2023)
di: Nguyen, Xuan-Bac, et al.
Pubblicazione: (2023)
Point-Supervised Skeleton-Based Human Action Segmentation
di: Wang, Hongsong, et al.
Pubblicazione: (2026)
di: Wang, Hongsong, et al.
Pubblicazione: (2026)
Action Draft and Verify: A Self-Verifying Framework for Vision-Language-Action Model
di: Zhao, Chen, et al.
Pubblicazione: (2026)
di: Zhao, Chen, et al.
Pubblicazione: (2026)
Protective and Restorative Effects of Biophilic Design in High School Indoor Environments on Stress and Cognitive Function
di: Li Mengqi, et al.
Pubblicazione: (2025)
di: Li Mengqi, et al.
Pubblicazione: (2025)
Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
di: Song, Python, et al.
Pubblicazione: (2025)
di: Song, Python, et al.
Pubblicazione: (2025)
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
di: Zhang, Chi, et al.
Pubblicazione: (2025)
di: Zhang, Chi, et al.
Pubblicazione: (2025)
Hg-I2P: Bridging Modalities for Generalizable Image-to-Point-Cloud Registration via Heterogeneous Graphs
di: An, Pei, et al.
Pubblicazione: (2026)
di: An, Pei, et al.
Pubblicazione: (2026)
VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to Rank
di: Wu, Tianhe, et al.
Pubblicazione: (2025)
di: Wu, Tianhe, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
di: Liang, Zichen, et al.
Pubblicazione: (2025) -
From Atoms to Chains: Divergence-Guided Reasoning Curriculum for Unlabeled LLM Domain Adaptation
di: Wang, Yongqi, et al.
Pubblicazione: (2026) -
Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
di: Lu, Haolang, et al.
Pubblicazione: (2025) -
Comment on “Association of Gastrointestinal Symptoms With Severity and Progression of Cognitive Impairment in Parkinson's Disease: A Systematic Review and Meta‐Analysis”
di: Sheng‐Nan Li, et al.
Pubblicazione: (2026) -
EasyMimic: A Low-Cost Framework for Robot Imitation Learning from Human Videos
di: Zhang, Tao, et al.
Pubblicazione: (2026)