CodePercept: Code-Grounded Visual STEM Perception for MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guan, Tongkun, Yang, Zhibo, Wan, Jianqiang, Yang, Mingkun, Guo, Zhengtao, Hu, Zijian, Luo, Ruilin, Chen, Ruize, Jiang, Songtao, Wang, Peng, Shen, Wei, Lin, Junyang, Yang, Xiaokang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning
von: Luo, Ruilin, et al.
Veröffentlicht: (2026)
von: Luo, Ruilin, et al.
Veröffentlicht: (2026)
PosFormer: Recognizing Complex Handwritten Mathematical Expression with Position Forest Transformer
von: Guan, Tongkun, et al.
Veröffentlicht: (2024)
von: Guan, Tongkun, et al.
Veröffentlicht: (2024)
Bridging Synthetic and Real Worlds for Pre-training Scene Text Detectors
von: Guan, Tongkun, et al.
Veröffentlicht: (2023)
von: Guan, Tongkun, et al.
Veröffentlicht: (2023)
Learning Transferable Temporal Primitives for Video Reasoning via Synthetic Videos
von: Jiang, Songtao, et al.
Veröffentlicht: (2026)
von: Jiang, Songtao, et al.
Veröffentlicht: (2026)
AdaCodec: A Predictive Visual Code for Video MLLMs
von: Hou, Haowen, et al.
Veröffentlicht: (2026)
von: Hou, Haowen, et al.
Veröffentlicht: (2026)
Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
von: Wang, Zining, et al.
Veröffentlicht: (2025)
von: Wang, Zining, et al.
Veröffentlicht: (2025)
A Token-level Text Image Foundation Model for Document Understanding
von: Guan, Tongkun, et al.
Veröffentlicht: (2025)
von: Guan, Tongkun, et al.
Veröffentlicht: (2025)
Linking Perception, Confidence and Accuracy in MLLMs
von: Du, Yuetian, et al.
Veröffentlicht: (2026)
von: Du, Yuetian, et al.
Veröffentlicht: (2026)
Exploring MLLMs Perception of Network Visualization Principles
von: Miller, Jacob, et al.
Veröffentlicht: (2025)
von: Miller, Jacob, et al.
Veröffentlicht: (2025)
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
von: Yang, Zhibo, et al.
Veröffentlicht: (2024)
von: Yang, Zhibo, et al.
Veröffentlicht: (2024)
The Solution for the ICCV 2023 Perception Test Challenge 2023 -- Task 6 -- Grounded videoQA
von: Zhang, Hailiang, et al.
Veröffentlicht: (2024)
von: Zhang, Hailiang, et al.
Veröffentlicht: (2024)
Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models
von: Jiang, Songtao, et al.
Veröffentlicht: (2024)
von: Jiang, Songtao, et al.
Veröffentlicht: (2024)
SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions
von: Tu, Jinzhe, et al.
Veröffentlicht: (2026)
von: Tu, Jinzhe, et al.
Veröffentlicht: (2026)
Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP Supervision
von: Yin, Kangsheng, et al.
Veröffentlicht: (2025)
von: Yin, Kangsheng, et al.
Veröffentlicht: (2025)
Semantic Voting: Execution-Grounded Consensus for LLM Code Generation
von: Jiang, Shan, et al.
Veröffentlicht: (2026)
von: Jiang, Shan, et al.
Veröffentlicht: (2026)
Neural Material Adaptor for Visual Grounding of Intrinsic Dynamics
von: Cao, Junyi, et al.
Veröffentlicht: (2024)
von: Cao, Junyi, et al.
Veröffentlicht: (2024)
Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
von: Zhu, Fangrui, et al.
Veröffentlicht: (2025)
von: Zhu, Fangrui, et al.
Veröffentlicht: (2025)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
HueManity: Probing Fine-Grained Visual Perception in MLLMs
von: Grover, Rynaa, et al.
Veröffentlicht: (2025)
von: Grover, Rynaa, et al.
Veröffentlicht: (2025)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
von: Wang, Ziwei, et al.
Veröffentlicht: (2025)
Shared and Unique Neural Codes for Biological Motion Perception in Humans and Macaque Monkeys
von: Yuhui Cheng, et al.
Veröffentlicht: (2025)
von: Yuhui Cheng, et al.
Veröffentlicht: (2025)
GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning
von: Chen, Guizhen, et al.
Veröffentlicht: (2025)
von: Chen, Guizhen, et al.
Veröffentlicht: (2025)
MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
von: Xu, Yunqiu, et al.
Veröffentlicht: (2024)
von: Xu, Yunqiu, et al.
Veröffentlicht: (2024)
SparseCoop: Cooperative Perception with Kinematic-Grounded Queries
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
"My productivity is boosted, but ..." Demystifying Users' Perception on AI Coding Assistants
von: Lyu, Yunbo, et al.
Veröffentlicht: (2025)
von: Lyu, Yunbo, et al.
Veröffentlicht: (2025)
Integrating Image Perception and Time‐to‐First‐Spike Coding in MoS2 Phototransistors for Spiking Neural Network
von: Xiangwei Su, et al.
Veröffentlicht: (2024)
von: Xiangwei Su, et al.
Veröffentlicht: (2024)
GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
von: Zheng, Guanghao, et al.
Veröffentlicht: (2025)
von: Zheng, Guanghao, et al.
Veröffentlicht: (2025)
GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
von: Li, Rang, et al.
Veröffentlicht: (2025)
von: Li, Rang, et al.
Veröffentlicht: (2025)
An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability
von: Wu, Daiqing, et al.
Veröffentlicht: (2025)
von: Wu, Daiqing, et al.
Veröffentlicht: (2025)
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
von: Zhang, Miaosen, et al.
Veröffentlicht: (2025)
von: Zhang, Miaosen, et al.
Veröffentlicht: (2025)
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
von: Jain, Jitesh, et al.
Veröffentlicht: (2024)
von: Jain, Jitesh, et al.
Veröffentlicht: (2024)
Evaluating and Achieving Controllable Code Completion in Code LLM
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
von: Zhang, Jiajun, et al.
Veröffentlicht: (2026)
Towards Better Correctness and Efficiency in Code Generation
von: Feng, Yunlong, et al.
Veröffentlicht: (2025)
von: Feng, Yunlong, et al.
Veröffentlicht: (2025)
Metal Oxide‐Based Neuromorphic Artificial Visual Perception Devices and Systems for Information Perception, Memory and Processing
von: Fan Yang, et al.
Veröffentlicht: (2025)
von: Fan Yang, et al.
Veröffentlicht: (2025)
Scaling Agentic Verifier for Competitive Coding
von: Ma, Zeyao, et al.
Veröffentlicht: (2026)
von: Ma, Zeyao, et al.
Veröffentlicht: (2026)
Explore the Hallucination on Low-level Perception for MLLMs
von: Sun, Yinan, et al.
Veröffentlicht: (2024)
von: Sun, Yinan, et al.
Veröffentlicht: (2024)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
The Relationship Between Black Girls' Perceptions of Their STEM Teachers and Their STEM Identity
von: Kristian Edosomwan, et al.
Veröffentlicht: (2025)
von: Kristian Edosomwan, et al.
Veröffentlicht: (2025)
GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding
von: Zhu, Jiashun, et al.
Veröffentlicht: (2026)
von: Zhu, Jiashun, et al.
Veröffentlicht: (2026)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
von: Li, Danyang, et al.
Veröffentlicht: (2025)
von: Li, Danyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning
von: Luo, Ruilin, et al.
Veröffentlicht: (2026) -
PosFormer: Recognizing Complex Handwritten Mathematical Expression with Position Forest Transformer
von: Guan, Tongkun, et al.
Veröffentlicht: (2024) -
Bridging Synthetic and Real Worlds for Pre-training Scene Text Detectors
von: Guan, Tongkun, et al.
Veröffentlicht: (2023) -
Learning Transferable Temporal Primitives for Video Reasoning via Synthetic Videos
von: Jiang, Songtao, et al.
Veröffentlicht: (2026) -
AdaCodec: A Predictive Visual Code for Video MLLMs
von: Hou, Haowen, et al.
Veröffentlicht: (2026)