E2E-GMNER: End-to-End Generative Grounded Multimodal Named Entity Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Meng, Ning, Jinzhong, Wu, Xiaolong, Lin, Hongfei, Zhang, Yijia
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911605589016576
author Zhang, Meng
Ning, Jinzhong
Wu, Xiaolong
Lin, Hongfei
Zhang, Yijia
author_facet Zhang, Meng
Ning, Jinzhong
Wu, Xiaolong
Lin, Hongfei
Zhang, Yijia
contents Grounded Multimodal Named Entity Recognition (GMNER) aims to jointly identify named entity mentions in text, predict their semantic types, and ground each entity to a corresponding visual region in an associated image. Existing approaches predominantly adopt pipeline-based architectures that decouple textual entity recognition and visual grounding, leading to error accumulation and suboptimal joint optimization. In this paper, we propose E2E-GMNER, a fully end-to-end generative framework that unifies entity recognition, semantic typing, visual grounding, and implicit knowledge reasoning within a single multimodal large language model. We formulate GMNER as an instruction-tuned conditional generation task and incorporate chain-of-thought reasoning to enable the model to adaptively determine when visual evidence or background knowledge is informative, reducing reliance on noisy cues. To further address the instability of generative bounding box prediction, we introduce Gaussian Risk-Aware Box Perturbation (GRBP), which replaces hard box supervision with probabilistically perturbed soft targets to improve robustness against annotation noise and discretization errors. Extensive experiments on the Twitter-GMNER and Twitter-FMNERG benchmarks demonstrate that E2E-GMNER achieves highly competitive performance compared with state of the art methods, validating the effectiveness of unified end-to-end optimization and noise-aware grounding supervision. Code is available at:https://github.com/Finch-coder/E2E-GMNER
format Preprint
id arxiv_https___arxiv_org_abs_2604_17319
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle E2E-GMNER: End-to-End Generative Grounded Multimodal Named Entity Recognition
Zhang, Meng
Ning, Jinzhong
Wu, Xiaolong
Lin, Hongfei
Zhang, Yijia
Computer Vision and Pattern Recognition
Computation and Language
Grounded Multimodal Named Entity Recognition (GMNER) aims to jointly identify named entity mentions in text, predict their semantic types, and ground each entity to a corresponding visual region in an associated image. Existing approaches predominantly adopt pipeline-based architectures that decouple textual entity recognition and visual grounding, leading to error accumulation and suboptimal joint optimization. In this paper, we propose E2E-GMNER, a fully end-to-end generative framework that unifies entity recognition, semantic typing, visual grounding, and implicit knowledge reasoning within a single multimodal large language model. We formulate GMNER as an instruction-tuned conditional generation task and incorporate chain-of-thought reasoning to enable the model to adaptively determine when visual evidence or background knowledge is informative, reducing reliance on noisy cues. To further address the instability of generative bounding box prediction, we introduce Gaussian Risk-Aware Box Perturbation (GRBP), which replaces hard box supervision with probabilistically perturbed soft targets to improve robustness against annotation noise and discretization errors. Extensive experiments on the Twitter-GMNER and Twitter-FMNERG benchmarks demonstrate that E2E-GMNER achieves highly competitive performance compared with state of the art methods, validating the effectiveness of unified end-to-end optimization and noise-aware grounding supervision. Code is available at:https://github.com/Finch-coder/E2E-GMNER
title E2E-GMNER: End-to-End Generative Grounded Multimodal Named Entity Recognition
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2604.17319