E2E-GMNER: End-to-End Generative Grounded Multimodal Named Entity Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Meng, Ning, Jinzhong, Wu, Xiaolong, Lin, Hongfei, Zhang, Yijia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
Grounding Language Models for Visual Entity Recognition
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation
von: You, Ling, et al.
Veröffentlicht: (2025)
von: You, Ling, et al.
Veröffentlicht: (2025)
Multi-Grained Query-Guided Set Prediction Network for Grounded Multimodal Named Entity Recognition
von: Tang, Jielong, et al.
Veröffentlicht: (2024)
von: Tang, Jielong, et al.
Veröffentlicht: (2024)
A Proposal-Free Query-Guided Network for Grounded Multimodal Named Entity Recognition
von: Li, Hongbing, et al.
Veröffentlicht: (2026)
von: Li, Hongbing, et al.
Veröffentlicht: (2026)
Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
von: Si, Chenglei, et al.
Veröffentlicht: (2024)
End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
von: Liang, Jianxin, et al.
Veröffentlicht: (2024)
von: Liang, Jianxin, et al.
Veröffentlicht: (2024)
Leveraging Entity Information for Cross-Modality Correlation Learning: The Entity-Guided Multimodal Summarization
von: Zhang, Yanghai, et al.
Veröffentlicht: (2024)
von: Zhang, Yanghai, et al.
Veröffentlicht: (2024)
E2E-MFD: Towards End-to-End Synchronous Multimodal Fusion Detection
von: Zhang, Jiaqing, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaqing, et al.
Veröffentlicht: (2024)
TextFormer: A Query-based End-to-End Text Spotter with Mixed Supervision
von: Zhai, Yukun, et al.
Veröffentlicht: (2023)
von: Zhai, Yukun, et al.
Veröffentlicht: (2023)
An Effective End-to-End Solution for Multimodal Action Recognition
von: Wang, Songping, et al.
Veröffentlicht: (2025)
von: Wang, Songping, et al.
Veröffentlicht: (2025)
Uncovering Entity Identity Confusion in Multimodal Knowledge Editing
von: Wu, Shu, et al.
Veröffentlicht: (2026)
von: Wu, Shu, et al.
Veröffentlicht: (2026)
End-to-End Agentic RAG System Training for Traceable Diagnostic Reasoning
von: Zheng, Qiaoyu, et al.
Veröffentlicht: (2025)
von: Zheng, Qiaoyu, et al.
Veröffentlicht: (2025)
Efficient End-to-End Visual Document Understanding with Rationale Distillation
von: Zhu, Wang, et al.
Veröffentlicht: (2023)
von: Zhu, Wang, et al.
Veröffentlicht: (2023)
EMMA: End-to-End Multimodal Model for Autonomous Driving
von: Hwang, Jyh-Jing, et al.
Veröffentlicht: (2024)
von: Hwang, Jyh-Jing, et al.
Veröffentlicht: (2024)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
von: Ahn, Young Jin, et al.
Veröffentlicht: (2024)
von: Ahn, Young Jin, et al.
Veröffentlicht: (2024)
Pose-Based Sign Language Spotting via an End-to-End Encoder Architecture
von: Johnny, Samuel Ebimobowei, et al.
Veröffentlicht: (2025)
von: Johnny, Samuel Ebimobowei, et al.
Veröffentlicht: (2025)
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
von: Zhang, Yaolun, et al.
Veröffentlicht: (2026)
von: Zhang, Yaolun, et al.
Veröffentlicht: (2026)
EAMA : Entity-Aware Multimodal Alignment Based Approach for News Image Captioning
von: Zhang, Junzhe, et al.
Veröffentlicht: (2024)
von: Zhang, Junzhe, et al.
Veröffentlicht: (2024)
ICR-Drive: Instruction Counterfactual Robustness for End-to-End Language-Driven Autonomous Driving
von: Hamid, Kaiser, et al.
Veröffentlicht: (2026)
von: Hamid, Kaiser, et al.
Veröffentlicht: (2026)
Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding
von: Ye, Junyi, et al.
Veröffentlicht: (2024)
von: Ye, Junyi, et al.
Veröffentlicht: (2024)
GIT-CXR: End-to-End Transformer for Chest X-Ray Report Generation
von: Sîrbu, Iustin, et al.
Veröffentlicht: (2025)
von: Sîrbu, Iustin, et al.
Veröffentlicht: (2025)
A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
End-to-End Chess Recognition
von: Masouris, Athanasios, et al.
Veröffentlicht: (2023)
von: Masouris, Athanasios, et al.
Veröffentlicht: (2023)
VP-MEL: Visual Prompts Guided Multimodal Entity Linking
von: Mi, Hongze, et al.
Veröffentlicht: (2024)
von: Mi, Hongze, et al.
Veröffentlicht: (2024)
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
von: Xie, Xudong, et al.
Veröffentlicht: (2024)
von: Xie, Xudong, et al.
Veröffentlicht: (2024)
E2E-GNet: An End-to-End Skeleton-based Geometric Deep Neural Network for Human Motion Recognition
von: Olaoluwa, Mubarak, et al.
Veröffentlicht: (2026)
von: Olaoluwa, Mubarak, et al.
Veröffentlicht: (2026)
Checkup2Action: A Multimodal Clinical Check-up Report Dataset for Patient-Oriented Action Card Generation
von: Xiang, Sike, et al.
Veröffentlicht: (2026)
von: Xiang, Sike, et al.
Veröffentlicht: (2026)
FEA-SLT: A Gloss-Free End-to-End Framework for Facial-Expression-Aware Sign Language Translation
von: Tu, Guobin, et al.
Veröffentlicht: (2026)
von: Tu, Guobin, et al.
Veröffentlicht: (2026)
Towards Fully Decoupled End-to-End Person Search
von: Zhang, Pengcheng, et al.
Veröffentlicht: (2023)
von: Zhang, Pengcheng, et al.
Veröffentlicht: (2023)
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
von: Goetting, Dylan, et al.
Veröffentlicht: (2024)
von: Goetting, Dylan, et al.
Veröffentlicht: (2024)
From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving
von: Zheng, Huan, et al.
Veröffentlicht: (2025)
von: Zheng, Huan, et al.
Veröffentlicht: (2025)
Generalizable Entity Grounding via Assistance of Large Language Model
von: Qi, Lu, et al.
Veröffentlicht: (2024)
von: Qi, Lu, et al.
Veröffentlicht: (2024)
Light Up the Shadows: Enhance Long-Tailed Entity Grounding with Concept-Guided Vision-Language Models
von: Zhang, Yikai, et al.
Veröffentlicht: (2024)
von: Zhang, Yikai, et al.
Veröffentlicht: (2024)
Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
GoalFlow: Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Driving
von: Xing, Zebin, et al.
Veröffentlicht: (2025)
von: Xing, Zebin, et al.
Veröffentlicht: (2025)
End-to-End 3D Spatiotemporal Perception with Multimodal Fusion and V2X Collaboration
von: Yang, Zhenwei, et al.
Veröffentlicht: (2025)
von: Yang, Zhenwei, et al.
Veröffentlicht: (2025)
RIG: Synergizing Reasoning and Imagination in End-to-End Generalist Policy
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2025)
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2025)
MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures
von: Strohmeyer, Tim, et al.
Veröffentlicht: (2026)
von: Strohmeyer, Tim, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition
von: Li, Jinyuan, et al.
Veröffentlicht: (2024) -
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
von: Li, Jinyuan, et al.
Veröffentlicht: (2024) -
Grounding Language Models for Visual Entity Recognition
von: Xiao, Zilin, et al.
Veröffentlicht: (2024) -
TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation
von: You, Ling, et al.
Veröffentlicht: (2025) -
Multi-Grained Query-Guided Set Prediction Network for Grounded Multimodal Named Entity Recognition
von: Tang, Jielong, et al.
Veröffentlicht: (2024)