InfoDet: A Dataset for Infographic Element Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Jiangning, Zhou, Yuxing, Wang, Zheng, Yao, Juntao, Gu, Yima, Yuan, Yuhui, Liu, Shixia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts
by: Xie, Tianchi, et al.
Published: (2025)
by: Xie, Tianchi, et al.
Published: (2025)
InfoAffect: Affective Annotations of Infographics in Information Spread
by: Fu, Zihang, et al.
Published: (2025)
by: Fu, Zihang, et al.
Published: (2025)
Structural-Entropy-Based Sample Selection for Efficient and Effective Learning
by: Xie, Tianchi, et al.
Published: (2024)
by: Xie, Tianchi, et al.
Published: (2024)
EnsemHalDet: Robust VLM Hallucination Detection via Ensemble of Internal State Detectors
by: Miyazato, Ryuhei, et al.
Published: (2026)
by: Miyazato, Ryuhei, et al.
Published: (2026)
InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models
by: Liu, Xinxin, et al.
Published: (2026)
by: Liu, Xinxin, et al.
Published: (2026)
LMM-Det: Make Large Multimodal Models Excel in Object Detection
by: Li, Jincheng, et al.
Published: (2025)
by: Li, Jincheng, et al.
Published: (2025)
OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
by: Zhao, Tiancheng, et al.
Published: (2022)
by: Zhao, Tiancheng, et al.
Published: (2022)
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
by: Shen, Yijun, et al.
Published: (2025)
by: Shen, Yijun, et al.
Published: (2025)
BizGen: Advancing Article-level Visual Text Rendering for Infographics Generation
by: Peng, Yuyang, et al.
Published: (2025)
by: Peng, Yuyang, et al.
Published: (2025)
Plain-Det: A Plain Multi-Dataset Object Detector
by: Shi, Cheng, et al.
Published: (2024)
by: Shi, Cheng, et al.
Published: (2024)
MemeMind: A Large-Scale Multimodal Dataset with Chain-of-Thought Reasoning for Harmful Meme Detection
by: Gu, Hexiang, et al.
Published: (2025)
by: Gu, Hexiang, et al.
Published: (2025)
CitDet: A Benchmark Dataset for Citrus Fruit Detection
by: James, Jordan A., et al.
Published: (2023)
by: James, Jordan A., et al.
Published: (2023)
VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection
by: Yao, Jianhang, et al.
Published: (2025)
by: Yao, Jianhang, et al.
Published: (2025)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
by: Zang, Yuan, et al.
Published: (2025)
by: Zang, Yuan, et al.
Published: (2025)
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
by: Lozano, Alejandro, et al.
Published: (2025)
by: Lozano, Alejandro, et al.
Published: (2025)
PedDet: Adaptive Spectral Optimization for Multimodal Pedestrian Detection
by: Zhao, Rui, et al.
Published: (2025)
by: Zhao, Rui, et al.
Published: (2025)
Support or Refute: Analyzing the Stance of Evidence to Detect Out-of-Context Mis- and Disinformation
by: Yuan, Xin, et al.
Published: (2023)
by: Yuan, Xin, et al.
Published: (2023)
DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video
by: Hou, Jiawei, et al.
Published: (2025)
by: Hou, Jiawei, et al.
Published: (2025)
Gaussian-Det: Learning Closed-Surface Gaussians for 3D Object Detection
by: Yan, Hongru, et al.
Published: (2024)
by: Yan, Hongru, et al.
Published: (2024)
RemDet: Rethinking Efficient Model Design for UAV Object Detection
by: Li, Chen, et al.
Published: (2024)
by: Li, Chen, et al.
Published: (2024)
GUIDE: A Guideline-Guided Dataset for Instructional Video Comprehension
by: Liang, Jiafeng, et al.
Published: (2024)
by: Liang, Jiafeng, et al.
Published: (2024)
Multimodal Human-AI Synergy for Medical Imaging Quality Control: A Hybrid Intelligence Framework with Adaptive Dataset Curation and Closed-Loop Evaluation
by: Qin, Zhi, et al.
Published: (2025)
by: Qin, Zhi, et al.
Published: (2025)
RadarGaussianDet3D: Gaussian Representation-based Real-time 3D Object Detection with 4D Automotive Radars
by: Xiong, Weiyi, et al.
Published: (2025)
by: Xiong, Weiyi, et al.
Published: (2025)
RGBX-DiffusionDet: A Framework for Multi-Modal RGB-X Object Detection Using DiffusionDet
by: Orfaig, Eliraz, et al.
Published: (2025)
by: Orfaig, Eliraz, et al.
Published: (2025)
DetCLIPv3: Towards Versatile Generative Open-vocabulary Object Detection
by: Yao, Lewei, et al.
Published: (2024)
by: Yao, Lewei, et al.
Published: (2024)
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
by: Zhou, Shijie, et al.
Published: (2024)
by: Zhou, Shijie, et al.
Published: (2024)
UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection
by: Zhang, Yanran, et al.
Published: (2026)
by: Zhang, Yanran, et al.
Published: (2026)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
by: Liu, Zheng, et al.
Published: (2024)
by: Liu, Zheng, et al.
Published: (2024)
PerioDet: Large-Scale Panoramic Radiograph Benchmark for Clinical-Oriented Apical Periodontitis Detection
by: Fang, Xiaocheng, et al.
Published: (2025)
by: Fang, Xiaocheng, et al.
Published: (2025)
LayoutLLM: Layout Instruction Tuning with Large Language Models for Document Understanding
by: Luo, Chuwei, et al.
Published: (2024)
by: Luo, Chuwei, et al.
Published: (2024)
GVSynergy-Det: Synergistic Gaussian-Voxel Representations for Multi-View 3D Object Detection
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool
by: Wang, Yan, et al.
Published: (2024)
by: Wang, Yan, et al.
Published: (2024)
MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models
by: Zhou, Keyan, et al.
Published: (2025)
by: Zhou, Keyan, et al.
Published: (2025)
MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning
by: Luo, Yuxuan, et al.
Published: (2025)
by: Luo, Yuxuan, et al.
Published: (2025)
MatchDet: A Collaborative Framework for Image Matching and Object Detection
by: Lai, Jinxiang, et al.
Published: (2023)
by: Lai, Jinxiang, et al.
Published: (2023)
ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
by: Hasegawa, Kimihiro, et al.
Published: (2025)
by: Hasegawa, Kimihiro, et al.
Published: (2025)
MutDet: Mutually Optimizing Pre-training for Remote Sensing Object Detection
by: Huang, Ziyue, et al.
Published: (2024)
by: Huang, Ziyue, et al.
Published: (2024)
VisText-Mosquito: A Unified Multimodal Dataset for Visual Detection, Segmentation, and Textual Explanation on Mosquito Breeding Sites
by: Islam, Md. Adnanul, et al.
Published: (2025)
by: Islam, Md. Adnanul, et al.
Published: (2025)
BIMCV-R: A Landmark Dataset for 3D CT Text-Image Retrieval
by: Chen, Yinda, et al.
Published: (2024)
by: Chen, Yinda, et al.
Published: (2024)
Similar Items
-
ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation
by: Li, Zhen, et al.
Published: (2025) -
InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts
by: Xie, Tianchi, et al.
Published: (2025) -
InfoAffect: Affective Annotations of Infographics in Information Spread
by: Fu, Zihang, et al.
Published: (2025) -
Structural-Entropy-Based Sample Selection for Efficient and Effective Learning
by: Xie, Tianchi, et al.
Published: (2024) -
EnsemHalDet: Robust VLM Hallucination Detection via Ensemble of Internal State Detectors
by: Miyazato, Ryuhei, et al.
Published: (2026)