GenExam: A Multidisciplinary Text-to-Image Exam
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Zhaokai, Yin, Penghao, Zhao, Xiangyu, Tian, Changyao, Qiao, Yu, Wang, Wenhai, Dai, Jifeng, Luo, Gen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
di: Luo, Gen, et al.
Pubblicazione: (2025)
di: Luo, Gen, et al.
Pubblicazione: (2025)
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
di: Wang, Zhaokai, et al.
Pubblicazione: (2025)
di: Wang, Zhaokai, et al.
Pubblicazione: (2025)
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
di: Lei, Zhenxin, et al.
Pubblicazione: (2025)
di: Lei, Zhenxin, et al.
Pubblicazione: (2025)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
di: Tian, Changyao, et al.
Pubblicazione: (2024)
di: Tian, Changyao, et al.
Pubblicazione: (2024)
Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
di: Luo, Gen, et al.
Pubblicazione: (2024)
di: Luo, Gen, et al.
Pubblicazione: (2024)
CoMemo: LVLMs Need Image Context with Image Memory
di: Liu, Shi, et al.
Pubblicazione: (2025)
di: Liu, Shi, et al.
Pubblicazione: (2025)
Parameter-Inverted Image Pyramid Networks
di: Zhu, Xizhou, et al.
Pubblicazione: (2024)
di: Zhu, Xizhou, et al.
Pubblicazione: (2024)
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
di: Tian, Changyao, et al.
Pubblicazione: (2025)
di: Tian, Changyao, et al.
Pubblicazione: (2025)
ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework
di: Chen, Guanzhou, et al.
Pubblicazione: (2026)
di: Chen, Guanzhou, et al.
Pubblicazione: (2026)
Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
di: Yang, Chenyu, et al.
Pubblicazione: (2024)
di: Yang, Chenyu, et al.
Pubblicazione: (2024)
Disentanglement and Assessment of Shortcuts in Ophthalmological Retinal Imaging Exams
di: Fernandes, Leonor, et al.
Pubblicazione: (2025)
di: Fernandes, Leonor, et al.
Pubblicazione: (2025)
ADDP: Learning General Representations for Image Recognition and Generation with Alternating Denoising Diffusion Process
di: Tian, Changyao, et al.
Pubblicazione: (2023)
di: Tian, Changyao, et al.
Pubblicazione: (2023)
Learning 1D Causal Visual Representation with De-focus Attention Networks
di: Tao, Chenxin, et al.
Pubblicazione: (2024)
di: Tao, Chenxin, et al.
Pubblicazione: (2024)
Bounding Box Stability against Feature Dropout Reflects Detector Generalization across Environments
di: Yang, Yang, et al.
Pubblicazione: (2024)
di: Yang, Yang, et al.
Pubblicazione: (2024)
EuraGovExam: A Multilingual Multimodal Benchmark from Real-World Civil Service Exams
di: Kim, Jaeseong, et al.
Pubblicazione: (2026)
di: Kim, Jaeseong, et al.
Pubblicazione: (2026)
Point or Line? Using Line-based Representation for Panoptic Symbol Spotting in CAD Drawings
di: Wei, Xingguang, et al.
Pubblicazione: (2025)
di: Wei, Xingguang, et al.
Pubblicazione: (2025)
Demystify Transformers & Convolutions in Modern Image Deep Networks
di: Hu, Xiaowei, et al.
Pubblicazione: (2022)
di: Hu, Xiaowei, et al.
Pubblicazione: (2022)
MathDoc: Benchmarking Structured Extraction and Active Refusal on Noisy Mathematics Exam Papers
di: Zhou, Chenyue, et al.
Pubblicazione: (2026)
di: Zhou, Chenyue, et al.
Pubblicazione: (2026)
DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving
di: Cui, Erfei, et al.
Pubblicazione: (2023)
di: Cui, Erfei, et al.
Pubblicazione: (2023)
Grading Handwritten Engineering Exams with Multimodal Large Language Models
di: Perš, Janez, et al.
Pubblicazione: (2026)
di: Perš, Janez, et al.
Pubblicazione: (2026)
GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing
di: Liu, Mingxin, et al.
Pubblicazione: (2026)
di: Liu, Mingxin, et al.
Pubblicazione: (2026)
Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation
di: Salazar, Israfel, et al.
Pubblicazione: (2025)
di: Salazar, Israfel, et al.
Pubblicazione: (2025)
PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models
di: Yang, Chenyu, et al.
Pubblicazione: (2024)
di: Yang, Chenyu, et al.
Pubblicazione: (2024)
LangBridge: Interpreting Image as a Combination of Language Embeddings
di: Liao, Jiaqi, et al.
Pubblicazione: (2025)
di: Liao, Jiaqi, et al.
Pubblicazione: (2025)
AutoOEP -- A Multi-modal Framework for Online Exam Proctoring
di: Naveen, Aryan Kashyap
Pubblicazione: (2025)
di: Naveen, Aryan Kashyap
Pubblicazione: (2025)
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
di: Duan, Yuchen, et al.
Pubblicazione: (2024)
di: Duan, Yuchen, et al.
Pubblicazione: (2024)
Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
di: Cui, Yiming, et al.
Pubblicazione: (2025)
di: Cui, Yiming, et al.
Pubblicazione: (2025)
ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution
di: Cui, Long, et al.
Pubblicazione: (2025)
di: Cui, Long, et al.
Pubblicazione: (2025)
MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity
di: Liu, Yangzhou, et al.
Pubblicazione: (2024)
di: Liu, Yangzhou, et al.
Pubblicazione: (2024)
The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
di: Wang, Weiyun, et al.
Pubblicazione: (2024)
di: Wang, Weiyun, et al.
Pubblicazione: (2024)
GenAD: Generalized Predictive Model for Autonomous Driving
di: Yang, Jiazhi, et al.
Pubblicazione: (2024)
di: Yang, Jiazhi, et al.
Pubblicazione: (2024)
Team PA-VCG's Solution for Competition on Understanding Chinese College Entrance Exam Papers in ICDAR'25
di: Wu, Wei, et al.
Pubblicazione: (2025)
di: Wu, Wei, et al.
Pubblicazione: (2025)
InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing
di: Tian, Changyao, et al.
Pubblicazione: (2026)
di: Tian, Changyao, et al.
Pubblicazione: (2026)
ACCSAMS: Automatic Conversion of Exam Documents to Accessible Learning Material for Blind and Visually Impaired
di: Wilkening, David, et al.
Pubblicazione: (2024)
di: Wilkening, David, et al.
Pubblicazione: (2024)
Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
di: Xiong, Yuwen, et al.
Pubblicazione: (2024)
di: Xiong, Yuwen, et al.
Pubblicazione: (2024)
Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
di: Li, Yan, et al.
Pubblicazione: (2025)
di: Li, Yan, et al.
Pubblicazione: (2025)
Trust Your Critic: Robust Reward Modeling and Reinforcement Learning for Faithful Image Editing and Generation
di: Zhao, Xiangyu, et al.
Pubblicazione: (2026)
di: Zhao, Xiangyu, et al.
Pubblicazione: (2026)
ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
di: Liao, Jiaqi, et al.
Pubblicazione: (2025)
di: Liao, Jiaqi, et al.
Pubblicazione: (2025)
FairGen: Enhancing Fairness in Text-to-Image Diffusion Models via Self-Discovering Latent Directions
di: Jiang, Yilei, et al.
Pubblicazione: (2024)
di: Jiang, Yilei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
di: Luo, Gen, et al.
Pubblicazione: (2025) -
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
di: Wang, Zhaokai, et al.
Pubblicazione: (2025) -
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
di: Li, Hao, et al.
Pubblicazione: (2024) -
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
di: Lei, Zhenxin, et al.
Pubblicazione: (2025) -
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
di: Tian, Changyao, et al.
Pubblicazione: (2024)