GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | Xia, Renqiu, Li, Mingsheng, Ye, Hancheng, Wu, Wenjie, Zhou, Hongbin, Yuan, Jiakang, Peng, Tianshuo, Cai, Xinyu, Yan, Xiangchao, Wang, Bin, He, Conghui, Shi, Botian, Chen, Tao, Yan, Junchi, Zhang, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy
by: Ye, Hancheng, et al.
Published: (2024)
by: Ye, Hancheng, et al.
Published: (2024)
SPOT: Scalable 3D Pre-training via Occupancy Prediction for Learning Transferable 3D Representations
by: Yan, Xiangchao, et al.
Published: (2023)
by: Yan, Xiangchao, et al.
Published: (2023)
StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding
by: Xia, Renqiu, et al.
Published: (2023)
by: Xia, Renqiu, et al.
Published: (2023)
Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving
by: Liu, Qi, et al.
Published: (2025)
by: Liu, Qi, et al.
Published: (2025)
TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
by: Fu, Daocheng, et al.
Published: (2025)
by: Fu, Daocheng, et al.
Published: (2025)
ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
by: Xia, Renqiu, et al.
Published: (2024)
by: Xia, Renqiu, et al.
Published: (2024)
ReSimAD: Zero-Shot 3D Domain Transfer for Autonomous Driving with Source Reconstruction and Target Simulation
by: Zhang, Bo, et al.
Published: (2023)
by: Zhang, Bo, et al.
Published: (2023)
SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing
by: Yan, Xiangchao, et al.
Published: (2025)
by: Yan, Xiangchao, et al.
Published: (2025)
GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
by: Feng, Yuan, et al.
Published: (2025)
by: Feng, Yuan, et al.
Published: (2025)
Chimera: Improving Generalist Model with Domain-Specific Experts
by: Peng, Tianshuo, et al.
Published: (2024)
by: Peng, Tianshuo, et al.
Published: (2024)
On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets
by: Liao, Ning, et al.
Published: (2023)
by: Liao, Ning, et al.
Published: (2023)
DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models
by: Xia, Renqiu, et al.
Published: (2024)
by: Xia, Renqiu, et al.
Published: (2024)
Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback
by: Yuan, Jiakang, et al.
Published: (2025)
by: Yuan, Jiakang, et al.
Published: (2025)
Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views
by: Zhang, Xiangdong, et al.
Published: (2025)
by: Zhang, Xiangdong, et al.
Published: (2025)
Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression
by: Ye, Hancheng, et al.
Published: (2024)
by: Ye, Hancheng, et al.
Published: (2024)
NITP: Next Implicit Token Prediction for LLM Pre-training
by: Zhang, Xiangdong, et al.
Published: (2026)
by: Zhang, Xiangdong, et al.
Published: (2026)
DriveVGGT: Calibration-Constrained Visual Geometry Transformers for Multi-Camera Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2025)
by: Jia, Xiaosong, et al.
Published: (2025)
UniQA: Unified Vision-Language Pre-training for Image Quality and Aesthetic Assessment
by: Zhou, Hantao, et al.
Published: (2024)
by: Zhou, Hantao, et al.
Published: (2024)
Learning Adaptive and Temporally Causal Video Tokenization in a 1D Latent Space
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
Centroid-centered Modeling for Efficient Vision Transformer Pre-training
by: Yan, Xin, et al.
Published: (2023)
by: Yan, Xin, et al.
Published: (2023)
Image Over Text: Transforming Formula Recognition Evaluation with Character Detection Matching
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
Anatomical Structure-Guided Medical Vision-Language Pre-training
by: Li, Qingqiu, et al.
Published: (2024)
by: Li, Qingqiu, et al.
Published: (2024)
Incorporating Pre-trained Diffusion Models in Solving the Schrödinger Bridge Problem
by: Tang, Zhicong, et al.
Published: (2025)
by: Tang, Zhicong, et al.
Published: (2025)
Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?
by: Wen, Zichen, et al.
Published: (2025)
by: Wen, Zichen, et al.
Published: (2025)
OmniCaptioner: One Captioner to Rule Them All
by: Lu, Yiting, et al.
Published: (2025)
by: Lu, Yiting, et al.
Published: (2025)
MathLearner: A Large Language Model Agent Framework for Learning to Solve Mathematical Problems
by: Xie, Wenbei, et al.
Published: (2024)
by: Xie, Wenbei, et al.
Published: (2024)
GASP: Unifying Geometric and Semantic Self-Supervised Pre-training for Autonomous Driving
by: Ljungbergh, William, et al.
Published: (2025)
by: Ljungbergh, William, et al.
Published: (2025)
Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training
by: Bawazir, Ameera, et al.
Published: (2024)
by: Bawazir, Ameera, et al.
Published: (2024)
Incorporating Graph Attention Mechanism into Geometric Problem Solving Based on Deep Reinforcement Learning
by: Zhong, Xiuqin, et al.
Published: (2024)
by: Zhong, Xiuqin, et al.
Published: (2024)
3D Scene Graph Guided Vision-Language Pre-training
by: Liu, Hao, et al.
Published: (2024)
by: Liu, Hao, et al.
Published: (2024)
Unified Batch Normalization: Identifying and Alleviating the Feature Condensation in Batch Normalization and a Unified Framework
by: Wang, Shaobo, et al.
Published: (2023)
by: Wang, Shaobo, et al.
Published: (2023)
FlatFusion: Delving into Details of Sparse Transformer-based Camera-LiDAR Fusion for Autonomous Driving
by: Zhu, Yutao, et al.
Published: (2024)
by: Zhu, Yutao, et al.
Published: (2024)
UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
by: Chen, Lan, et al.
Published: (2025)
by: Chen, Lan, et al.
Published: (2025)
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets
by: Wang, Pengyu, et al.
Published: (2025)
by: Wang, Pengyu, et al.
Published: (2025)
Formula-Supervised Visual-Geometric Pre-training
by: Yamada, Ryosuke, et al.
Published: (2024)
by: Yamada, Ryosuke, et al.
Published: (2024)
SuPreME: A Supervised Pre-training Framework for Multimodal ECG Representation Learning
by: Cai, Mingsheng, et al.
Published: (2025)
by: Cai, Mingsheng, et al.
Published: (2025)
VLP: A Survey on Vision-Language Pre-training
by: Chen, Feilong, et al.
Published: (2022)
by: Chen, Feilong, et al.
Published: (2022)
DriveTransformer: Unified Transformer for Scalable End-to-End Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2025)
by: Jia, Xiaosong, et al.
Published: (2025)
AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D Vision
by: Cheng, Xiaoya, et al.
Published: (2026)
by: Cheng, Xiaoya, et al.
Published: (2026)
Similar Items
-
Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy
by: Ye, Hancheng, et al.
Published: (2024) -
SPOT: Scalable 3D Pre-training via Occupancy Prediction for Learning Transferable 3D Representations
by: Yan, Xiangchao, et al.
Published: (2023) -
StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding
by: Xia, Renqiu, et al.
Published: (2023) -
Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving
by: Liu, Qi, et al.
Published: (2025) -
TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
by: Fu, Daocheng, et al.
Published: (2025)