Synthesizing Multimodal Geometry Datasets from Scratch and Enabling Visual Alignment via Plotting Code
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Haobo, Bai, Tianyi, Chen, Chen, Zhang, Jiajun, Zeng, Bohan, Zhang, Wentao, Yuan, Binhang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
TAG: Thinking with Action Unit Grounding for Facial Expression Recognition
von: Lin, Haobo, et al.
Veröffentlicht: (2026)
von: Lin, Haobo, et al.
Veröffentlicht: (2026)
CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
von: Duan, Chengqi, et al.
Veröffentlicht: (2025)
von: Duan, Chengqi, et al.
Veröffentlicht: (2025)
Towards Visual Query Localization in the 3D World
von: Peng, Liang, et al.
Veröffentlicht: (2026)
von: Peng, Liang, et al.
Veröffentlicht: (2026)
Hierarchical Context Alignment with Disentangled Geometric and Temporal Modeling for Semantic Occupancy Prediction
von: Li, Bohan, et al.
Veröffentlicht: (2024)
von: Li, Bohan, et al.
Veröffentlicht: (2024)
WideRange4D: Enabling High-Quality 4D Reconstruction with Wide-Range Movements and Scenes
von: Yang, Ling, et al.
Veröffentlicht: (2025)
von: Yang, Ling, et al.
Veröffentlicht: (2025)
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
von: E, Shaojun, et al.
Veröffentlicht: (2025)
von: E, Shaojun, et al.
Veröffentlicht: (2025)
VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
von: Zhao, Xuanle, et al.
Veröffentlicht: (2025)
von: Zhao, Xuanle, et al.
Veröffentlicht: (2025)
Widget2Code: From Visual Widgets to UI Code via Multimodal LLMs
von: Zhang, Houston H., et al.
Veröffentlicht: (2025)
von: Zhang, Houston H., et al.
Veröffentlicht: (2025)
Defending Multimodal Backdoored Models by Repulsive Visual Prompt Tuning
von: Zhang, Zhifang, et al.
Veröffentlicht: (2024)
von: Zhang, Zhifang, et al.
Veröffentlicht: (2024)
SPA++: Generalized Graph Spectral Alignment for Versatile Domain Adaptation
von: Xiao, Zhiqing, et al.
Veröffentlicht: (2025)
von: Xiao, Zhiqing, et al.
Veröffentlicht: (2025)
Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge
von: Liang, Hao, et al.
Veröffentlicht: (2025)
von: Liang, Hao, et al.
Veröffentlicht: (2025)
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge
von: Zhao, Yaqi, et al.
Veröffentlicht: (2024)
von: Zhao, Yaqi, et al.
Veröffentlicht: (2024)
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos
von: Feng, Hengyi, et al.
Veröffentlicht: (2026)
von: Feng, Hengyi, et al.
Veröffentlicht: (2026)
See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMs
von: Zhang, Yongchang, et al.
Veröffentlicht: (2026)
von: Zhang, Yongchang, et al.
Veröffentlicht: (2026)
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
A Visual Question Answering Method for SAR Ship: Breaking the Requirement for Multimodal Dataset Construction and Model Fine-Tuning
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
Any2AnyTryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing Tasks
von: Guo, Hailong, et al.
Veröffentlicht: (2025)
von: Guo, Hailong, et al.
Veröffentlicht: (2025)
QuadBox: Accelerating 3D Gaussian Splatting with Geometry-Aware Boxes
von: Li, Xinze, et al.
Veröffentlicht: (2026)
von: Li, Xinze, et al.
Veröffentlicht: (2026)
UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations
von: Zhao, Yaqi, et al.
Veröffentlicht: (2026)
von: Zhao, Yaqi, et al.
Veröffentlicht: (2026)
SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
von: Yin, Yuanyang, et al.
Veröffentlicht: (2024)
von: Yin, Yuanyang, et al.
Veröffentlicht: (2024)
ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch
von: Liu, Zheng, et al.
Veröffentlicht: (2026)
von: Liu, Zheng, et al.
Veröffentlicht: (2026)
Thinking with Spatial Code for Physical-World Video Reasoning
von: Chen, Jieneng, et al.
Veröffentlicht: (2026)
von: Chen, Jieneng, et al.
Veröffentlicht: (2026)
When Pedestrian Detection Meets Multi-Modal Learning: Generalist Model and Benchmark Dataset
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaxin, et al.
Veröffentlicht: (2024)
GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes
von: Wang, Di, et al.
Veröffentlicht: (2025)
von: Wang, Di, et al.
Veröffentlicht: (2025)
Enhancing Visual Programming for Visual Reasoning via Probabilistic Graphs
von: Wan, Wentao, et al.
Veröffentlicht: (2025)
von: Wan, Wentao, et al.
Veröffentlicht: (2025)
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
von: Wang, Qiuchen, et al.
Veröffentlicht: (2026)
von: Wang, Qiuchen, et al.
Veröffentlicht: (2026)
VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation
von: Yuan, Jiayi, et al.
Veröffentlicht: (2026)
von: Yuan, Jiayi, et al.
Veröffentlicht: (2026)
A Survey of Multimodal Large Language Model from A Data-centric Perspective
von: Bai, Tianyi, et al.
Veröffentlicht: (2024)
von: Bai, Tianyi, et al.
Veröffentlicht: (2024)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
von: Wu, Size, et al.
Veröffentlicht: (2025)
von: Wu, Size, et al.
Veröffentlicht: (2025)
Beyond Modality Collapse: Representations Blending for Multimodal Dataset Distillation
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
Trans4D: Realistic Geometry-Aware Transition for Compositional Text-to-4D Synthesis
von: Zeng, Bohan, et al.
Veröffentlicht: (2024)
von: Zeng, Bohan, et al.
Veröffentlicht: (2024)
Transformer-Based Visual Segmentation: A Survey
von: Li, Xiangtai, et al.
Veröffentlicht: (2023)
von: Li, Xiangtai, et al.
Veröffentlicht: (2023)
POINTS-Seeker: Towards Training a Multimodal Agentic Search Model from Scratch
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
Semantic Score Distillation Sampling for Compositional Text-to-3D Generation
von: Yang, Ling, et al.
Veröffentlicht: (2024)
von: Yang, Ling, et al.
Veröffentlicht: (2024)
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers
von: Bing, Zhaodong, et al.
Veröffentlicht: (2025)
von: Bing, Zhaodong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
von: Bai, Tianyi, et al.
Veröffentlicht: (2025) -
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
von: Bai, Tianyi, et al.
Veröffentlicht: (2025) -
TAG: Thinking with Action Unit Grounding for Facial Expression Recognition
von: Lin, Haobo, et al.
Veröffentlicht: (2026) -
CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
von: Duan, Chengqi, et al.
Veröffentlicht: (2025) -
Towards Visual Query Localization in the 3D World
von: Peng, Liang, et al.
Veröffentlicht: (2026)