PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Zilu, Lin, Hongbin, Yuan, Zhihao, Zheng, Chaoda, Qiu, Pengshuo, Jiang, Dongzhi, Zhang, Renrui, Feng, Chun-Mei, Li, Zhen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
Towards Flexible 3D Perception: Object-Centric Occupancy Completion Augments 3D Object Detection
by: Zheng, Chaoda, et al.
Published: (2024)
by: Zheng, Chaoda, et al.
Published: (2024)
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
by: Zhang, Renrui, et al.
Published: (2024)
by: Zhang, Renrui, et al.
Published: (2024)
DriveFlow: Rectified Flow Adaptation for Robust 3D Object Detection in Autonomous Driving
by: Lin, Hongbin, et al.
Published: (2025)
by: Lin, Hongbin, et al.
Published: (2025)
Scenarios and Approaches for Situated Natural Language Explanations
by: Qiu, Pengshuo, et al.
Published: (2024)
by: Qiu, Pengshuo, et al.
Published: (2024)
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
by: Yuan, Zhihao, et al.
Published: (2025)
by: Yuan, Zhihao, et al.
Published: (2025)
Frequency-domain general synthetic iterative scheme for efficient simulation of oscillatory rarefied gas flows
by: Li, Pengshuo, et al.
Published: (2026)
by: Li, Pengshuo, et al.
Published: (2026)
MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
by: Zhang, Renrui, et al.
Published: (2024)
by: Zhang, Renrui, et al.
Published: (2024)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
by: Chen, Xinyan, et al.
Published: (2025)
by: Chen, Xinyan, et al.
Published: (2025)
Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding
by: Yuan, Zhihao, et al.
Published: (2023)
by: Yuan, Zhihao, et al.
Published: (2023)
A Possible Mechanism to Explain the Prograde Equatorial Jet of a Jupiter-like Gaseous Giant
by: Lian, Yuchen, et al.
Published: (2026)
by: Lian, Yuchen, et al.
Published: (2026)
TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
by: Fu, Daocheng, et al.
Published: (2025)
by: Fu, Daocheng, et al.
Published: (2025)
Stepwise Stiffening Chromophore Strategy Realizes a Series of Ultralong Blue Room‐Temperature Phosphorescent Materials
by: Zhihao Guan, et al.
Published: (2024)
by: Zhihao Guan, et al.
Published: (2024)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
Benchmarking the Robustness of LiDAR Semantic Segmentation Models
by: Yan, Xu, et al.
Published: (2023)
by: Yan, Xu, et al.
Published: (2023)
Empowering Large Language Models with 3D Situation Awareness
by: Yuan, Zhihao, et al.
Published: (2025)
by: Yuan, Zhihao, et al.
Published: (2025)
Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understanding
by: Chen, Junhan, et al.
Published: (2026)
by: Chen, Junhan, et al.
Published: (2026)
Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation
by: He, Jun, et al.
Published: (2026)
by: He, Jun, et al.
Published: (2026)
Understanding and treating intrauterine adhesions: Insights into molecular mechanisms and innovative therapies
by: Dongzhi Gou, et al.
Published: (2025)
by: Dongzhi Gou, et al.
Published: (2025)
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
by: Chen, Zhen, et al.
Published: (2025)
by: Chen, Zhen, et al.
Published: (2025)
Notes on Laver Tables
by: Qi, Renrui
Published: (2025)
by: Qi, Renrui
Published: (2025)
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
by: Jiang, Dongzhi, et al.
Published: (2025)
by: Jiang, Dongzhi, et al.
Published: (2025)
Role-Augmented Intent-Driven Generative Search Engine Optimization
by: Chen, Xiaolu, et al.
Published: (2025)
by: Chen, Xiaolu, et al.
Published: (2025)
SPRINT: an integrated software for automated, AI-guided dispensing and precision, high-throughput single-cell sorting
by: Ye, Zilu
Published: (2026)
by: Ye, Zilu
Published: (2026)
FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model
by: Lin, Hongbin, et al.
Published: (2025)
by: Lin, Hongbin, et al.
Published: (2025)
Language-Assisted 3D Scene Understanding
by: Wu, Yanmin, et al.
Published: (2023)
by: Wu, Yanmin, et al.
Published: (2023)
Fractal Correction Engine Applied to the Casimir Effect: Pi-Curvature Analysis of Quantum Vacuum Fluctuations
by: McEvoy, Adam L
Published: (2026)
by: McEvoy, Adam L
Published: (2026)
Self-assembly of supermolecular species directed by hydrogen bonding and aromatic Pi-Pi stacking interactions
by: X.H. Li
Published: (2008)
by: X.H. Li
Published: (2008)
Saten: Sparse Augmented Tensor Networks for Post-Training Compression of Large Language Models
by: Solgi, Ryan, et al.
Published: (2025)
by: Solgi, Ryan, et al.
Published: (2025)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Empowering sustainability practices through energy transition: The role of digital economy and technological innovation among BRICS economies
by: Muhammad Awais Baloch, et al.
Published: (2024)
by: Muhammad Awais Baloch, et al.
Published: (2024)
DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generation
by: Lin, Hongbin, et al.
Published: (2025)
by: Lin, Hongbin, et al.
Published: (2025)
Enhancing Label-efficient Medical Image Segmentation with Text-guided Diffusion Models
by: Feng, Chun-Mei
Published: (2024)
by: Feng, Chun-Mei
Published: (2024)
MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
by: Li, Yizhi, et al.
Published: (2023)
by: Li, Yizhi, et al.
Published: (2023)
Scalable In-Context Learning on Tabular Data via Retrieval-Augmented Large Language Models
by: Wen, Xumeng, et al.
Published: (2025)
by: Wen, Xumeng, et al.
Published: (2025)
PiPa++: Towards Unification of Domain Adaptive Semantic Segmentation via Self-supervised Learning
by: Chen, Mu, et al.
Published: (2024)
by: Chen, Mu, et al.
Published: (2024)
LargePiG: Your Large Language Model is Secretly a Pointer Generator
by: Sun, Zhongxiang, et al.
Published: (2024)
by: Sun, Zhongxiang, et al.
Published: (2024)
SA-GS: Scale-Adaptive Gaussian Splatting for Training-Free Anti-Aliasing
by: Song, Xiaowei, et al.
Published: (2024)
by: Song, Xiaowei, et al.
Published: (2024)
Mitigating Hallucinated Translations in Large Language Models with Hallucination-focused Preference Optimization
by: Tang, Zilu, et al.
Published: (2025)
by: Tang, Zilu, et al.
Published: (2025)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
by: Jiang, Dongzhi, et al.
Published: (2025)
by: Jiang, Dongzhi, et al.
Published: (2025)
Similar Items
-
MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
by: Jiang, Dongzhi, et al.
Published: (2024) -
Towards Flexible 3D Perception: Object-Centric Occupancy Completion Augments 3D Object Detection
by: Zheng, Chaoda, et al.
Published: (2024) -
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
by: Zhang, Renrui, et al.
Published: (2024) -
DriveFlow: Rectified Flow Adaptation for Robust 3D Object Detection in Autonomous Driving
by: Lin, Hongbin, et al.
Published: (2025) -
Scenarios and Approaches for Situated Natural Language Explanations
by: Qiu, Pengshuo, et al.
Published: (2024)