LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866908636278685696 |
|---|---|
| author | Zhang, Xingxuan Ren, Gang Yu, Han Yuan, Hao Wang, Hui Li, Jiansheng Wu, Jiayun Mo, Lang Mao, Li Hao, Mingchao Dai, Ningbo Xu, Renzhe Li, Shuyang Zhang, Tianyang He, Yue Wang, Yuanrui Zhang, Yunjia Xu, Zijing Li, Dongzhe Gao, Fang Zou, Hao Liu, Jiandong Liu, Jiashuo Xu, Jiawei Cheng, Kaijie Li, Kehan Zhou, Linjun Li, Qing Fan, Shaohua Lin, Xiaoyu Han, Xinyan Li, Xuanyue Lu, Yan Xue, Yuan Jiang, Yuanyuan Wang, Zimu Wang, Zhenlei Cui, Peng |
| author_facet | Zhang, Xingxuan Ren, Gang Yu, Han Yuan, Hao Wang, Hui Li, Jiansheng Wu, Jiayun Mo, Lang Mao, Li Hao, Mingchao Dai, Ningbo Xu, Renzhe Li, Shuyang Zhang, Tianyang He, Yue Wang, Yuanrui Zhang, Yunjia Xu, Zijing Li, Dongzhe Gao, Fang Zou, Hao Liu, Jiandong Liu, Jiashuo Xu, Jiawei Cheng, Kaijie Li, Kehan Zhou, Linjun Li, Qing Fan, Shaohua Lin, Xiaoyu Han, Xinyan Li, Xuanyue Lu, Yan Xue, Yuan Jiang, Yuanyuan Wang, Zimu Wang, Zhenlei Cui, Peng |
| contents | We argue that progress toward general intelligence requires complementary foundation models grounded in language, the physical world, and structured data. This report presents LimiX-16M and LimiX-2M, two instantiations of our large structured-data models (LDMs). Both models treat structured data as a joint distribution over variables and missingness, thus capable of addressing a wide range of tabular tasks through query-based conditional prediction via a single model. They are pretrained using masked joint-distribution modeling with an episodic, context-conditional objective, supporting rapid, training-free adaptation at inference. We evaluate LimiX models across 11 large structured-data benchmarks with broad regimes of sample size, feature dimensionality, class number, categorical-to-numerical feature ratio, missingness, and sample-to-feature ratios. LimiX-16M consistently surpasses strong baselines, as shown in Figure 1 and Figure 2. The superiority holds across a wide range of tasks, such as classification, regression, missing value imputation, and data generation, often by substantial margins, while avoiding task-specific architectures or bespoke training per task. Notably, LimiX-2M delivers strong results under tight compute and memory budgets. We also present the first scaling law study for LDMs, revealing how data and model scaling jointly influence downstream performance and offering quantitative guidance for tabular foundation modeling. All LimiX models are publicly accessible under Apache 2.0. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_03505 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence Zhang, Xingxuan Ren, Gang Yu, Han Yuan, Hao Wang, Hui Li, Jiansheng Wu, Jiayun Mo, Lang Mao, Li Hao, Mingchao Dai, Ningbo Xu, Renzhe Li, Shuyang Zhang, Tianyang He, Yue Wang, Yuanrui Zhang, Yunjia Xu, Zijing Li, Dongzhe Gao, Fang Zou, Hao Liu, Jiandong Liu, Jiashuo Xu, Jiawei Cheng, Kaijie Li, Kehan Zhou, Linjun Li, Qing Fan, Shaohua Lin, Xiaoyu Han, Xinyan Li, Xuanyue Lu, Yan Xue, Yuan Jiang, Yuanyuan Wang, Zimu Wang, Zhenlei Cui, Peng Machine Learning Artificial Intelligence Computation and Language We argue that progress toward general intelligence requires complementary foundation models grounded in language, the physical world, and structured data. This report presents LimiX-16M and LimiX-2M, two instantiations of our large structured-data models (LDMs). Both models treat structured data as a joint distribution over variables and missingness, thus capable of addressing a wide range of tabular tasks through query-based conditional prediction via a single model. They are pretrained using masked joint-distribution modeling with an episodic, context-conditional objective, supporting rapid, training-free adaptation at inference. We evaluate LimiX models across 11 large structured-data benchmarks with broad regimes of sample size, feature dimensionality, class number, categorical-to-numerical feature ratio, missingness, and sample-to-feature ratios. LimiX-16M consistently surpasses strong baselines, as shown in Figure 1 and Figure 2. The superiority holds across a wide range of tasks, such as classification, regression, missing value imputation, and data generation, often by substantial margins, while avoiding task-specific architectures or bespoke training per task. Notably, LimiX-2M delivers strong results under tight compute and memory budgets. We also present the first scaling law study for LDMs, revealing how data and model scaling jointly influence downstream performance and offering quantitative guidance for tabular foundation modeling. All LimiX models are publicly accessible under Apache 2.0. |
| title | LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2509.03505 |