LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Xingxuan, Ren, Gang, Yu, Han, Yuan, Hao, Wang, Hui, Li, Jiansheng, Wu, Jiayun, Mo, Lang, Mao, Li, Hao, Mingchao, Dai, Ningbo, Xu, Renzhe, Li, Shuyang, Zhang, Tianyang, He, Yue, Wang, Yuanrui, Zhang, Yunjia, Xu, Zijing, Li, Dongzhe, Gao, Fang, Zou, Hao, Liu, Jiandong, Liu, Jiashuo, Xu, Jiawei, Cheng, Kaijie, Li, Kehan, Zhou, Linjun, Li, Qing, Fan, Shaohua, Lin, Xiaoyu, Han, Xinyan, Li, Xuanyue, Lu, Yan, Xue, Yuan, Jiang, Yuanyuan, Wang, Zimu, Wang, Zhenlei, Cui, Peng
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908636278685696
author Zhang, Xingxuan
Ren, Gang
Yu, Han
Yuan, Hao
Wang, Hui
Li, Jiansheng
Wu, Jiayun
Mo, Lang
Mao, Li
Hao, Mingchao
Dai, Ningbo
Xu, Renzhe
Li, Shuyang
Zhang, Tianyang
He, Yue
Wang, Yuanrui
Zhang, Yunjia
Xu, Zijing
Li, Dongzhe
Gao, Fang
Zou, Hao
Liu, Jiandong
Liu, Jiashuo
Xu, Jiawei
Cheng, Kaijie
Li, Kehan
Zhou, Linjun
Li, Qing
Fan, Shaohua
Lin, Xiaoyu
Han, Xinyan
Li, Xuanyue
Lu, Yan
Xue, Yuan
Jiang, Yuanyuan
Wang, Zimu
Wang, Zhenlei
Cui, Peng
author_facet Zhang, Xingxuan
Ren, Gang
Yu, Han
Yuan, Hao
Wang, Hui
Li, Jiansheng
Wu, Jiayun
Mo, Lang
Mao, Li
Hao, Mingchao
Dai, Ningbo
Xu, Renzhe
Li, Shuyang
Zhang, Tianyang
He, Yue
Wang, Yuanrui
Zhang, Yunjia
Xu, Zijing
Li, Dongzhe
Gao, Fang
Zou, Hao
Liu, Jiandong
Liu, Jiashuo
Xu, Jiawei
Cheng, Kaijie
Li, Kehan
Zhou, Linjun
Li, Qing
Fan, Shaohua
Lin, Xiaoyu
Han, Xinyan
Li, Xuanyue
Lu, Yan
Xue, Yuan
Jiang, Yuanyuan
Wang, Zimu
Wang, Zhenlei
Cui, Peng
contents We argue that progress toward general intelligence requires complementary foundation models grounded in language, the physical world, and structured data. This report presents LimiX-16M and LimiX-2M, two instantiations of our large structured-data models (LDMs). Both models treat structured data as a joint distribution over variables and missingness, thus capable of addressing a wide range of tabular tasks through query-based conditional prediction via a single model. They are pretrained using masked joint-distribution modeling with an episodic, context-conditional objective, supporting rapid, training-free adaptation at inference. We evaluate LimiX models across 11 large structured-data benchmarks with broad regimes of sample size, feature dimensionality, class number, categorical-to-numerical feature ratio, missingness, and sample-to-feature ratios. LimiX-16M consistently surpasses strong baselines, as shown in Figure 1 and Figure 2. The superiority holds across a wide range of tasks, such as classification, regression, missing value imputation, and data generation, often by substantial margins, while avoiding task-specific architectures or bespoke training per task. Notably, LimiX-2M delivers strong results under tight compute and memory budgets. We also present the first scaling law study for LDMs, revealing how data and model scaling jointly influence downstream performance and offering quantitative guidance for tabular foundation modeling. All LimiX models are publicly accessible under Apache 2.0.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03505
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence
Zhang, Xingxuan
Ren, Gang
Yu, Han
Yuan, Hao
Wang, Hui
Li, Jiansheng
Wu, Jiayun
Mo, Lang
Mao, Li
Hao, Mingchao
Dai, Ningbo
Xu, Renzhe
Li, Shuyang
Zhang, Tianyang
He, Yue
Wang, Yuanrui
Zhang, Yunjia
Xu, Zijing
Li, Dongzhe
Gao, Fang
Zou, Hao
Liu, Jiandong
Liu, Jiashuo
Xu, Jiawei
Cheng, Kaijie
Li, Kehan
Zhou, Linjun
Li, Qing
Fan, Shaohua
Lin, Xiaoyu
Han, Xinyan
Li, Xuanyue
Lu, Yan
Xue, Yuan
Jiang, Yuanyuan
Wang, Zimu
Wang, Zhenlei
Cui, Peng
Machine Learning
Artificial Intelligence
Computation and Language
We argue that progress toward general intelligence requires complementary foundation models grounded in language, the physical world, and structured data. This report presents LimiX-16M and LimiX-2M, two instantiations of our large structured-data models (LDMs). Both models treat structured data as a joint distribution over variables and missingness, thus capable of addressing a wide range of tabular tasks through query-based conditional prediction via a single model. They are pretrained using masked joint-distribution modeling with an episodic, context-conditional objective, supporting rapid, training-free adaptation at inference. We evaluate LimiX models across 11 large structured-data benchmarks with broad regimes of sample size, feature dimensionality, class number, categorical-to-numerical feature ratio, missingness, and sample-to-feature ratios. LimiX-16M consistently surpasses strong baselines, as shown in Figure 1 and Figure 2. The superiority holds across a wide range of tasks, such as classification, regression, missing value imputation, and data generation, often by substantial margins, while avoiding task-specific architectures or bespoke training per task. Notably, LimiX-2M delivers strong results under tight compute and memory budgets. We also present the first scaling law study for LDMs, revealing how data and model scaling jointly influence downstream performance and offering quantitative guidance for tabular foundation modeling. All LimiX models are publicly accessible under Apache 2.0.
title LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.03505