NVIDIA releases Kumo Tabular, an open foundation model for tabular data that tops four benchmarks

NVIDIA released Kumo Tabular, an open-source model that predicts tabular outcomes without training. It ranks first on four benchmarks and runs 17 times faster than competitor LimiX-2.

NVIDIA releases Kumo Tabular, an open foundation model for tabular data that tops four benchmarks

NVIDIA released Kumo Tabular today, an open-source foundation model for tabular data that predicts outcomes for new rows in a single forward pass without training, tuning, or feature engineering. The model ranks first on four major benchmarks, including TabArena and BeyondArena, while running 17 times faster than competitor LimiX-2 on standard hardware. For industries relying on structured data, this shift eliminates the weeks-long cycle of model development that has defined enterprise machine learning for the past two decades.

Breaking the gradient-boosted tree cycle

Tabular data underpins most enterprise machine learning tasks, from predicting customer churn and credit default in finance to forecasting demand in retail and assessing risk in insurance. For years, the industry standard has been gradient-boosted trees. These models perform well but require a heavy operational lift: every new question demands collecting labels, engineering features, searching for hyperparameters, and deploying a model that starts from zero knowledge.

Kumo Tabular changes this workflow by applying in-context learning to tables. Much like a large language model can solve a new task by reading examples in a prompt, Kumo reads a labeled table as its context and directly predicts labels for new rows. The model was pretrained exclusively on artificial data, making it agnostic to specific industry datasets while retaining a general understanding of tabular structure. It comes in three sizes, ranging from 28 million to 215 million parameters, and is released under the OpenMDW-1.1 license, which permits commercial use.

The architecture is built to handle the specific mechanics of a spreadsheet. It uses a Transformer with column, row, and in-context attention. Cell embedding processes numerical and categorical values through Fourier features, treating missing values without requiring imputation. Row embedding then compresses these values using alternating column and row attention, allowing the model to understand both individual column distributions and complex feature interactions.

A final stage handles in-context learning, where the model relates labeled context rows to unlabeled query rows. This design ensures that predictions for one row do not depend on which other rows are being scored simultaneously, maintaining consistency regardless of batch size. The model also employs length-aware attention temperature, which adjusts sharpness as the table grows, preventing accuracy drops when processing datasets larger than those seen during training.

Performance and practical limitations

In evaluations against tuned gradient-boosted trees and AutoGluon, Kumo Tabular achieved an ELO of 1950 on the TabArena leaderboard. It also established a new state-of-the-art on the accuracy-efficiency Pareto front, meaning it delivers higher accuracy with significantly lower computational cost. On TALENT, it took the top overall ranking across classification accuracy, log-loss, and regression RMSE. The model supports up to 10 classes in a single pass, with the library extending this to any number of classes using error-correcting output codes.

Despite these gains, Kumo Tabular has constraints. It currently works only on numerical and categorical columns. Text, images, or timestamps must be converted into features via preprocessing steps before input. Additionally, accuracy may degrade if the data distribution for new queries differs significantly from the context rows, a common issue in shifting market conditions. NVIDIA advises validating accuracy and calibration on held-out data before deployment.

The release aligns with broader efforts to make AI Machine Learning Training more accessible to data scientists who traditionally spend most of their time on feature engineering rather than model architecture. By automating the structural understanding of tables, models like Kumo allow practitioners to focus on interpretation and business logic.

Why this matters for finance and healthcare

For professionals in education, finance, healthcare, insurance, and research, Kumo Tabular offers a faster path to predictive insight. Insurance actuaries can test risk models on new claim datasets instantly, bypassing the retraining pipeline. Healthcare researchers can analyze patient records for outcome prediction without needing to engineer features for every new variable set. In finance, traders and analysts can model price movements or default risks on fresh data feeds with minimal latency. The technology reduces the barrier to entry for advanced predictive analytics, allowing non-specialists to derive value from structured data while retaining the speed and accuracy previously reserved for custom-built solutions.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)