Lifestyle

NVIDIA releases Kumo Tabular, an open tabular foundation model it says ranks first on four benchmarks

On September 29, 2026, NVIDIA released Kumo Tabular, a free-to-use AI model that makes predictions from spreadsheet-style data without being trained for each task. It could make prediction work faster for companies and researchers. Here is how NVIDIA says it works, the results NVIDIA reports and the limits NVIDIA itself names.

About 7 min read

NVIDIA releases Kumo Tabular, an open tabular foundation model it says ranks first on four benchmarks
Image: Mokaair (Original editorial artwork)

What happened

On September 29, 2026, NVIDIA released NVIDIA Kumo Tabular on the Hugging Face blog. It belongs to the NVIDIA Kumo Structured model family and is positioned as an open foundation model for tabular data. It handles two kinds of prediction: classification (choosing a category, such as whether a customer will leave) and regression (predicting a number, such as a price). According to NVIDIA, the model was pretrained only on synthetic data and comes in three sizes, from 28 million to 215 million parameters (the internal values a model learns). It is released under the OpenMDW-1.1 license, which permits commercial use. The code is on GitHub and the model weights are on Hugging Face. The model runs through structured-data-models, a GPU-native library that NVIDIA has just released.

Tabular data is information arranged in rows and columns, like a spreadsheet. Examples include customer records, transactions, sensor logs and orders. NVIDIA notes that for two decades this kind of prediction has mostly relied on gradient-boosted trees, a widely used method that combines many simple decision rules. With that approach, every new question means collecting labels (known answers) and doing feature engineering (hand-crafting useful input columns). It also means searching hyperparameters (settings chosen before training), then validating and deploying a new model. Kumo Tabular instead borrows "in-context learning" from large language models. It reads a labeled table as a set of examples and predicts the unknown rows directly, in a single forward pass (one run through the model, with no learning step).

NVIDIA releases Kumo Tabular, an open tabular foundation model it says ranks first on four benchmarks
Mokaair editorial verification flow · Image: Mokaair (Original editorial artwork)
Read the full description

Sources are collected, independently checked, then reviewed by Jev.

How it differs from the traditional approach

The right column reflects NVIDIA's own description of Kumo Tabular
AspectTraditional gradient-boosted tree workflowKumo Tabular (as described by NVIDIA)
New tasksA model must be trained from scratch for each problemReads labeled rows and predicts in a single forward pass
Feature engineering and tuningRequires designing features and searching hyperparametersNVIDIA says no training, tuning or feature engineering is needed
Missing valuesUsually need handling or imputation (filling in gaps)NVIDIA says no imputation is needed; the model treats missing values specially
Regression outputUsually a single valueOutputs 999 quantiles, giving a point prediction and an estimate of uncertainty
Supported column typesDepends on the model and preprocessingOnly numeric and categorical columns; other types must be converted by built-in preprocessing

How the model works and was trained

NVIDIA says Kumo Tabular is a Transformer, the type of neural network behind large language models, designed around the structure of a table. It uses the column, row and in-context attention mechanisms introduced in TabICL and TabPFN. In plain terms, it works in three steps. First, it works out what each value means within its column. Next, it looks at how the columns in a row relate to each other. Finally, it compares the rows with known answers to the rows it needs to predict.

NVIDIA says each training table was sampled from a structural causal model (SCM), a simulated web of cause-and-effect relationships. NVIDIA deliberately built in flaws common in real data. These include values missing in several patterns, duplicate rows with conflicting labels, and columns with many categories. They also include regression targets with occasional extreme values. Training ran in three stages. The first stage used tables of 1,024 rows and up to 100 columns. The second expanded the examples from 400 to 10,240 rows, and the third extended them to 60,000 rows. NVIDIA says the Small, Medium and Large versions saw roughly 35 million, 71 million and 137 million synthetic tables respectively. It also says the training method and data generators will be released later.

Benchmark results published by NVIDIA

Benchmarks are shared tests used to compare models. Several of them report an ELO score, a rating built from head-to-head comparisons, similar to chess ratings.

  • TabArena: NVIDIA says it ranks first overall with an ELO of 1950. It also says the model runs faster than LimiX-2 when both are tested on a single RTX 6000 Pro GPU with the same settings.
  • BeyondArena: NVIDIA says it ranks first with an ELO of 1418 and an Improvability score of 7.78%.
  • TALENT: NVIDIA says it achieves the best overall ranking on classification accuracy, classification log loss and regression RMSE, with average ranks of 6.67, 3.98 and 4.22 respectively.
  • ScoringBench: NVIDIA says the Large and Medium versions rank first and second by average rank.

What it means for general readers

Most people will never use this kind of model directly. But many everyday decisions rely on predictions from tables, such as whether a customer will leave, the risk of a loan default, expected demand and pricing. If NVIDIA's claims hold up, building such prediction models could take companies and researchers less effort and time. That could make it easier for small teams to try. However, faster predictions are not necessarily correct predictions. Where people's rights and interests are at stake, validation, calibration and human oversight remain important. This article is for information only and does not constitute a recommendation to buy any product or hardware.

Frequently asked questions

What is Kumo Tabular?

According to NVIDIA, it is an open foundation model for classification and regression on tabular data. It is part of the NVIDIA Kumo Structured model family and was released on Hugging Face on September 29, 2026.

Does it really need no training?

NVIDIA says users supply rows with known answers and the rows to be predicted. The model then returns predictions in a single forward pass, with no task-specific training, tuning or feature engineering. The model itself was pretrained in advance on synthetic tables.

Can it be used commercially?

NVIDIA says the model is released under the OpenMDW-1.1 license and can be used commercially. Read the license terms yourself before using it.

Should I trust the first-place rankings?

The rankings come from NVIDIA's own announcement. NVIDIA itself recommends checking accuracy and calibration on your own held-out data before deployment.

Can it handle text or images?

NVIDIA says the model directly handles only numeric and categorical columns. Text, images or timestamps must first be converted into features through the library's built-in preprocessing.

Browse the latest news in this topic

Latest travel guides

Sources

Lifestyle