NVIDIA Kumo Tabular Targets a New Accuracy-Efficiency Frontier for Tabular Prediction
NVIDIA Kumo Tabular Targets a New Accuracy Efficiency Frontier for Tabular Prediction Published September 29, 2026 Highlights NVIDIA Kumo Tabular, part of the NVIDIA Kumo Struct...
By Hardware Team
NVIDIA Kumo Tabular Targets a New Accuracy-Efficiency Frontier for Tabular Prediction
Published September 29, 2026
Highlights
NVIDIA Kumo Tabular, part of the NVIDIA Kumo Structured model collection, is an open foundation model for tabular data. Available on Hugging Face, it predicts labels for new rows from a table of labeled examples in a single forward pass. The model supports classification and regression without task-specific training, tuning, or feature engineering.
Kumo Tabular was pretrained entirely on artificial data. It is available in three sizes, ranging from 28 million to 215 million parameters, runs through NVIDIA's open-source structured-data-models library, and is released under the OpenMDW-1.1 license for commercial use.
The model ranks first on four benchmarks: TabArena, BeyondArena, TALENT, and ScoringBench.
- Model code: https://github.com/NVIDIA/structured-data-models
- Model weights: https://huggingface.co/nvidia/Kumo-Tabular
The Shift to Tabular Foundation Models
Tabular data supports many enterprise machine learning workloads. Customer records, transactions, sensor logs, claims, and orders are commonly stored in tables, where models are used to predict outcomes such as churn, default, demand, and price.
For about two decades, gradient-boosted trees have handled many of these tasks effectively. However, the surrounding workflow has remained largely unchanged. Each new prediction problem generally requires collecting labels, engineering features, searching hyperparameters, validating results, and deploying a model that learns the task from scratch.
Large language models introduced another approach. With in-context learning, a pretrained model can receive examples in a prompt and solve a task without changing its weights. The same principle can be applied to tables: a model pretrained on many tables can use labeled rows as context and predict labels for new rows directly.
NVIDIA Kumo Tabular is an open foundation model for tabular classification and regression. Given a table containing labeled rows and rows requiring predictions, it returns class probabilities or numerical predictions in a single forward pass.
How Kumo Tabular Works
Kumo Tabular is a Transformer designed around table structure. Its architecture uses column, row, and in-context attention, drawing on techniques introduced in TabICL and TabPFN.
The model must address three questions:
- What does each value mean within its column?
- How do the columns in a row interact?
- How are labeled context rows related to query rows with unknown labels?
Cell Embedding
Groups of cells are converted into tokens. Numerical and categorical values pass through Fourier features, including sines and cosines of learned frequencies, with separate weights for each value type. Missing values require no imputation and receive special handling. Each token in the context also receives a label embedding.
Row Embedding
Rows are converted into embeddings through repeated alternation between two attention mechanisms.
Column attention examines a single column and learns how a value relates to that column's distribution. For example, it can determine whether 42 is typical or extreme. This mechanism uses induced self-attention, so its cost grows linearly with the number of rows.
Row attention examines the tokens within one row and learns how features interact. Rotary positions distinguish the columns. Four learnable [CLS] tokens are added to each row and provide the final row representation. After this compression, the cost of the final stage no longer depends on the number of columns.
In-Context Learning
A final Transformer operates on the row embeddings. Context rows attend to one another, while query rows attend only to context rows. Each prediction therefore depends on the context and the individual query row, rather than on which other rows are evaluated at the same time.
Because the context does not attend to query rows, its keys and values can be computed once and reused for later predictions. Query rows use Test-GQA, which reduces the cache read by each prediction.
For classification, a prediction head produces class probabilities. For regression, it produces 999 quantiles, which are used to derive a point prediction and an uncertainty estimate.
Length-Aware Attention Temperature
Softmax attention can become more diffuse as the number of keys increases. Attention that remains focused across a few hundred rows may spread across tens of thousands of rows, particularly when an inference table is much larger than the tables seen during training.
Kumo Tabular addresses this by scaling each query with a temperature that grows logarithmically with the number of keys. The coefficient is learned separately for each attention head. This is intended to keep attention focused as tables become longer or wider.
How Kumo Tabular Was Built
Kumo Tabular was pretrained entirely on artificial tables. Each training table is sampled from a Structural Causal Model (SCM).
The generator first samples a configuration for the table, including its size, task, mechanisms, and missingness patterns. A random causal graph then connects hidden variables. Nodes are evaluated from root to leaf using randomly selected functions, including linear maps, small neural networks, trees, and Gaussian processes.
Some nodes become numerical or categorical columns, one becomes the target, and the remaining nodes stay hidden, representing unmeasured causes that can exist in real datasets. Post-processing correlates groups of columns, clips outliers, and introduces missing values. A rapid tree-ensemble check removes tables without a learnable signal.
Because the generator is procedural rather than a trained model, it can produce an ongoing supply of tables with different graphs and mechanisms.
The generated tables include several characteristics of real-world data. Values can be missing in different patterns, features can be coarsened so that duplicate rows have different labels, categorical columns can contain many levels, and regression targets can have heavy-tailed distributions. Training on these examples allows the model to handle such conditions without data cleanup.
For each artificial table, the model receives most rows with their labels as context and predicts labels for the remaining rows. Classification uses cross-entropy loss, while regression uses quantile loss. Classification and regression are trained as separate models.
As with TabICLv2, training takes place in three stages:
- The first and longest stage uses tables with 1,024 rows and up to 100 columns.
- The second stage varies the context from 400 to 10,240 rows.
- The third stage extends the context to 60,000 rows, still with up to 100 columns.
Kumo Tabular-Small, Kumo Tabular-Medium, and Kumo Tabular-Large saw approximately 35 million, 71 million, and 137 million artificial tables, respectively.
NVIDIA stated that the training recipe and artificial data generators would be released later.
Performance
NVIDIA evaluated all three Kumo Tabular sizes with default settings against the full TabArena leaderboard. The evaluation included tuned gradient-boosted trees, AutoGluon, and recent tabular foundation models.
Kumo Tabular ranked first overall with an ELO of 1950 and ran 17 times faster than LimiX-2 under a uniform evaluation setup using a single RTX 6000 Pro. Across all three model sizes, it established the reported state of the art on the accuracy-efficiency Pareto front.
Additional evaluations covered BeyondArena, TALENT, and ScoringBench.
- On BeyondArena, Kumo Tabular achieved an ELO of 1418 and an Improvability score of 7.78%, placing first on the leaderboard.
- On TALENT, it achieved the top overall ranking across classification accuracy, classification log-loss, and regression RMSE. Its average ranks were 6.67, 3.98, and 4.22, respectively.
- On ScoringBench, which evaluates predictive distributions, Kumo Tabular-Large and Kumo Tabular-Medium ranked first and second by average rank.
Limitations
Kumo Tabular supports numerical and categorical columns. Text, images, and timestamps can be converted into features using built-in preprocessing recipes.
A single forward pass supports up to 10 classes. The library extends this to larger numbers of classes through error-correcting output codes.
Accuracy can decline when tables extend well beyond the training ranges or when query rows come from a different distribution than the context rows. As with other predictive models, accuracy and calibration should be evaluated on held-out data before deployment.
Demo
Kumo Tabular runs through NVIDIA's GPU-native structured-data-models library. The library downloads model weights from the Hub on first use and provides the preprocessing, ensembling, and many-class handling used in the evaluations.
The following example shows the basic flow from a pandas.DataFrame to a prediction:
import sdm # structured-data-models
## Tensorize tabular data:
table = sdm.TableTensor.from_pandas(pd.load_csv(...), device="cuda")
na_mask = table["target"].isnan()
model = sdm.models.KumoTabular(device="cuda")
pred = model(
# In-context examples (features/targets):
x_context=table[~na_mask].drop_columns("target"),
y_context=table[~na_mask, "target"],
# Prediction examples (features):
x_query=table[na_mask].drop_column("target"),
)
Availability
Kumo Tabular is released under the OpenMDW License Agreement, version 1.1.
- Model code: https://github.com/NVIDIA/structured-data-models
- Model weights: https://huggingface.co/nvidia/Kumo-Tabular
NVIDIA's release materials state that developers should assess whether the model meets the requirements of their industry and use case, including potential misuse. Reports concerning model quality, risk, security vulnerabilities, or NVIDIA AI issues can be submitted through the project's GitHub issue tracker.
Acknowledgements
NVIDIA credited David Holzmüller with contributing ideas and ablations to Kumo Tabular, and Vignesh Kothapalli with assisting on the project during his internship.