DLRMv4 Brings Sequence-Based Recommendation Training to MLPerf
Recommendation workloads have been part of MLPerf Training through the DLRM benchmark, which measures how long a system needs to train a personalized recommendation model to a d...
By AI Engineering Team
Recommendation workloads have been part of MLPerf Training through the DLRM benchmark, which measures how long a system needs to train a personalized recommendation model to a defined quality target. DLRMv2 [1] recommends content from user behavior and item features, using deep learning to model their interactions for services such as e-commerce platforms and streaming applications.
Recent sequence-modeling advances, including large language models, have changed how hyperscale recommendation systems represent user activity. Instead of compressing behavior into aggregated dense features, many production systems represent recent interactions as sequences of per-event tokens [2][3]. These models continue to improve as their compute budgets and parameter counts increase, while earlier feature-interaction approaches tend to plateau. MLPerf Inference already includes an HSTU-based benchmark for this workload class [4], but MLPerf Training has lacked a comparable training benchmark.
DLRMv4 addresses that gap by replacing the traditional feature-interaction stack with HSTU [2], a transducer that processes user interaction histories as sequences. It also introduces Yambda-5B [5], a multi-behavior dataset of music-listening activity with a production-scale embedding footprint of approximately 560 GB. Using HSTU keeps the training benchmark aligned with the related inference benchmark. DLRMv4 therefore reflects both the computational demands of attention over long user histories and the memory demands of embedding tables that must be distributed across modern accelerators.
Model Architecture
Figure 1: Traditional DLRM architecture (a) compared with HSTU (b)
Traditional DLRM systems use an encode-then-interaction pipeline. Heterogeneous inputs first pass through feature-extraction components: bottom MLPs process numerical features, while embedding lookups process categorical IDs. The resulting representations are sent to a feature-interaction module, such as a cross network or factorization machine, and then to a task head that produces the prediction.
This interaction module does not process the raw behavior stream. Instead, it operates on engineered and pre-aggregated inputs, including decayed counters, ratios, and pooled embeddings.
HSTU removes this separation. Categorical and slowly changing features are organized into a unified chronological sequence of items and actions. A stack of identical transduction layers then processes the sequence end to end. Each layer combines the functions that DLRM handles separately:
- A projection generates the internal representations required by the layer.
- Pointwise attention aggregates information across the sequence using a bias based on positional and temporal distance.
- A gated transformation allows the attention-pooled features to interact directly, replacing both the explicit cross network and the feed-forward block.
Because the same block is repeated throughout the model, its main scaling dimensions are depth, width, and sequence length. The HSTU research reports a power-law relationship between compute and quality across three orders of magnitude, while the encode-then-interaction baseline reaches a flatter scaling regime.
A single training example illustrates the design. The model may score one candidate item for a user on a content platform. The user history consists of interleaved past items and actions, such as engagements, approvals, and skips, truncated to a fixed budget of several thousand events. Each event becomes a token containing the item, the action, and the event time. A negative signal such as a skip therefore remains an explicit event instead of being averaged into a counter.
Slowly changing context, including user identity and precomputed cross features, is compressed into several prefix tokens. The candidate item is appended to the sequence, allowing the model to compare it with the history in one pass. The HSTU stack processes the complete sequence, and the representation at the candidate position is passed to a small prediction head.
Processing history, context, and the candidate jointly in one causal block makes HSTU relevant to production systems and computationally demanding to train. The same configuration underlies the MLPerf DLRMv3 inference benchmark [4].
Dataset
DLRMv4 uses Yambda-5B [5], an open recommendation dataset containing real multi-behavior user interactions released by Yandex Music. The benchmark uses the full 5B-event variant.
| Property | Value |
|---|---|
| Source | Yandex Music public release |
| Interactions | Approximately 4.79B |
| Users | 1M |
| Items | 9.39M |
| Behavior types | 5: listen, like, dislike, unlike, undislike |
| Raw size | Approximately 38 GB for the Yambda-5B variant |
| Per-event size | Approximately 20 bytes |
Table 1: Yambda-5B dataset properties
Yambda is well suited to HSTU because it provides a per-user timeline. Each interaction records the user, item, action type, and timestamp, allowing histories to be reconstructed chronologically and passed to the transducer as sequences. With approximately 4.79 billion interactions, the dataset can sustain a full-scale training workload.
| Criterion | Criteo 1TB (DLRMv2) | Yambda-5B (DLRMv4) |
|---|---|---|
| Real data | Yes | Yes |
| Interaction volume | 4.2B rows | 4.79B rows |
| Per-user history | No, flat features only | Yes, multi-behavior pools |
| Raw on-disk size | Approximately 4 TB, denormalized | Approximately 38 GB, normalized |
| Embedding table size | Approximately 100 GB | Approximately 560 GB after adding cross-feature tables |
| HSTU-aligned | No sequence axis | Yes |
Table 2: Yambda-5B and Criteo 1TB comparison
The reference implementation extends Yambda’s four native sparse features, item, artist, album, and uid, with hashed cross-product tables. Industrial recommendation systems commonly represent combinations such as user and artist or item and hour of day. Including these features makes the benchmark more representative of production models and increases the embedding footprint to approximately 560 GiB, using a 512-dimensional embedding with fp32 precision.
| Category | Type | Cardinality | Size, 512 dimensions, fp32 |
|---|---|---|---|
| item | Native | 9.39M | 19.2 GB |
| artist | Native | 1.29M | 2.6 GB |
| album | Native | 3.37M | 6.9 GB |
| uid | Native | 1M | 2.0 GB |
| user × artist | Cross feature | 100M | 204.8 GB |
| user × album | Cross feature | 40M | 81.9 GB |
| user × hour | Cross feature | 24M | 49.2 GB |
| item × hour | Cross feature | 40M | 81.9 GB |
| artist × hour | Cross feature | 32M | 65.5 GB |
| user × is_organic | Cross feature | 2M | 4.1 GB |
| user × artist × hour | Cross feature | 40M | 81.9 GB |
| Total | Approximately 560 GB |
Table 3: Yambda-5B feature expansion with cross-product tables
Recommendation traffic is highly skewed: a small number of items account for a large share of interactions. An analysis of approximately 205,000 sampled interaction events measured reads from the item, artist, and album embedding tables. All three distributions closely followed a Zipf pattern.
| Table | Zipf exponent | Fit quality, R² | Top 1% share | Top 10% share |
|---|---|---|---|---|
| item | 0.53 | 0.99 | 16% | 50% |
| album | 0.60 | 0.98 | 18% | 54% |
| artist | 0.83 | 0.96 | 24% | 67% |
Table 4: Popularity skew in Yambda-5B native ID tables
The skew is even stronger from the hardware’s perspective because each event is read through multiple overlapping user histories. The top 1% of accessed rows account for 52% to 68% of embedding lookups, while the top 10% account for more than 90%. This distribution creates a significant embedding-sharding workload at this scale.
Reference Implementation
The reference implementation extends Meta’s open-source HSTU implementation [6], which also supports the DLRMv3 inference benchmark. It integrates that code with Yambda-5B through a preprocessing and streaming pipeline that converts the raw multi-behavior event log into per-user sequences. Cross-feature tables built from the dataset’s native vocabularies expand the embedding footprint to production scale.
HSTU Hyperparameters
The reference model uses three HSTU layers, four attention heads, a model dimension of 512, attention linear and query/key dimensions of 128, and a maximum sequence length of 4,096. Training uses bf16 mixed precision, with attention computed by a fused jagged-attention Triton kernel.
Dense parameters and embedding tables use separate optimizers. Adam updates the dense parameters, while row-wise Adagrad updates the sparse embedding tables. The learning rate warms up to a target scaled according to global batch size.
| Hyperparameter | Value |
|---|---|
| HSTU attention layers | 3 |
| Maximum sequence length | 4,096 |
| Attention heads | 4 |
| Model, or transducer, dimension | 512 |
| Per-head dimension | 128 |
| Precision | bf16 mixed precision |
| Dense optimizer | Adam |
| Sparse optimizer | Row-wise Adagrad |
| Learning-rate warmup | 24,000 steps |
| Target learning rate, 8K global batch | 1e-6 |
| Target learning rate, 16K global batch | 2e-6 |
| Target learning rate, 32K global batch | 4e-6 |
| Embedding table dimension | 512 |
Table 5: Main HSTU hyperparameters
Computation Cost Analysis
The cost analysis uses the following parameters:
| Symbol | Value | Meaning |
|---|---|---|
| B | 1,024 | Per-GPU batch size |
| L | 3 | HSTU layers |
| h | 4 | Attention heads per layer |
| d_qk | 128 | Per-head query/key dimension |
| d_v | 128 | Per-head value/update dimension |
| D | 512 | Model dimension |
| S | 4,096 | Maximum sequence length |
Table 6: Cost-analysis symbols
| Component | Forward plus backward cost expression | Cost per GPU step |
|---|---|---|
| UVQK GEMMs | B · L · 3 · 2 · S · D · (2·d_qk + 2·d_v) · h | 79.16 TFLOPS |
| Output projection GEMMs | B · L · 3 · 2 · S · (3 · h · d_v) · D | 59.37 TFLOPS |
| HSTU causal attention | B · L · 3.5 · 2 · (S² / 2) · h · (d_qk + d_v) | 184.72 TFLOPS |
| Total | 323.26 TFLOPS |
Table 7: HSTU computation-cost breakdown
Weight GEMMs use a factor of three for forward and backward computation because the backward pass produces both input and weight gradients. Attention uses a factor of 3.5 because its backward pass produces dQ, dK, and dV through roughly five GEMMs, compared with two in the forward pass. Causal attention also contributes a factor of one-half because the kernels evaluate only the lower triangle.
Convergence Evaluation
The first convergence experiment compared learning rates using an 8K global batch size.
| Learning rate | Examples to AUC 0.75 | Training time to AUC 0.75 |
|---|---|---|
| 1e-6 | 68.8M | 125.4 minutes |
| 1e-7 | 217.6M | 396.3 minutes |
Table 8: Learning-rate comparison for an 8K batch size
Lower learning rates generally produced more stable convergence across runs, but they could require substantially more examples to reach the same AUC target. Reducing the learning rate from 1e-6 to 1e-7 required roughly three times as many training examples to reach an AUC of 0.75, increasing training time from 125 minutes to 396 minutes at the same throughput.
The reference configuration therefore uses 1e-6 for an 8K batch size, with linear scaling for larger batches and learning-rate warmup to reduce variability between runs with different seeds.
The next experiment evaluated three global batch sizes across 20 random seeds, for a total of 60 runs on MI350 GPUs. Evaluation occurred every 0.1% of the training data. Convergence was defined as the first evaluation that reached an AUC of 0.75 on held-out data.
| Global batch size | GPUs | Samples to AUC 0.75, min/mean/max | Wall-clock time, min/mean/max | CV |
|---|---|---|---|---|
| 8,192 | 8 | 61.9M / 69.3M / 80.3M | 2h22m / 2h39m / 2h59m | 5.6% |
| 16,384 | 16 | 78.0M / 87.7M / 100.9M | 1h39m / 1h52m / 2h12m | 7.2% |
| 32,768 | 32 | 98.6M / 113.0M / 128.5M | 1h17m / 1h23m / 1h36m | 6.7% |
Table 9: Convergence results by global batch size
Figure 2: Convergence curves for 8K, 16K, and 32K global batch sizes
All runs reached an AUC of 0.75. The coefficient of variation, calculated as the sample standard deviation of the 20 per-seed sample counts divided by their mean, ranged from 5.6% to 7.2%. This variation was sufficiently limited for a single run to represent each batch size.
Scaling the learning rate linearly with batch size reduced the sample penalty for larger batches. Quadrupling the global batch size from 8,192 to 32,768 required 1.63 times as many samples on average, increasing from 69.3 million to 113.0 million. Convergence occurred within the first 3% to 5% of an epoch over the 2.29-billion-sample training set.
Conclusion
DLRMv4 updates MLPerf Training for the way modern production recommendation systems are built: sequence models operating on real user histories, attention over thousands of events, and embedding tables measured in hundreds of gigabytes.
The benchmark defines a common model, dataset, quality target, and evaluation process for this class of workloads. It therefore provides a consistent basis for measuring training performance as sequence-based recommendation systems continue to expand.
References
- DLRM-DCNv2, MLPerf Training recommendation benchmark.
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations.
- OneTrans: Unified Feature Interaction and Sequence Modeling with One Transformer in Industrial Recommender.
- DLRMv3, generative recommendation benchmark in MLPerf Inference.
- Yambda-5B: A Large-Scale Multi-modal Dataset for Ranking and Retrieval.
- Meta, Generative Recommenders.
- Wide & Deep Learning for Recommender Systems.
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank Systems.
- MemoNet: Memorizing All Cross Features’ Representations Efficiently via Multi-Hash Codebook Network for CTR Prediction.