NVIDIA BlueField-4 Enables Scale-In Networking for Agentic AI Factories
Traditional cloud infrastructure was built for predictable, general purpose workloads and standardized interfaces. Agentic AI factories connect users, agents, applications, data...
By Hardware Team
Traditional cloud infrastructure was built for predictable, general-purpose workloads and standardized interfaces. Agentic AI factories connect users, agents, applications, data sources, and storage systems to massively accelerated computing at multi-terabit bandwidth per server. At this scale, dedicated data processing unit (DPU) capabilities are needed to handle networking, storage, and security at line rate.
NVIDIA’s Scale-In network infrastructure is designed as the fifth pillar of its AI networking architecture. It provides purpose-built acceleration for securing, managing, and operating agentic AI factories.
Powered by NVIDIA BlueField-4 and NVIDIA DOCA, and connected through NVIDIA Spectrum-X Ethernet, Scale-In accelerates the services that secure the AI stack and move application, data, and storage traffic throughout the AI factory. Host-independent processing keeps these services off host CPUs, helping prevent security, data-access, and operational functions from becoming bottlenecks as AI compute expands.
This article describes the BlueField-4 architecture and how Scale-In supports secure access, high-performance data movement, tenant isolation, streamlined operations, and predictable performance for agentic AI factories.
Scale-In accelerates north-south AI factory infrastructure
AI factories have distinct infrastructure requirements at different scales:
- Scale-Up: NVIDIA NVLink combines GPUs into a coherent accelerator.
- Scale-Out: NVIDIA Spectrum-X Ethernet and NVIDIA Quantum InfiniBand connect servers across an AI factory.
- Scale-Across: NVIDIA Spectrum-XGS Ethernet connects distributed AI factories.
- Context Memory: NVIDIA CMX, built on the NVIDIA STX modular foundation for AI-native storage, provides shared KV-cache storage within the factory.
- Scale-In: NVIDIA BlueField-4, NVIDIA DOCA, and NVIDIA Spectrum-X Ethernet accelerate access, security, data movement, and infrastructure operations around AI compute.
North-south networks provide the path into and out of a data center. They connect users, applications, data sources, storage systems, and services to individual systems. Traditional cloud data centers built these networks around software-defined infrastructure, composability, and elasticity, allowing resources and access to be provisioned and scaled as demand changed.
Agentic AI increases the demands on this infrastructure. Software-defined networking, composability, and elasticity remain important, but they are no longer sufficient on their own. AI factories combine large-scale accelerated compute with growing numbers of users, agents, applications, enterprise data sources, and storage systems that interact continuously.
Security, multi-tenant networking, data and storage access, and infrastructure operations cannot depend solely on software running on general-purpose host CPUs. They also cannot remain isolated layers managed independently. These functions must operate as a unified infrastructure domain, giving operators consistent control over access, data movement, security, provisioning, and observability.
Scale-In addresses these requirements by extending north-south access into a unified, accelerated infrastructure domain. BlueField-4 supplies host-independent acceleration and offload, DOCA provides a common programming and operating model for infrastructure services, and Spectrum-X Ethernet provides high-performance connectivity across the Scale-In access path, including external storage and data sources.
Together, these technologies provide consistent control over access, data movement, security, provisioning, and observability across the infrastructure supporting accelerated compute and data storage.
Figure 1. Scale-Up, Scale-Out, and Scale-Across expand compute connectivity, while Scale-In connects, secures, provisions, and observes the infrastructure surrounding compute.
AI factories use complementary storage infrastructures for persistent data and inference context. Scale-In provides high-performance access to AI-native storage for training, inference, and analytics, as well as enterprise AI data systems used for retrieval and multimodal indexing. Separately, CMX provides pod-level storage for shared KV cache and reusable inference state inside the factory.
BlueField-4 serves as the infrastructure processor for Scale-In and, separately, as the data and storage processor for CMX. Scale-In connects AI factory and enterprise data to AI compute, while CMX preserves and shares context for faster and more efficient inference.
These five infrastructure pillars address different requirements across the AI factory. Scaling GPUs, racks, and data centers is effective only when data access, storage, cybersecurity, and operations scale with them.
BlueField-4 powers Scale-In infrastructure
BlueField-4 accelerates Scale-In services across GPU servers, agentic CPU systems, AI factory storage systems, and cloud services. In NVIDIA Vera Rubin NVL72, NVIDIA ConnectX-9 SuperNICs carry tenant workload traffic over the Scale-Out network. BlueField-4 runs and accelerates the infrastructure services that connect, secure, and manage each server.
BlueField-4 creates a separate infrastructure-processing domain outside the tenant host. Operators can manage security policies, service state, and telemetry without relying on the host CPU or using its resources.
NVIDIA BlueField Astra extends trusted control from north-south access into the east-west Scale-Out fabric. It provides service providers with a unified, host-independent control point for provisioning, tenant isolation, and network policy across both domains.
BlueField-4 installs and updates policies and monitors telemetry, while ConnectX-9 enforces those policies directly in the data path. This closed-loop model provides consistent control across Scale-In and Scale-Out without placing device management in the tenant host.
BlueField-4 combines programmable control-plane processing with accelerated data-path processing. Software makes infrastructure decisions, while inline engines enforce them locally without sending the work back to the host CPU. Policies can therefore be coordinated across the factory and enforced locally at line rate.
| Component | Role | Benefit |
|---|---|---|
| 64-core NVIDIA Grace CPU | Runs policy, provisioning, telemetry, and infrastructure-orchestration software | Provides 6x more compute than its predecessor, enabling multiple infrastructure services to run concurrently |
| Inline acceleration engines | Process packets, RDMA, storage protocols, encryption, firewall rules, and policy enforcement | Handle operations at up to 800 Gb/s while reducing the processing burden on the Grace CPU and host CPU |
| LPDDR5X memory subsystem | Supplies data and service state to infrastructure software | Provides high memory bandwidth and power-efficient operation without creating a memory bottleneck |
| PCIe Gen6 host connection | Connects BlueField-4 to the host server | Provides a high-bandwidth path between the host and the Scale-In infrastructure-processing domain |
| 800 Gb/s network interface | Connects the server to the Scale-In fabric | Supports high-throughput access, security, data movement, and storage traffic |
Table 1. BlueField-4 components combine control-plane processing, data-path acceleration, memory access, host I/O, and network bandwidth across the Scale-In processing path.
Compared with BlueField-3, BlueField-4 provides 4x more memory bandwidth and 2x more network bandwidth. The additional capacity supports more concurrent services, larger policy and telemetry datasets, and increased traffic-processing and security throughput.
NVIDIA DOCA programs Scale-In services
NVIDIA DOCA turns BlueField-4 hardware accelerators into programmable, deployable Scale-In services. Production-ready, containerized DOCA microservices can run directly on BlueField-4. DOCA libraries and SDKs provide access to accelerated networking, security, storage, and telemetry capabilities for custom services.
- DOCA Flow programs hardware packet-processing pipelines.
- DOCA PCC supports programmable congestion behavior.
- DOCA Telemetry exposes device and service health.
- DOCA Platform Framework (DPF) manages provisioning, deployment, and updates.
BlueField-4’s multiservice architecture and native service function chaining direct each traffic flow through the required sequence of services. These capabilities provide a unified software model for operating services across BlueField-4 systems instead of managing separate server pipelines.
NVIDIA Spectrum-X Ethernet connects the Scale-In fabric
NVIDIA Spectrum-X Ethernet provides the high-performance fabric across the Scale-In access path, including external storage. BlueField-4 processes infrastructure services at each system, while Spectrum-X Ethernet carries traffic between the AI factory and its users, applications, data sources, services, and external storage.
Spectrum-X Ethernet addresses load-balancing conflicts and congestion at scale, improves resource utilization, helps maintain high effective bandwidth, and isolates concurrent traffic. This supports more consistent and predictable performance for access and storage flows.
BlueField-4 DPUs, DOCA microservices, and Spectrum-X Ethernet networking are co-designed with NVIDIA Vera Rubin. Processor performance, memory bandwidth, PCIe connections, network bandwidth, acceleration, and software are designed to operate together. The resulting Scale-In path handles transport, service processing, and control without consuming resources assigned to AI workloads.
Key Scale-In use cases for the agentic AI factory
Agentic AI factories must share accelerated infrastructure safely, provide access to enterprise data, bring systems online, and continuously monitor performance. BlueField-4 acceleration and DOCA services address these operational requirements in several areas.
Build isolated AI factory virtual private clouds
An AI factory virtual private cloud (VPC) isolates tenant and application traffic on shared physical infrastructure. DOCA Host-Based Networking (HBN) accelerates north-south Layer 3 routing and multi-tenant isolation on BlueField-4.
DOCA Flow programs traffic classification and access-control rules, while DOCA-accelerated Open vSwitch (OVS-DOCA) applies the policies on east-west interfaces. BlueField Astra extends the same VPC policy across east-west Scale-Out interfaces, allowing one policy model to govern both access and Scale-Out traffic.
In Vera Rubin, this coordinated model spans 7.2 Tb/s of aggregate interface bandwidth. The total includes 800 Gb/s on the north-south BlueField-4 path and four 1.6 Tb/s east-west paths per compute tray.
This provides consistent policy coverage across access and Scale-Out traffic without routing all east-west traffic through the 800 Gb/s interface. Operators can centrally provision isolated AI factory VPCs while tenants continue using the accelerated fabric. The architecture provides cloud-like elasticity, consistent isolation, reduced host CPU overhead, and more efficient use of shared AI infrastructure.
Enforce security in silicon
BlueField-4 places the enforcement point in hardware, outside the host operating system. Tenant software cannot disable or bypass these controls. Offloaded security processing preserves host CPU cycles and avoids a separate software hop.
BlueField-4 accelerates threat detection and enforces network, file, and object access policies at line rate:
- DOCA Argus provides runtime threat detection.
- DOCA Vault enforces file-access policy.
- DOCA Flow programs line-rate network enforcement.
- BlueField Astra coordinates encryption, isolation, and policy across different networks on each server.
Astra synchronizes policy, telemetry, keys, and enforcement across north-south and east-west traffic. This keeps privileged security controls outside tenant hosts, provides consistent protection for shared AI services, and preserves host resources for AI workloads.
Accelerate storage access
Scale-In connects AI compute to training data, model assets, enterprise knowledge, and application data from external storage. If storage protocol processing, storage virtualization, or data movement cannot keep pace, GPUs may wait for data despite available compute and Scale-Up or Scale-Out network bandwidth.
BlueField-4 accelerates NVMe-oF, file and object storage protocols over RDMA and TCP, storage virtualization, and data movement on dedicated infrastructure. Spectrum-X Ethernet provides the AI-optimized Ethernet fabric across the storage path. Congestion management and performance isolation help sustain high effective bandwidth across concurrent storage flows, delivering up to 1.45x more storage throughput than off-the-shelf Ethernet.
This approach reduces host CPU overhead, helps keep AI compute supplied with data, and provides more consistent access to enterprise data for training, retrieval, and inference. The external-data path complements CMX.
Figure 2. The co-designed NVIDIA BlueField-4 and Spectrum-X Ethernet Scale-In storage path delivers up to 1.45x the throughput of off-the-shelf Ethernet.
Run the AI factory control plane
AI factory nodes must be provisioned, configured, secured, and connected before they can run AI workloads. BlueField-4 onboards nodes, provisions networking and storage resources, starts servers, and loads the host operating system over the network.
Because this control plane operates independently of the host, it can establish policies and provision resources before tenant software starts. DOCA Platform Framework is a Kubernetes-native orchestration framework that provisions, manages, and scales BlueField DPUs as Kubernetes nodes.
It manages DPU discovery, provisioning, service deployment, and updates across the AI factory. This can shorten deployment, improve configuration consistency, preserve host CPU resources, and bring new AI capacity online sooner.
Observe and optimize AI operations
Operating an AI factory at scale requires continuous visibility into network traffic, storage access, service health, performance, and resource utilization. DOCA Telemetry collects and exports this information to network-monitoring platforms. DOCA libraries also enable observability providers to integrate the same infrastructure signals into their tools.
By gathering telemetry through BlueField-4, operators retain visibility independently of tenant-host software. When infrastructure telemetry is combined with GPU and network utilization data, operators can identify constraints involving access, traffic, storage, policy enforcement, or workload placement.
Fleet-wide visibility helps detect anomalies, locate bottlenecks, and resolve incidents before they affect AI workloads.
Scale-In connects scaled compute to AI factory operations
Scale-In evolves north-south networks into a coordinated, accelerated infrastructure domain for data access, security, and operations. NVIDIA BlueField-4 provides networking acceleration and operates the services programmed through DOCA. Spectrum-X Ethernet supplies high effective bandwidth and performance isolation across Scale-In access and storage paths.
Together, these components provide the infrastructure processing, data movement, security, and operational control required to support an AI factory built on scaled compute.