Cisco and Nvidia expand secure AI factory to rack-scale systems as market enters execution era

Cisco and Nvidia are moving their joint AI factory to rack-scale production systems, orderable through Cisco starting in September. The shift targets token generation efficiency, with Nvidia reporting inference costs cut 10x to 30x through software optimization.

Categorized in: AI News Operations
Published on: Aug 26, 2026
Cisco and Nvidia expand secure AI factory to rack-scale systems as market enters execution era

Cisco and Nvidia are moving their joint AI factory offering from design to production. The companies expanded the Cisco Secure AI Factory with Nvidia to rack-scale systems that combine liquid-cooled compute, AI-optimized networking, and validated designs. The full rack-scale solution will be orderable through Cisco in September.

The shift reflects a broader market transition. AI infrastructure conversations are no longer about acquiring graphics processing units. They're about building complete systems that generate tokens reliably and efficiently at scale. GPUs may be the engine, but an AI factory is a system - compute, networking, storage, cooling, software, and operations must work together from day one. If one component fails, expensive capacity sits idle.

In interviews on theCUBE, Cisco's Will Eatherton, senior vice president and head of networking engineering, and Nvidia's Gilad Shainer and Marc Hamilton discussed how neoclouds, sovereign AI programs, and enterprises are moving infrastructure into production.

Time to first token becomes the new metric

Neoclouds illustrate the urgency. Many have customers lined up before GPUs arrive, which means deployment delays translate directly into deferred revenue. Sovereign AI programs face additional pressure because they must balance performance, data control, and national infrastructure requirements.

"When we work with neoclouds, from the moment that they put in the PO for the GPU, they already have their end customers lined up," said Eatherton. "And one of the challenges is just the speed, the expectation that from the moment that is all the project planned, that the GPUs have to come back, that then has to go through all of the final deployment aspects of software and then a bring up and hand over to their income."

The economic shift is fundamental. Traditional enterprise infrastructure was frequently viewed as a cost center to be optimized downward. An AI factory manufactures a digital product: tokens. Those tokens power applications and business processes, giving infrastructure a direct connection to revenue. Utilization, availability, tokens per second, and tokens per watt become business metrics.

"The real cost savings in an AI factory is not about cost savings, but it is about token generation and how do you drive that revenue - so having a repeatable way to do it," said Hamilton.

A rack of components is not an AI factory

An AI factory operates as one enormous computing system assembled from thousands of components. GPUs, network interface cards, switches, cables, storage systems, models, and software libraries must operate in concert. Nvidia describes the AI factory as a five-layer cake encompassing the data center's land, power, and shell; chips; infrastructure; models; and applications.

"Building an AI factory, it's not connecting components and hoping for the best," said Shainer. "Building an AI factory means that you need to build a supercomputer and a supercomputer that needs to be built quickly, needs to be built fast and needs to provide the highest numbers of tokens per second, the highest numbers of tokens per power, and so forth."

Cisco's expanded solution supports HGX and MGX form factors, including Nvidia NVL72 systems and a path toward the Vera Rubin platform. The networking architecture combines Nvidia Spectrum-X Ethernet with Cisco Silicon One. Spectrum-X provides adaptive routing, congestion control, and lossless capabilities for distributed AI computing. Cisco wraps those systems with networking, software, sales, and support.

Reference architectures reduce financial and operational risk

An Nvidia Cloud Partner reference architecture establishes how the entire system should be constructed, tested, and operated. Cisco Validated Designs adapt those requirements to Cisco networking and management technologies. The goal is repeatability - customers should not have to reinvent the system every time they deploy a cluster.

The certification has financial implications. Infrastructure lenders want confidence that the assets they finance will deliver the utilization and performance required to support the investment. Following an established reference architecture can affect financing terms.

"We've actually had some of the largest finance lenders in the world say that they will only finance at preferred rates customers that are following that NCP reference architecture," Hamilton said. "So, this is now a huge selling point for any Cisco customer, any Cisco sales rep."

Availability is where the risk becomes visible. A multibillion-dollar AI factory running at 50% or 60% availability because of cabling, firmware, or software issues destroys the underlying economics. The first token matters, but the billionth token matters more.

Day 2 operations will separate the winners

Getting an AI factory online is only the beginning. Models change rapidly, inference software improves, and organizations continually introduce new workloads. The system must be upgraded and optimized without sacrificing availability.

"If you look at the Blackwell generation of GPUs, over the lifetime of that product, we've driven down the inference costs by x factors," Hamilton said. "Not 1 or 2x, but 10x, 20x, 30x by going through and doing software optimization. So, being able to continuously upgrade that AI factory once you install it is super important."

Cisco is positioning Nexus One and Cisco Cloud Control as a common management layer across routing, front-end networking, storage networks, and the Spectrum-X backend. The addition of AgenticOps creates an opportunity to apply AI-assisted monitoring and lifecycle management across the infrastructure.

"A lot of the industry focus is up to the point that you light up the cluster and you get your first token out," Eatherton said. "That's been a big focus. That's great. But the Day 2 aspects around monitoring and health and availability and software upgrades, and these are things that from a Cisco standpoint, we've put a lot of focus on here over the years."

This becomes more important as inference moves across on-premises systems, neoclouds, and the edge. A common architecture can make those boundaries less visible. An enterprise should be able to run sensitive inference locally, burst into a neocloud, and extend intelligence to edge environments.

Why this matters for operations professionals

The AI infrastructure buildout may be one of the largest capital deployment cycles in computing history, but capital alone will not decide the outcome. Execution will. For operations teams, this means the skills that matter are shifting from component-level management to systems-level orchestration.

Operations professionals who understand validated architectures, deployment automation, and continuous optimization will be the ones who turn expensive infrastructure into productive capacity. The market is moving from GPU scarcity to systems engineering, where networking, software, and operations determine how much useful intelligence an organization produces from every dollar and watt. Time to first token gets an AI factory into the race. Continuous optimization and availability win it - and that work falls squarely on operations. AI for Operations training and an AI Learning Path for Operations Managers can help teams build those capabilities.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)