On September 10, 2026 NVIDIA and Palantir announced a joint deployment of artificial intelligence across critical supply chains. This is not a laboratory pilot: the first customer is NVIDIA itself, which uses the stack to decide how to allocate materials, manufacturing capacity and scarce components along a supply chain the company describes as one of the most complex in the world, spanning millions of parts and thousands of suppliers. The announcement lands as the AI race enters a phase where the competitive constraint shifts from the power of a single chip to the ability to build, distribute and power it.
From wafer to token: what the stack actually does
The architecture brings NVIDIA Nemotron open models into Palantir Foundry and AIP, Palantir's artificial intelligence platform, grounded in the Ontology: the layer that organises scattered data and logic into objects, links and actions representing the real world. The stated goal is to create supply chain visibility, identify constraints earlier, continuously codify operational expertise and guide decisions at machine speed, while retaining control and ownership of proprietary data. Post-trained models recommend actions, explain tradeoffs and flag emerging risks, but supply chain experts keep control of final decisions. Inside AIP, NVIDIA cuOpt software unlocks optimisation and scenario planning, so teams can model supply constraints and evaluate the operational impact of allocation choices.
The mathematics of allocation: cuOpt and Time of Ownership
According to NVIDIA's technical documentation, the chain is measured from wafer-out to first token and splits into two legs: time-to-rack, from silicon leaving the fab to an assembled system arriving on a data center floor, and time-to-token, everything thereafter, including power, cooling, networking and the software stack. NVIDIA's supply chain team built a command centre with Palantir, the Digital Supply Chain Intelligence centre, unifying every input to an allocation decision inside Foundry. The quantitative side runs on cuOpt, which solves a weekly mixed-integer linear programme minimising Time of Ownership, the time a manufacturing site holds material before it leaves as a sub-assembly or product. The formulation also maps the dependency graph: a compute tray is blocked by its scarcest input rather than its average one, and the binding constraint shifts from week to week.
One figure conveys the physical complexity: a GB200 NVL72 compute tray requires two Grace CPUs, four Blackwell GPUs and thirty-two HBM3e stacks. A Vera Rubin rack, the next-generation platform, adds up to 1.3 million parts across compute, memory, networking, power, cooling and mechanical components.
The most interesting, and most candid, detail in the technical documentation is another one: human planners consistently outperformed the quantitative model, because they factored in emails, weather forecasts, geopolitical events and supplier debriefs that the solver could not see. That is precisely why the project is not about replacing people, but about codifying their expertise.
Post-trained Nemotron: 86.7% on allocation decisions
NVIDIA therefore post-trained Nemotron 3.5 Lightning, a 30-billion-parameter model, on captured allocation decisions, rationales and outcomes, using NeMo Anonymizer, Data Designer and AutoModel inside a governed lifecycle on Palantir Autopilot. The specialised model reached 86.7% allocation-decision accuracy on the development benchmark: 31.2 percentage points ahead of the larger Nemotron 3 Ultra and 69.2 points ahead of its own base model. Accepted, edited and overridden recommendations feed back into the Ontology and drive further governed retraining runs: institutional knowledge becomes an accumulating asset that shortens onboarding.
Not the first step
The partnership did not start now. On October 28, 2025 the two companies announced a first integrated technology stack for operational AI, with Lowe's among the first adopters, building a digital replica of its global logistics network. The perimeter has widened since: the new reference architecture, the Palantir Sovereign AI Operating System supported by Dell Technologies and Cisco, is designed for sovereign and on-premise deployments, and the model is being offered to manufacturing, energy, healthcare, automotive and aerospace.
The market gives no discounts
Wall Street shrugged. According to 24/7 Wall St., NVIDIA was down 1.35% and Palantir 1.88% in premarket trading; Giocare in Borsa reported NVIDIA at 218.92 dollars, down 2.12% over 24 hours. The reason is that the deal does not change the numbers that matter right now. Vera Rubin shipments began in August 2026 and the platform should account for roughly 20% of data center revenue in the current quarter, with the revenue opportunity per gigawatt rising to 40 billion dollars from 25 billion on the Blackwell generation. Data center revenue reached 89.02 billion dollars in the second quarter, up 117% year on year, and guidance for the current quarter is 108 billion. But the company guided to a fourth-quarter non-GAAP gross margin of 71-72%, squeezed by extreme pricing conditions in memory, and carries 279 billion dollars of supply obligations tied largely to memory procurement for Vera Rubin. Palantir closed the second quarter with revenue of 1.94 billion dollars, up 92.8%, US commercial revenue of 764 million, up 149%, full-year guidance of about 8.15 billion and a Rule of 40 of 155%. Valuations remain stretched: NVIDIA is worth around 5.4 trillion dollars, Palantir about 390 billion.
The real bottleneck: chips, power or infrastructure?
The three answers are not alternatives: they move. On memory, a KB Securities note reported by TechRadar points to a severe supply shortage in 2027, while infrastructure investment has grown 60% year on year to 1.3 trillion dollars, burning memory as fuel. On the power side the problem is not national generation but its local concentration. TechRepublic reports figures from the International Energy Agency: global data center electricity demand grew 17% in 2025, with consumption at AI-focused data centers jumping 50%, and total use is projected to rise from 485 terawatt-hours in 2025 to 950 in 2030. RAND estimates that US plans for 2030, 151 gigawatts of announced front-of-the-meter capacity plus 149 gigawatts behind the meter, translate into about 82 gigawatts of net deliverable capacity, concentrated mostly in Texas. A July 2024 incident in Northern Virginia, recalled by Harvard's Belfer Center, disconnected sixty data centers at once, creating a 1,500 megawatt surplus that forced emergency grid adjustments: a concentrated-load stability problem, not a generation shortfall. And the delivery bill is already steep: Goldman Sachs estimates only 50-60% of planned US capacity will open on schedule, hit by grid delays and equipment shortages. The bottleneck, in other words, is shifting from silicon to the layer that organises and powers silicon. Which is exactly the ground on which the September 10 announcement is being played out.
Sources