Models

NVIDIA Combines Palantir and cuOpt to Optimize Global Supply Chain Operations

NVIDIA is leveraging Palantir Foundry and cuOpt to streamline hardware supply chain decisions across its worldwide manufacturing network, measuring performance from chip production through system deployment.

·3 min read
Palantir Foundry and cuOpt drive NVIDIA supply chain allocation
Palantir Foundry and cuOpt drive NVIDIA supply chain allocation

NVIDIA has deployed Palantir Foundry and cuOpt to optimize hardware supply chain allocation decisions at its global manufacturing facilities.

The chipmaker tracks operational performance across the entire journey from wafer-out to first token. This span divides into two segments: time-to-rack, which measures the movement from fabrication output to a finished data centre system, and time-to-token, encompassing power delivery, thermal management, network setup, and initial software deployment.

Managing NVL72 and Vera Rubin component flows

Expanding hardware production has intensified supply chain pressures. Each NVIDIA Grace Blackwell NVL72 rack incorporates 18 compute trays, and every tray demands two Grace CPUs, four Blackwell GPUs, and 32 HBM3e memory packages distributed among thousands of suppliers, OEMs, and contract design partners.

The emerging supply network for NVIDIA's Vera Rubin architecture will be double the scale of the existing Grace Blackwell infrastructure.

Production cannot advance until components arrive through three distinct pathways: direct inventory, consignment arrangements, and third-party suppliers. Earlier deliveries must pause for slower-arriving parts, prolonging what NVIDIA calls 'Time of Ownership'—the interval between material receipt at a facility and the departure of completed sub-assemblies.

Manufacturing site assignments undergo revision every seven days across rolling two-quarter planning windows to balance component availability, production throughput, and customer delivery commitments.

Mixed-integer linear programming via cuOpt

To manage these interconnected factors, NVIDIA's operations group constructed the 'Digital Supply Chain Intelligence' command centre leveraging Palantir Foundry. Foundry's Ontology represents manufacturing sites, supplier commitments, inventory levels, and production objectives as linked entities and relationships.

NVIDIA cuOpt, a publicly available library for GPU-powered optimization algorithms, accesses this operational framework directly. By framing allocation as a mixed-integer linear program aimed at reducing TOO, the optimization engine assesses component limitations throughout every level of the bill of materials.

In addition to generating weekly production schedules, cuOpt pinpoints binding factory constraints, such as regional assembly capacity restrictions relative to available memory supplies.

Training Nemotron on qualitative operational records

Quantitative optimization methods alone proved insufficient to incorporate unstructured operational factors that human planners routinely observe, including supplier call recordings, regional climate data, partner communications, and international political circumstances.

NVIDIA addressed this limitation by fine-tuning Nemotron 3.5 Lightning, an open-weight mixture-of-experts architecture with 30 billion total parameters and roughly three billion active parameters per inference step.

The preparation workflow runs historical records through NeMo Anonymizer to remove confidential operational information, NeMo Data Designer to equilibrate training samples with artificial supply disruption cases, and NeMo AutoModel to implement low-rank adaptation (LoRA) modifications while preserving original model parameters. Palantir Autopilot oversees data provenance, model versioning, and suggestion delivery.

Production benchmarks and future reinforcement learning

When tested against historical allocation information, the fine-tuned Nemotron 3.5 Lightning variant demonstrated 86.7 percent decision accuracy, surpassing 55.5 percent for the larger Nemotron 3 Ultra and 17.5 percent for the baseline Lightning model without customization.

The customized model reached 58.6 percent balanced accuracy and 57.5 percent macro-F1 score, exceeding Nemotron 3 Ultra's 42 percent balanced accuracy and 39.5 percent macro-F1 score.

Model customization ran on two NVIDIA B200 GPUs in just minutes. Specialized fine-tuning enhanced allocation recommendations, though forecasting production hazards further ahead remained challenging.

Allocation determinations, human planner modifications, manual interventions, and actual factory results are consistently fed back into the Palantir Ontology.

NVIDIA stated this information will serve as preference pairs for reinforcement learning algorithms—assessing suggestions on allocation effectiveness, regulatory adherence, and reasoning transparency—with operational systems staying strictly separated from uncontrolled live retraining.