How to Size and Architect Storage for a 40,000-GPU AI Data Center
  • Posted On :2026-10-05
  • Category :Data Center
  • By :Ahmad Tamim

How to Size and Architect Storage for a 40,000-GPU AI Data Center


Article Summary
  • How to estimate storage capacity and bandwidth for a 40,000-GPU AI environment.

  • Why local NVMe, high-performance shared storage, object storage, and checkpoints need to work together.

  • How to design storage networking so that data bottlenecks don't leave expensive GPUs waiting.

What happens when 40,000 GPUs are ready to train, but the storage system can't feed them fast enough?

At small scale, storage can seem like a secondary consideration. At AI data-center scale, it can become one of the most important parts of the architecture.

For organizations planning large AI deployments, Viperatech recommends thinking beyond the question of “how many terabytes do we need?” The better question is: how quickly must data move, where should it live, and how often will the GPUs access it?


How Much Storage Does a 40,000-GPU AI Data Center Need?

There is no universal storage-per-GPU number. Requirements depend heavily on the workloads, datasets, models, checkpoint strategy, and expected growth.

A practical design may need several storage layers:

Storage layer

Primary purpose

Local NVMe

Cache, scratch space and staging

High-performance shared storage

Active training datasets and checkpoints

Object storage

Large data lakes and model artifacts

Backup/archive

Long-term retention and recovery

NVIDIA's current DGX architecture similarly uses local NVMe alongside high-performance shared storage, allowing frequently reused data to be cached closer to the GPUs. 


How Much Storage Bandwidth Do 40,000 GPUs Need?

Capacity tells you how much data you can store. Bandwidth tells you how quickly you can deliver it.

For example, if a workload effectively requires:

  • 1 GB/s per GPU → 40 TB/s aggregate

  • 2 GB/s per GPU → 80 TB/s aggregate

  • 5 GB/s per GPU → 200 TB/s aggregate

Storage requirements also need to be considered alongside the compute hardware and workload. For professional AI environments, GPUs such as NVIDIA RTX PRO 6000 Blackwell GPUs can be part of a broader infrastructure strategy where GPU performance, memory, networking and storage are designed together.

These are illustrative planning calculations, not universal requirements. Actual storage bandwidth depends on the workload and how much data can be served from memory or local NVMe cache.

NVIDIA's SuperPOD guidance specifically notes that storage performance requirements vary by model and dataset. Some computer-vision workloads, for example, can require substantially more read bandwidth than workloads where datasets fit into local cache. 


Why Isn't “TB per GPU” Enough for AI Storage Sizing?

A large storage system isn't automatically a fast storage system.

Imagine a warehouse containing millions of products. The warehouse may have enormous capacity, but if only a few loading bays are available, customers still wait.

AI systems have a similar problem.

Storage planning needs to consider:

  • Sequential and random I/O

  • Read and write throughput

  • IOPS

  • Metadata operations

  • Number of concurrent jobs

  • Dataset reuse

  • Checkpoint traffic

  • Network bandwidth

For AI training, data is often read repeatedly across training iterations, making caching and sustained read performance particularly important. 


What Storage Architecture Should a 40,000-GPU AI Cluster Use?

A tiered architecture is generally more practical than putting everything into one storage system.

A simplified design looks like this:

Object Storage / Data Lake → High-Performance Shared Storage → Local NVMe Cache → GPU Cluster

The ideal architecture also depends on the workload and industry requirements; for example, organizations handling sensitive healthcare data may take a different approach, as discussed in Viperatech's guide to private AI infrastructure for healthcare organizations.

The object layer can hold large datasets and long-term data. High-performance shared storage serves active workloads, while local NVMe can cache or stage frequently accessed information closer to the compute nodes.

This reduces unnecessary network traffic and helps prevent the shared storage layer from becoming a bottleneck. NVIDIA's current systems use local NVMe for caching and staging and high-performance shared storage for large AI datasets. 


NVMe vs. Parallel File Systems vs. Object Storage: What Should You Use?

The answer is often all three, rather than choosing one.

Technology

Best suited for

NVMe SSDs

Local cache, scratch workloads and fast staging

Parallel file systems

High-throughput shared AI training

Object storage

Large datasets, data lakes and model artifacts

Archive storage

Long-term retention

The important point is matching each storage tier to the job it performs.

For example, an organization training large models may keep its master datasets in object storage while staging frequently used training data onto high-performance shared storage and local NVMe.


How Should Checkpoint Storage Be Designed for Large AI Models?

AI training periodically saves the model's state in a checkpoint so training can recover after a failure.

At large scale, checkpoints can become significant write workloads. NVIDIA notes that checkpoint files can reach terabytes and that checkpoint operations can affect training progress when they are performed synchronously. 

A simple way to estimate the required write rate is:

Checkpoint bandwidth = checkpoint size ÷ desired write window

For example, writing a 5 TB checkpoint within 60 seconds requires approximately 83 GB/s of sustained write bandwidth before accounting for overhead.

Local NVMe can also be used to accelerate temporary checkpointing before data is moved to shared storage. 


How Do You Design the Storage Network for 40,000 GPUs?

The storage network needs to be designed alongside the servers and storage, not as an afterthought.

Depending on the architecture, planners may evaluate:

  • High-speed Ethernet or InfiniBand

  • RDMA

  • GPUDirect Storage

  • NIC bandwidth

  • PCIe topology

  • Oversubscription

  • Redundant paths

  • Separate storage and compute traffic

NVIDIA GPUDirect Storage provides a direct data path between storage and GPU memory, reducing reliance on CPU memory as an intermediate “bounce buffer.” This can improve bandwidth and reduce CPU utilization for suitable workloads. 


What Does a Practical 40,000-GPU AI Storage Architecture Look Like?

A hypothetical large-scale design could look like:

40,000 GPUs

↓

Local NVMe caching and staging

↓

High-performance parallel/shared storage

↓

Object-storage data lake

↓

Backup and archive

Storage is only one part of the physical infrastructure. The GPU servers, PCIe architecture, networking and expansion capabilities must also support the storage design. For the server layer, see Viperatech's guide to Supermicro servers for AI and data centers.

The exact number of SSDs, storage nodes, petabytes and network links should be calculated from the organization's actual workloads.

That is the key principle: size the storage system around workload behavior, not simply GPU count.

A deployment built for computer vision, LLM training, inference, scientific computing, or enterprise fine-tuning can have very different storage requirements.


Frequently Asked Questions

How much storage does a 40,000-GPU AI cluster need?

There is no single answer. Capacity depends on datasets, checkpoints, scratch space, model artifacts, retention policies and future growth. Large deployments typically use multiple storage tiers rather than relying on one system.

How do you calculate storage bandwidth for an AI GPU cluster?

Start with the expected effective I/O demand per GPU, multiply it by the number of GPUs, and then account for caching, workload behavior and network overhead. Benchmarking the actual workload is preferable to relying solely on a theoretical number.

Is NVMe storage better than object storage for AI training?

They serve different purposes. NVMe is excellent for local caching and high-speed temporary workloads, while object storage is well suited to large-scale datasets and data lakes. A tiered architecture can use both.

What is the best storage architecture for large-scale AI?

For very large clusters, a combination of local NVMe, high-performance shared storage, object storage and durable backup/archive can provide a better balance of speed, capacity and resilience than a single storage tier.


Conclusion

A 40,000-GPU AI data center is not simply a larger version of a small GPU cluster. At this scale, storage capacity, throughput, caching, checkpointing, networking and resilience have to be designed as one system.

The right question isn't “How many terabytes should we buy per GPU?” It's “What does our AI workload need, and how quickly must the infrastructure deliver it?”

Whether you're building AI infrastructure for model training, enterprise workloads or specialized applications, Viperatech can help organizations evaluate the broader GPU, server, storage and data-center infrastructure requirements.