how to programmatically determine available GPU memory with tensorflow?
ML System Design practice on Codemia
Design recommenders, ranking systems and training pipelines the way ML interviews actually ask for them, with worked solutions.
Introduction
Programmatically checking "available GPU memory" with TensorFlow is trickier than it sounds because there are two different questions you might mean. One is how much memory TensorFlow's allocator is currently using. The other is how much free memory the GPU has overall at the device level. TensorFlow can report allocator statistics directly, but device-wide free memory is usually easier to query through NVIDIA's management library.
TensorFlow Allocator Memory Info
In TensorFlow 2, the most direct built-in API is tf.config.experimental.get_memory_info.
The returned dictionary typically includes fields such as:
- '
currentfor memory currently used by TensorFlow's allocator' - '
peakfor peak allocator usage'
This is useful when you want to understand TensorFlow's own memory behavior during model creation or training.
It is not the same as asking the GPU driver how much memory the whole device still has free for every process on the system.
Device-Wide Free Memory With NVML
For actual free GPU memory on NVIDIA devices, NVML is usually the more accurate tool. In Python, the common wrapper is pynvml.
This reports device-level totals from the NVIDIA driver, which is closer to what most people mean by available GPU memory.
If your TensorFlow program shares the GPU with other processes, NVML gives a better picture of global pressure than TensorFlow's allocator stats alone.
Using Both Together
The most informative setup is often to read both views.
This lets you compare:
- memory managed by TensorFlow
- memory occupied on the device overall
That distinction matters because TensorFlow may reserve memory or share the GPU with other libraries and processes.
Controlling Memory Growth
TensorFlow can reserve GPU memory aggressively unless configured otherwise. Enabling memory growth often makes diagnostics easier.
With memory growth enabled, TensorFlow allocates GPU memory more incrementally instead of grabbing most of it upfront.
That does not directly answer the memory query question, but it changes how the reported numbers behave and often prevents confusion when monitoring allocations.
A Practical Helper Function
Here is a small helper that returns both TensorFlow allocator stats and NVML stats when available.
This is useful in training scripts, experiments, and debugging notebooks.
Common Pitfalls
A common mistake is assuming TensorFlow's get_memory_info reports total free GPU memory for the whole system. It usually reports allocator information, not a full device-wide view.
Another mistake is forgetting that TensorFlow may reserve memory early unless memory growth is enabled. That can make the GPU appear more occupied than expected.
Developers also sometimes query GPU state before TensorFlow has initialized the device, then compare it to values gathered later under a different memory policy.
Finally, remember that these techniques are most practical for NVIDIA GPUs. The NVML approach is vendor-specific.
Summary
- TensorFlow can report its own allocator memory with
tf.config.experimental.get_memory_info. - That built-in API is not the same as global device-wide free memory.
- For NVIDIA GPUs, NVML via
pynvmlis the usual way to get total, used, and free device memory. - Using both TensorFlow stats and NVML gives the clearest picture.
- Memory growth changes TensorFlow allocation behavior and can make diagnostics easier.
- Be explicit about whether you want allocator usage or whole-device availability.
Related reading
- How to Properly Combine TensorFlow's Dataset API and Keras?
- How to properly feed specific tensor to keras model
- How to properly manage memory and batch size with TensorFlow
- How to properly reduce the size of a tensorflow savedmodel?
- How to properly set steps_per_epoch and validation_steps in Keras?
- How to properly set steps_per_epoch and validation_steps in Keras?
- How to properly use from_logits in keras loss function for binary classification?
- How to properly use tf.function while subclassing keras Layer/Model?
.png&w=3840&q=75)
Tackling System Design Interview Problems
A short course that equips you with the skills to approach system design interviews methodically.
Start the free courseTrack what you have practised
A free account saves your progress, solutions and study plan across every problem on Codemia.
ML System Design practice on Codemia
Design recommenders, ranking systems and training pipelines the way ML interviews actually ask for them, with worked solutions.