How to run tensorflow with gpu support in docker-compose?
ML System Design practice on Codemia
Design recommenders, ranking systems and training pipelines the way ML interviews actually ask for them, with worked solutions.
Introduction
Running TensorFlow with GPU support in Docker Compose depends on more than the container image. The host must expose an NVIDIA GPU correctly, Docker must be able to pass that device through, and the Compose file must explicitly reserve GPU access for the service.
If any one of those pieces is missing, TensorFlow usually starts successfully but sees only CPU devices. A good setup verifies GPU visibility at each layer before you try to train a real model.
Check Host Prerequisites First
Before touching Compose, confirm that the host itself can see the GPU. On a Linux host, nvidia-smi should work outside containers.
You also need the NVIDIA Container Toolkit so Docker can expose GPU devices to containers. A quick smoke test is running a CUDA image directly:
If this command fails, Compose will fail too. Fix the host driver or container runtime first instead of debugging TensorFlow.
Use a GPU-Capable TensorFlow Image
Your container image must include GPU-enabled TensorFlow and compatible CUDA libraries. A common choice is the official GPU-tagged TensorFlow image.
The important part is the devices reservation under deploy.resources.reservations. The capabilities field is required. Use count when you want a number of GPUs, or device_ids when you want specific ones. Do not set both in the same reservation.
Verify TensorFlow Sees the GPU
Create a minimal script and run it through Compose before adding training code.
Then start the service:
If the output shows at least one physical GPU, your container is wired correctly.
Target a Specific GPU
On multi-GPU machines, pin the service to one device so jobs do not fight over the same hardware.
You can get device IDs from nvidia-smi. This is useful for shared servers where one Compose project should not consume every GPU on the host.
Keep the Compose File Simple
Once GPU access works, keep the service definition boring. Mount your project, set a working directory, and run one deterministic command. For example:
The less extra logic you put into startup, the easier it is to tell whether a failure comes from Docker, TensorFlow, or your application code.
Distinguish Compose Problems from TensorFlow Problems
If docker run --gpus all ... nvidia-smi works but TensorFlow still sees no GPU, the issue is usually image compatibility. If even the CUDA container cannot see a GPU, the problem is lower in the stack and unrelated to TensorFlow.
That separation saves time. Too many people start by changing Python packages when the actual issue is that the Docker runtime never exposed the device in the first place.
Common Pitfalls
- Skipping the host-level
nvidia-smicheck and trying to debug everything from inside TensorFlow. - Using a CPU-only TensorFlow image and expecting Compose GPU settings to fix it.
- Omitting
capabilities: [gpu], which is required in the device reservation. - Setting both
countanddevice_idsin the same GPU reservation. - Running large training jobs before confirming with a tiny script that TensorFlow can see the GPU.
Summary
- Verify the NVIDIA driver and container runtime on the host before using Compose.
- Use a TensorFlow image that actually includes GPU support.
- Reserve GPU devices explicitly in the Compose service definition.
- Confirm GPU visibility with a short TensorFlow script before training real models.
- On multi-GPU hosts, pin services to specific devices when isolation matters.
Related reading
- How to save and restore Keras LSTM model?
- How to save and restore partitioned variable in Tensorflow
- How to save final model using keras?
- How To Save Keras Regressor Model?
- how to save a scikit-learn pipline with keras regressor inside to disk?
- How to save a trained tensorflow model for later use for application?
- How to sample batch from only one class at each iteration
- How to sample large database and implement K-means and K-nn in R?

System Design Fundamentals
Build a strong foundation in designing scalable, reliable distributed systems.
View the courseTrack what you have practised
A free account saves your progress, solutions and study plan across every problem on Codemia.
ML System Design practice on Codemia
Design recommenders, ranking systems and training pipelines the way ML interviews actually ask for them, with worked solutions.