Training broke with ResourceExausted error
ML System Design practice on Codemia
Design recommenders, ranking systems and training pipelines the way ML interviews actually ask for them, with worked solutions.
Introduction
In the world of machine learning and deep learning, one common stumbling block practitioners encounter is the `ResourceExhaustedError`. This error typically occurs when a system runs out of resources during training, most often due to insufficient GPU memory. Understanding why this happens and how to mitigate it is crucial for the development of efficient and scalable AI models.
Understanding ResourceExhaustedError
A `ResourceExhaustedError` is thrown when the resource requirements of a machine learning model exceed the available limits in the development environment. In the context of deep learning, this primarily refers to memory exhaustion on a GPU. Let's break down the technical aspects:
Technical Explanation
- Memory Requirements: Each operation in a neural network, such as matrix multiplication or activation functions, allocates a portion of the GPU's memory. A deep model with millions of parameters can require a significant amount of memory to store these parameters, as well as intermediate results during forward and backward passes.
- Batch Size: One of the primary contributors to memory consumption is the batch size. Larger batches require more memory, as each input in the batch has its forward and backward computation graphs and gradients stored temporarily.
- Model Complexity: The number of layers, and neurons within each layer, also affects the memory usage. More complex models with deeper architectures will have a higher memory footprint.
Example Scenario
Assume you're training a convolutional neural network (CNN) for image classification. Your model experiences a `ResourceExhaustedError` during training with the following specifications:
- Model: CNN with 5 convolutional layers, 3 fully-connected layers.
- Input size: 256x256 RGB images.
- Batch size: 128.
In this setup, the error might indicate that your current GPU cannot handle the memory requirements to store the gradients, activations, and model parameters for a batch size of 128.
Common Situations and Solutions
Causes
Here's a summarized list of common causes and possible solutions:
| Cause | Solution |
| Large Model Complexity | Solution: Use model pruning or transfer learning to reduce parameters. |
| Large Batch Size | Solution: Decrease batch size. Enable gradient accumulation to simulate a larger batch size. |
| High Resolution of Input | Solution: Downscale images before feeding them to the model. |
| Limited GPU Memory | Solution: Use multi-GPU setups, or consider using paging from CPU to facilitate larger models without fitting them entirely in memory. |
| Memory Intensive Operations | Solution: Replace operations using more efficient implementations such as CuDNN for deep learning operations. |
Advanced Techniques
Gradient Checkpointing
Gradient checkpointing is a technique aimed at saving memory by trading compute for memory. During the forward pass, instead of storing all intermediate activations necessary for the backpropagation step, the model is subdivided into segments. Only the inputs and outputs of each segment are kept, and other intermediates are recomputed in the backward phase, reducing memory usage significantly.
Mixed Precision Training
Mixed precision training involves using half-precision (16-bit) floating-point numbers instead of full precision (32-bit) for computations and storage. This reduces memory usage and increases training performance. The reduction in precision can be managed with loss scaling to maintain numerical stability.
Dealing with Resource Constraints
When faced with resource limitations that lead to `ResourceExhaustedError`, consider the following:
- Profile Resource Usage: Use profiling tools to monitor memory usage and identify operations that are particularly memory intensive. Frameworks like TensorFlow and PyTorch come with built-in profiling utilities.
- Model Optimization Libraries: Libraries such as `TensorRT` or `NNAPI` can optimize models for lower memory usage and deliver improved performance by optimizing the computation graph and selectively employing precision techniques.
- Cloud Solutions: Consider leveraging cloud-based platforms that offer scalable resources. Rapid scaling to more capable hardware can often circumvent local hardware constraints.
Conclusion
Dealing with resource exhaustion due to insufficient memory in training environments is a common challenge faced in deep learning. Understanding the underlying causes and employing strategies such as reducing batch sizes, using advanced techniques like gradient checkpointing and mixed precision training, or optimizing model architectures can alleviate these issues. As models continue to grow in complexity, being equipped with strategies to manage scarce resources efficiently will remain a key competency for AI practitioners.
Related reading
- training by batches leads to more over-fitting
- Training custom dataset with translate model
- Training darknet finishes immediately
- Training data for sentiment analysis
- Training of keras model get's slower after each repetition
- Traveling salesman example with known global optimum
- TransactionManagementError You can''t execute queries until the end of the ''atomic'' block while using signals, but only during Unit Testing
- Transformers model from Hugging-Face throws error that specific classes couldn t be loaded

DSA Fundamentals
Master algorithmic patterns and data structures through hands-on LeetCode-style problems - from arrays and hashing to dynamic programming and advanced graphs.
View the courseTrack what you have practised
A free account saves your progress, solutions and study plan across every problem on Codemia.
ML System Design practice on Codemia
Design recommenders, ranking systems and training pipelines the way ML interviews actually ask for them, with worked solutions.