How to test a custom loss function in keras?
ML System Design practice on Codemia
Design recommenders, ranking systems and training pipelines the way ML interviews actually ask for them, with worked solutions.
Introduction
Testing a custom Keras loss function should happen before you trust any training result that uses it. A loss can compile successfully and still be mathematically wrong, numerically unstable, or incompatible with the tensor shapes your model actually produces.
The safest approach is to test the loss in layers: exact values on tiny tensors, shape behavior, gradients, and then one small integration run through model.compile and fit. That gives you fast feedback before a bug burns hours of training time.
Start With a Tiny Deterministic Example
Before involving a model, call the loss directly on a case where you can compute the answer by hand:
For this example, the squared errors are 0.25 and 0.25. After adding 0.1 to each, the mean is 0.35. That makes the expected answer obvious.
Turn that into an assertion:
This simple test catches a surprising number of mistakes immediately.
Check Shape and Reduction Behavior
Many custom loss bugs come from reducing along the wrong axis or returning a scalar when Keras expects a per-sample value. Test the shape explicitly:
If the result shape is not what you intended, the model may still run, but the training semantics will be off.
Verify Gradients Exist and Are Finite
A mathematically reasonable loss is still useless if TensorFlow cannot differentiate through it correctly. GradientTape is the fastest way to verify that:
If gradients are None, NaN, or Inf, training will either fail or behave unpredictably.
Add a Small Integration Test
After unit-testing the loss in isolation, prove that it works inside a real Keras training loop:
This test does not prove the loss is scientifically correct, but it does prove Keras can compile, backpropagate, and train with it.
Test Numerical Edge Cases
If your loss uses division, logarithms, exponentials, or square roots, add tests for boundary values:
- zeros
- tiny positive values
- very large values
- invalid negative values if the math forbids them
These cases are where silent instability usually lives. An epsilon term may be necessary, but test that the stabilized version still behaves the way you intend.
Common Pitfalls
The biggest mistake is testing only model.compile and assuming the loss math is correct because Keras accepted it. Compilation is a very weak test.
Another common problem is forgetting to test gradients. A loss that returns a scalar number is not automatically differentiable in a useful way.
People also mix up per-sample and globally reduced losses. That can change optimization behavior even when the model appears to train normally.
Finally, do not ignore numerical edge cases. If the loss can explode on a rare batch, it will eventually do so in training.
Summary
- Test custom losses directly on tiny tensors with hand-computed expected answers.
- Verify shape and reduction behavior explicitly.
- Check gradients with
GradientTapeand reject non-finite results. - Add one small compile-and-fit integration test.
- Include numerical edge-case tests for losses that use unstable operations.
Related reading
- How to test tensorflow cifar10 cnn tutorial model?
- How to train a customized transformer model with custom dataset formatting
- How to train a model with only an Embedding layer in Keras and no labels
- How to train a `RNN` with LSTM cells for time series prediction
- How to tie word embedding and softmax weights in keras?
- How to train a model in nodejs tensorflow.js?
- How to test if a kernel is a valid kernel
- How to test one single image in pytorch
.png&w=3840&q=75)
Tackling System Design Interview Problems
A short course that equips you with the skills to approach system design interviews methodically.
Start the free courseTrack what you have practised
A free account saves your progress, solutions and study plan across every problem on Codemia.
ML System Design practice on Codemia
Design recommenders, ranking systems and training pipelines the way ML interviews actually ask for them, with worked solutions.