Sequential Fitting: A Different Perspective on the Spectral Bias of Neural Networks
by Conor Rowan and Finn Murphy-Blanchard
Introduction
Evidenced by their success with complex tasks such as image classification, autonomy, and language modeling, neural networks are spectacularly good at fitting high-dimensional, nonlinear functions from data. In fact, neural networks have such robust representational capabilities that they can achieve zero training error on images with randomized class labels, meaning there is no structure in the training data which the network can exploit. Despite this flexibility, the neural network model class appears to provide useful inductive bias for many real-world tasks, as neural networks often generalize to unseen test data better than other model types. Yet, regression with neural networks suffers a serious drawback, which has become known as the “spectral bias” in the literature.
Popularized in 2019, the spectral bias states that neural networks fit regression targets from low to high frequencies. As shown in Figure 1, the neural network first learns the low-frequency content of the function, before refining the fit to capture the higher frequencies. As is standard in this literature, we understand the “frequency content” of the regression target to be provided by its Fourier transform.
Figure 1: Rahaman et al. showed empirically that a neural network (green) fits its regression target (blue) in order of increasing frequency.
Because networks fit the target function in order of increasing frequency, learning high-frequency functions is often quite slow, requiring a large number of training epochs. Subsequent works have corroborated the difficulties networks face in fitting high-frequency functions and have offered explanations for this intriguing phenomenon. Some authors have explained the spectral bias by studying the Fourier spectrum of popular activation functions (e.g., ReLU, hyperbolic tangent, sigmoid, etc.), noting that their spectra decay rapidly at high frequencies, and thus the network is inherently biased toward learning low frequencies.
An influential approach called the Neural Tangent Kernel (NTK) offers an elegant explanation of the spectral bias by showing that, in the limit of an infinite-width network, the network output evolves according to a linear dynamical system. Using the theory of linear dynamical systems to decompose the network output into orthogonal modes, it is shown that the convergence rate is inversely proportional to the frequency content of the mode.
A number of other works have explored the spectral bias across different network architectures and optimization algorithms. For example, one work showed that for wide two-layer networks with ReLU activation, the training process can be interpreted as a constrained optimization problem in which high-frequency components of the solution are more heavily penalized. More recently, it was shown from both an empirical and theoretical perspective that second-order quasi-Newton optimization strategies can mitigate the spectral bias for neural networks used in scientific machine learning applications.
While much attention has been paid to understanding the origins of the spectral bias, several researchers have proposed strategies to remedy it. One such strategy is using second-order optimization, while others involve modifications to the architecture of the network. Replacing standard activations with periodic functions like sinusoids is one architectural modification known as a SIREN network. Another popular architecture is the Fourier feature network, which, instead of modifying the activation functions, lifts the input to a higher-dimensional space with periodic embeddings at random frequencies.
In the context of scientific machine learning, Fourier features have been shown to improve performance for multi-scale partial differential equations. The success of standard neural network architectures in mainstream machine learning suggests that fitting high frequencies is not a bottleneck for many application areas. However, for scientific applications, the inability to robustly or efficiently fit high-frequency functions can be problematic, especially where multi-scale and wave propagation problems rely heavily on oscillatory solution fields.
Though the Fourier spectrum of the activation function offers insight into the origin of spectral bias, we believe a more intuitive understanding is possible. In this article, we argue that neural networks fit their target function beginning from the boundary and then progressing into the domain, building one oscillation of the target function at a time. We show that this behavior holds on several example problems in one and two spatial dimensions and find evidence for a “boundary effect.”
One-dimensional regression
Sequential fitting
In the following examples, we work with two hidden-layer MLP neural networks with hyperbolic tangent activation functions. We can write the network explicitly as:
u(x;θ)=w3⋅tanh(w2(tanh(w1x+b1))+b2),θ=[w3,w2,b2,w1,b1],
where θ is the collection of all the trainable parameters of the network and x∈Ω is the spatial coordinate. The widths of the two hidden-layers are taken to be equivalent, and we denote this width as H.
To demonstrate the phenomenon that we call sequential fitting, we begin with a one-dimensional regression problem on the unit domain, where the target function is given by v(x)=sin(26πx). The regression problem is solved with ADAM optimization and follows the specified settings.
Figure 2: Illustration of sequential fitting for a sinusoidal target function.
A second example shows how the envelope of the oscillatory function influences the process. If sequential fitting begins at the boundaries, we hypothesize that the behavior of the function near the boundaries may have an effect on training. In particular, we test the case where an envelope function drives the amplitude of the oscillations to zero on one end of the domain. Our target function is v(x)=x−−√sin(26πx);
Figure 3: Sequential fitting with asymmetric behavior.
Boundary effects
The previous example illustrated that the behavior of the target function near the boundary influences the fitting process. A striking demonstration is seen when fitting the target function v(x)=4x(1−x)sin(26πx). Here, as illustrated, the network fails to make progress toward representing the target due to small oscillations at both boundaries.
Figure 4: Lack of fitting progress with oscillations suppressed near the boundaries.
Basis function perspective
Returning to the example shown, we plot the set of basis functions at discrete training epochs, with the opacity proportional to the coefficient corresponding to each basis function. The results show that each basis function is a smoothed step function, providing insight into the sequential fitting process.
Figure 7: Relevant basis functions as training progresses.
As a final remark, if we switch to a SIREN network, replacing the tanh(⋅) activation with sin(2(⋅)), the basis functions themselves exhibit oscillatory behavior, eliminating the sequential fitting phenomenon.
Figure 8: Basis functions from a SIREN network.
Two-dimensional regression
We now extend our study of the spectral bias to two-dimensional regression problems. We set our computational domain as the unit square and perform midpoint integration. We investigate whether the sequential fitting phenomenon also holds in two spatial dimensions, where the network first becomes non-zero near a boundary and then works inward, as demonstrated in the following figures.
Figure 9: Training progress in two-dimensional regression.
The basis functions built by the network have similar properties to the one-dimensional case, comprising smoothed step functions.
Figure 10: Top basis functions from the two-dimensional case.
Conclusion
After reviewing some of the standard perspectives on spectral bias, we have offered an alternative understanding. We argue that MLP networks fit high-frequency functions from the boundaries inward, learning to represent one oscillation at a time. Our findings suggest that the behavior of the target function near the boundary significantly influences the training process, which we believe is a novel insight. Furthermore, we showed that the MLP networks iteratively built step-like basis functions, contrasting with the oscillatory behavior of those built by a SIREN network. Future studies might investigate the effect of network depth and activation function on sequential fitting and the robustness of the observed boundary effect.