Abstract
One of the central challenges in modern machine learning is understanding how neural networks generalize knowledge learned from training data to unseen test data. While numerous empirical techniques have been proposed to improve generalization, a theoretical understanding of its mechanism remains elusive. Here we introduce the concept of Boltzmann entropy into neural networks. By employing molecular simulation algorithms, we compute entropy landscapes as functions of both training loss and test accuracy (or test loss) across four distinct machine learning tasks. Our results reveal the existence of high-entropy advantage, wherein high-entropy network states generally outperform those reached via conventional training techniques like stochastic gradient descent. This entropy advantage provides a thermodynamic explanation for neural network generalizability: the generalizable states occupy a larger part of the parameter space than its non-generalizable analog at low train loss. We also find this advantage more pronounced in narrower neural networks.
© The Author(s) 2026
© The Author(s) 2026
| Original language | English |
|---|---|
| Article number | 44 |
| Number of pages | 9 |
| Journal | npj Artificial Intelligence |
| Volume | 2 |
| DOIs | |
| Publication status | Published - 16 Apr 2026 |
Funding
The authors thank National Natural Science Foundation of China for supporting this research (Grant 12405043). We also thank computational resources provided by Bridges-2 at Pittsburgh Supercomputing Center through ACCESS allocation CIS230096.
Publisher's Copyright Statement
- This full text is made available under CC-BY-NC-ND 4.0. https://creativecommons.org/licenses/by-nc-nd/4.0/
Fingerprint
Dive into the research topics of 'High-entropy advantage in neural networks' generalizability'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver