Skip to main navigation Skip to search Skip to main content

High-entropy advantage in neural networks' generalizability

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

5 Downloads (CityUHK Scholars)

Abstract

One of the central challenges in modern machine learning is understanding how neural networks generalize knowledge learned from training data to unseen test data. While numerous empirical techniques have been proposed to improve generalization, a theoretical understanding of its mechanism remains elusive. Here we introduce the concept of Boltzmann entropy into neural networks. By employing molecular simulation algorithms, we compute entropy landscapes as functions of both training loss and test accuracy (or test loss) across four distinct machine learning tasks. Our results reveal the existence of high-entropy advantage, wherein high-entropy network states generally outperform those reached via conventional training techniques like stochastic gradient descent. This entropy advantage provides a thermodynamic explanation for neural network generalizability: the generalizable states occupy a larger part of the parameter space than its non-generalizable analog at low train loss. We also find this advantage more pronounced in narrower neural networks.

© The Author(s) 2026
Original languageEnglish
Article number44
Number of pages9
Journalnpj Artificial Intelligence
Volume2
DOIs
Publication statusPublished - 16 Apr 2026

Funding

The authors thank National Natural Science Foundation of China for supporting this research (Grant 12405043). We also thank computational resources provided by Bridges-2 at Pittsburgh Supercomputing Center through ACCESS allocation CIS230096.

Publisher's Copyright Statement

  • This full text is made available under CC-BY-NC-ND 4.0. https://creativecommons.org/licenses/by-nc-nd/4.0/

Fingerprint

Dive into the research topics of 'High-entropy advantage in neural networks' generalizability'. Together they form a unique fingerprint.

Cite this