Abstract
Transformer networks excel in scientific applications. We explore two scenarios in ultra-high-energy cosmic ray simulations to examine what these network architectures learn. First, we investigate the trained positional encodings in air showers which are azimuthally symmetric. Second, we visualize the attention values assigned to cosmic particles originating from a galaxy catalog. In both cases, the Transformers learn plausible, physically meaningful features.
Author supplied keywords
Cite
CITATION STYLE
Erdmann, M., Langner, N., Schulte, J., & Wirtz, D. (2025). What Exactly Did the Transformer Learn from Our Physics Data? Computing and Software for Big Science, 9(1). https://doi.org/10.1007/s41781-025-00145-4
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.