Abstract
Modern OpenMP threading techniques are used to convert the MPI-only Hartree-Fock code in the GAMESS program to a hybrid MPI/OpenMP algorithm. Two separate implementations that differ by the sharing or replication of key data structures among threads are considered, density and Fock matrices. All implementations are benchmarked on a super-computer of 3,000 Intel® Xeon Phi™ processors. With 64 cores per processor, scaling numbers are reported on up to 192,000 cores. The hybrid MPI/OpenMP implementation reduces the memory footprint by approximately 200 times compared to the legacy code. The MPI/OpenMP code was shown to run up to six times faster than the original for a range of molecular system sizes. CCS CONCEPTS •Theory of computation → Massively parallel algorithms; •Computing methodologies → Rantum mechanic simulation; Massively parallel and high-performance simulations;
Author supplied keywords
Cite
CITATION STYLE
Mironov, V., Alexeev, Y., Keipert, K., D’Mello, M., Moskovsky, A., & Gordon, M. S. (2017). An efficient MPI/OpenMP parallelization of the Hartree-Fock method for the second generation of Intel® Xeon PhiTMprocessor. In International Conference for High Performance Computing, Networking, Storage and Analysis, SC (Vol. 2017-November). IEEE Computer Society. https://doi.org/10.1145/3126908.3126956
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.