Abstract
Non-Local Means algorithm (NLM) is a prominent image denoising algorithm. One of the major limitations of NLM algorithm and its variants is the time requirement. In this era of high performance computing, an efficient alternative to reduce the time complexity of any algorithm is its parallelization. In this paper, a parallelized version of basic NLM algorithm using CUDA architecture is proposed. The algorithm is developed on NVIDIA GeForce 940M GPU which follows Maxwell architecture with 3 SMs and 384 CUDA cores. Experiments are carried out using selected set of natural and medical images of various sizes. Our proposed parallelized version of NLM algorithm reduces the time requirement approximately by 50% in comparison to its basic version and also achieves comparable denoising performance in terms of PSNR, SSIM and FSIM evaluation metrics. The proposal is a model which can be customized for newer GPU architectures.
Author supplied keywords
Cite
CITATION STYLE
Wahid, F. F., Sugandhi, K., & Raju, G. (2020). Cuda implementation of non-local means algorithm for GPU processors. Indian Journal of Computer Science and Engineering, 11(1), 66–75. https://doi.org/10.21817/indjcse/2020/v11i1/201101057
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.