Abstract
This work describes a data-level parallelization strategy to accelerate the discrete wavelet transform (DWT) which was implemented and compared in two multi-threaded architectures, both with shared memory. The first considered architecture was a multi-core server and the second one was a graphics processing unit (GPU). The main goal of the research is to improve the computation times for popular DWT algorithms for representative modern GPU architectures. Comparisons were based on performance metrics (i.e., execution time, speedup, efficiency, and cost) for five decomposition levels of the DWT Daubechies db6 over random arrays of lengths (Formula presented.), (Formula presented.), (Formula presented.), (Formula presented.), (Formula presented.), (Formula presented.), and (Formula presented.). The execution times in our proposed GPU strategy were around (Formula presented.) s, compared to (Formula presented.) s for the sequential implementation. On the other hand, the maximum achievable speedup and efficiency were reached by our proposed multi-core strategy for a number of assigned threads equal to 32.
Author supplied keywords
Cite
CITATION STYLE
Rodriguez-Martinez, E., Benavides-Alvarez, C., Aviles-Cruz, C., Lopez-Saca, F., & Ferreyra-Ramirez, A. (2023). Improved Parallel Implementation of 1D Discrete Wavelet Transform Using CPU-GPU. Electronics (Switzerland), 12(16). https://doi.org/10.3390/electronics12163400
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.