An Outlier Detection Algorithm Based on Probability Density Clustering

2Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.

Abstract

Outlier detection for batch and streaming data is an important branch of data mining. However, there are shortcomings for existing algorithms. For batch data, the outlier detection algorithm, only labeling a few data points, is not accurate enough because it uses histogram strategy to generate feature vectors. For streaming data, the outlier detection algorithms are sensitive to data distance, resulting in low accuracy when sparse clusters and dense clusters are close to each other. Moreover, they require tuning of parameters, which takes a lot of time. With this, the manuscript per the authors propose a new outlier detection algorithm, called PDC which use probability density to generate feature vectors to train a lightweight machine learning model that is finally applied to detect outliers. PDC takes advantages of accuracy and insensitivity-to-data-distance of probability density, so it can overcome the aforementioned drawbacks.

Cite

CITATION STYLE

APA

Wang, W., Ren, Y., Zhou, R., & Zhang, J. (2023). An Outlier Detection Algorithm Based on Probability Density Clustering. International Journal of Data Warehousing and Mining, 19(1). https://doi.org/10.4018/IJDWM.333901

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free