NETS: Extremely Fast Outlier Detection from a Data Stream via Set-Based Processing

Cited 0 time in webofscience Cited 0 time in scopus
  • Hit : 45
  • Download : 0
This paper addresses the problem of efficiently detecting outliers from a data stream as old data points expire from and new data points enter the window incrementally. The proposed method is based on a newly discovered characteristic of a data stream that the change in the locations of data points in the data space is typically very insignificant. This observation has led to the finding that the existing distance-based outlier detection algorithms perform excessive unnecessary computations that are repetitive and/or canceling out the effects. Thus, in this paper, we propose a novel set-based approach to detecting outliers, whereby data points at similar locations are grouped and the detection of outliers or inliers is handled at the group level. Specifically, a new algorithm NETS is proposed to achieve a remarkable performance improvement by realizing set-based early identification of outliers or inners and taking advantage of the "net effect" between expired and new data points. Additionally, NETS is capable of achieving the same efficiency even for a high-dimensional data stream through two-level dimensional filtering. Comprehensive experiments using six real-world data streams show 5 to 25 times faster processing time than state-of-the-art algorithms with comparable memory consumption. We assert that NETS opens a new possibility to real-time data stream outlier detection.
Publisher
ASSOC COMPUTING MACHINERY
Issue Date
2019-07
Language
English
Article Type
Article
Citation

PROCEEDINGS OF THE VLDB ENDOWMENT, v.12, no.11, pp.1303 - 1315

ISSN
2150-8097
DOI
10.14778/3342263.3342269
URI
http://hdl.handle.net/10203/269074
Appears in Collection
IE-Journal Papers(저널논문)
Files in This Item
There are no files associated with this item.

qr_code

  • mendeley

    citeulike


rss_1.0 rss_2.0 atom_1.0