|
IGI Global
Main Office
701 E. Chocolate Avenue
Hershey, PA 17033, USA
Tel: 717-533-8845 x100
Toll Free: 1-866-342-6657
Fax: 717-533-8661
or 717-533-7115
|
|
|
Classification of Imbalanced Data with Random sets and Mean-Variance Filtering:
| Our Price: |
$30.00 US |
| Article #: |
ITJ4204 |
| Number of pages: |
63-78 pages |
| Source: |
International Journal of Data Warehousing and Mining, Vol. 4, Issue 2 |
| Author(s): |
Nikulin, Vladimir |
| Affiliation(s): |
Suncorp, Australia |
Order Now!
This document will be delivered electronically. Terms of Delivery |
|
Description
Imbalanced data represent a significant problem because the corresponding classifier has a tendency to ignore patterns which have smaller representation in the training set. We propose to consider a large number of balanced training subsets where representatives from the larger pattern are selected randomly. As an outcome, the system will produce a matrix of linear regression coefficients where rows represent random subsets and columns represent features. Based on the above matrix we make an assessment of the stability of the influence of the particular features. It is proposed to keep in the model only features with stable influence. The final model represents an average of the single models, which are not necessarily a linear regression. The above model had proven to be efficient and competitive during the PAKDD-2007 Data Mining Competition. |