1Department of Mathematics, TMG College of Arts and Science, Chennai, E-mail: manimannang@gmail.com
2Department of Statistics, Dr. Ambedkar Government Arts, College, Vyasarpadi, Chennai, priyagayu2006@gmail.com
Online published on 10 September, 2021.
This study attempts to identify COVID-19 pandemic affected districts of Tamilnadu using different data mining techniques and visualize with the help of Geographical Information System (GIS). The secondary data sources were collected from Health and Family Welfare Department, Tamilnadu. In this research paper Diagonosed cases, Deaths, Recovered cases, Active Cases, Population and Total case per 1, 000, 000 population were used as parameters. The main objective of this paper is to classify the districts of Tamilnadu based on the above parameters and database. Weka (Waikato Environment for Knowledge Analysis) is a very powerful data mining tool in the area of Data Science. Initially, the given databases are subject to normalize using Apriori algorithm and subsequently applying simple k-means algorithm to identify the groups and are named as three zones, Green (Low Spread Districts), Orange (Moderate Spread Districts) and Red (High Spread Districts). Finally, to cross validate the results of three clusters using various data mining model. In addition, the summary statistics, confusion matrix, visualization of k-mean cluster and classification of various states based on GIS has been included. The correct classification results achieved hundred percent of data mining techniques. The different data mining techniques are explained in the part of methodology.
Normalization, K-means, Naive Baye's Classifier, Multi Layer Perceptron (MLP) and Random Forest