Bonanza Offer FLAT 20% off & $20 sign up bonus Order Now
ALY6010
US
Northeastern University
This project aims to use an exploratory data analysis (EDA) to analyze the dataset, explore relationships in one variable with multiple variables, and explore the Winetasting dataset for distributions, visualization, outliers, and summarise essential characters. To do EDA, the wine testing dataset is used. It contains the survey of different countries and from different sources like twitter, from other regions, and the taster's name to do analysis.
The dataset includes 1000 observations and 14 variables. The variable country describes different types of the country taken the wine. There are six types of country, and some other country tested wine. The variable description and designation tell details of how they have taken the wine and their designation. Points variables tell about the given points by the taster. Price for the price of the wine. Region_1 and region_2 is the region from where the taster belongs. Taster_name is the name of the taster tastes the wine, and taster_twitter_handle is the taster's Twitter address. The title is the title of the wine, and the variety is the variety of wine.
The purpose of the dataset is the detailed analysis of testing review to determine which country wine tastes wine more. What are the price rate and the name of the person who tastes more wine, the points they have given for the taste. Visualize the analysis in different types of charts for detail visualization. To import the dataset, the command read.csv() and View() command is used to view the import dataset to do the analysis.
The data included for analysis are both numerical and character data type. The variable points and price are a numerical variable. Except for these two variables, others are categorical variable. To check the kind of data, the str() command is used. It tells detail of the structure of the data present in the dataset. (Appendix 1)
Class() command is for checking the class of the data set. Here the class is data.frame. dim() command is to check how many observation present in the dataset and the number of variables present in the dataset. names() command is to display the name of the variable, summary() command display the details of the variable on the basis of variable, it gives the mean, median and mode, quartile and class of the variable. (Appendix 2)
There are 1000 rows and 14 variables present in the dataset. To check the field of the dataset the str() command is run. The output displays the class and level of the dataset. To check the lavels of the variable for categorical variable convert character data type into factor.
To do data cleaning, first need to check null values. To check null values is.na() command is run. There is 57 null value present in the price variable, and it replaces by the mean of that variable. The mean is 32.44 for the price column. To check the outliers present in the dataset, the box plot is plotted. The outliers are present in the price and points column. The outliers present in the price are replaced with 3rd quartile value, and outliers present in points are replaced with the 3rd and 2nd quartile value. So there is no missing and outliers value present in the dataset.
The histogram is plotted for a continuous variable, In the wine tasting dataset the continuous variable is the price and points. The command plot_histogram to plot the histogram. The histogram shows that the value 20 tastes more than others. More than 100 frequency are for value 20. The histogram for points shows that more than 400 people give points from 0 to 89, and more than 100 peoples for 90 points.
The next step is to check the variety of wine. The library dplyr command is to group the variety of wine, group() command to group the variety, and summarise the grouped variety. The variety named chardonnay has 79 variety. Then group the Twitter handle of the wine taster. The Twitter handle named @vossroger has the highest review in the Twitter, 170 review comes from these Twitter account. The next command is to check region 2, in which region most of the wine taster belongs. There is 78 wine taster from the central coast region and 68 from the Sonoma region. 633 taster in region 2 does not mention their exact region name. The next is group 1, 165 taster does not mention their region and 24 tasters from Alsace in region1.
The country named the US subtracted from the wine data and then visualize to do detail analysis. Different types of plot are used to visualize the US dataset. (Appendix 5)
From the analysis of the wine tasting dataset, the conclusion is the wine is good in taste. The points give to wine is between 87 to 94. The price is on average, not more than 76, so lots of countries wine price is 76. There are lots of variety of wine. The taster gives a good review of the wine.
Are you in dire need of assignment help in the UK? Can’t figure out who can help you whenever you find yourself thinking, “Wouldn’t it be great if I could pay someone to do my assignment?” With Myassignmenthelp.co.uk, you can fulfil your desires without any hassle.
Send us your requirements, and our paper writers will take care of your assignment worries quickly. So now, you don't have to worry about, "Where can I find someone to do my assignment for me in the UK?” Instead, let our experts provide you with the best assignment help in London, Bristol, Manchester, Liverpool and more!
Upload your Assignment and improve Your Grade
Boost Grades