Bonanza Offer FLAT 20% off & $20 sign up bonus Order Now
ITEC202
AU
Australian Catholic University
The dataset used in this project is the American Housing dataset. This dataset contains 81 variables of interest and contains a total of 1460 observations(houses). The original purpose of the dataset was for tax assessment services, but the dataset could also be used to help in computing the assessed values for individual residential properties and was gathered between the years 2006 and 2010. This is because the variables contained in this dataset mostly represent exactly what a home buyer would inquire about a house before purchasing it. The chief custodian of the dataset is Iowa’s assessor’s office.
23 Nominal variables
23 Ordinal Variables
14 Discrete variables
20 Continuous variables.
For this project, the target variable we are interested in is the Sales price variable which represents the sale value of the particular house. The rest of our variables will be predictor variables
In analyzing the dataset, I made a few changes concerning variable transformation as shown below:
All discrete variables which indicated years were transformed to age, this transformation was done by taking subtracting the year variables from the current year 2021, since it is more valid to model the relationship between the age (e.g., say of the house and its sales price), than modelling the year the house was built and its sales price. The age of the house has a severe impact on the sales price of the house more than the year the house was built.
A few questions of interest that emerge are in our analysis are:
Data preprocessing is the technique that involves taking data that is in a messy, undesirable and unorganized format and incomplete, and transforming it into an understandable format. In our analysis here, data preprocessing will involve dealing with missing values appropriately, converting variables from one type into the desired type, dealing with outliers, and perform standardization and transformation where possible. This will enable us to end up with a dataset that is in proper order for modelling.
The missing values in continuous variables were replaced by the mean of the variable as a whole or the mean of the grouped variable as shown:
The missing values for the nominal and ordinal categorical variables also had to be inputted. Based on a background study that I did on the dataset, the majority of the missing values in the categorical variables were inputted with None as shown below:
The fireplace Quality variable had to input with No fireplace, those in Pool quality had to be inputted with No pool, and those in Fence quality had to be inputted with No fence.
The missing values in discrete variables such as years were transformed by including the Age variable. This is because for a missing value in a variable such as GarageYrBlt - The year the garage was built, we would include the age of that observation as 0 - Implying that there is no garage.
Since in our importation of datasets to our R environment, we included the as.is=T command, all categorical variables are by default character data type, thus we can coerce the ordinal and nominal variables to be factor data type variables, as fits.
Based on my view, some data values are outliers. In the data, there were observations of houses that had abnormally large Above ground living areas. I omitted these observations which had greater than 4000 square feet of the above-ground living area since that seems unreasonable.
Now, this data contains no missing values and possible outliers and is ready for visualization.
In this section, we use visualizations to better understand the relationship between the sales price and other variables. We will need ggplot2 and dplyr packages for visualization and data manipulation respectively.
Based on this plot on the relationship between the building class and sales price, we see that Building class 60 and 120, attract higher sales price, due to their higher median values than the rest of the classes. Building class 060 represents two-story buildings(Newer), and 120 represents 1-story buildings(Newer). The lowest sales price are observed in class 030 and 180 which represent older buildings(1945). The Tukey's honesty significance test however reveals that the majority of the Building classes do not have a difference in impacting sales prices, thus we could still omit this variable in our modelling process.
From this plot, we see that zone FV and RL representing Floating Village residential and Residential Low density have higher sales price than the other classification zones such as RH and RM for Residential High density and Residential Medium density. This could probably be because houses in floating villages and residential low density are mostly retirement houses, while in high and medium density they are mostly simple apartments for the majority working class.
From this plot, we see that there is a positive linear relationship between lot frontage area and sale price, with the blue line showing the general trend.
In this plot, we see that the relationship between the Lot area and the sales price is a positive one, but the relationship is not linear-linear, instead, it is a linear-quadratic relationship, as evidenced by the correlation test. The correlation between Salespriceand sqrt(LotArea) is stronger than the correlation between Salesprice and LotArea. This implies that in our modelling process, we will include the square rooted term instead of the linear term.
From the plot on a=street and alley access, it is evident in both that houses which have paved road and alley access fetch higher sales price than those which are accessible on gravel. Houses with No alley access also have higher sales price generally than houses with Graveled-alley access.
Generally, Lot shapes with regular shapes fetch lower sale prices than irregularly shaped lots. This can be shown by the median sales price of the irregularly shaped lotsIR3, IR2, IR1 being higher than the median price of regularly shaped lots. It is also evident from the above plot, that Hillside contours HLS, have generally higher sales prices, while Banked BNK land contours have lower sales prices.
It is evident that in the dataset, only a single house had No sewerage and No water supply, and it has a lower sales price than the majority of the houses with all utility present (Electricity, Gas, Water and Sewerage utilities).
The majority of houses, with CulDSac lot configuration, fetched higher sales price, than the rest of the lot configurations, with the lowest sales price emerging from Inside Lot configurations
Generally, all the land slope types generate fairly equal sales price, making it not so helpful in determining which property yields higher sales price based on land slopes.
A glance at this plot indicates that the majority of the houses nearest to Northridge, Stony Brook and North Ridge Heights have higher sales price than the rest. This can be explained because…
From this plot above, it can be seen that houses with proximity to a positive off-site feature PosA and PosN have higher sales price. While houses proximate to Arterial and feeder streets generally attract lower house sales price …
Generally, 1 Family dwelling and TownHouse dwelling have higher sales price than the rest of the types of dwellings. Although all the types of dwellings have fairly equal prices an analysis of variance test shows that the Building type used has a significant effect on the sales price.
The sales price, for Dwelling styles with finished levels, is higher than the corresponding sales price for unfinished levels, such as The sales price for two and a half level Finished 2.5Fin is generally higher than the sales price for unfinished two and a half building styles 2.5Unf, and similarly for other building types.
It is evident that the sales price is an increasing function of the overall material and finishes quality of the house such that the higher the overall quality, the higher the sales price. The overall condition varies with sales price, but generally, conditions 5 through to 9 have higher sales prices than conditions 1 to 4.
From the scatterplots generated, it is evident that newer house attracts higher sales price, This can be shown by the negative relationship between house age and sales price, with older houses having lower sales prices.
It is noticeable, that houses with roof materials involving wood, such as Wood shakes and wood shingles WdShakes, WdShngl have higher sales prices than the rest. From the boxplot, it appears that all roof styles have fairly equal sales price, this can also be shown by Tukey’s Honesty Significance Distance test which reveals that the majority of the different roof styles are not significantly different from one another.
The sales price associated with Exterior coverings such as stone, ImStucc-Imitation structure, and VinylSd-Vinyl siding is higher than the other sales prices. This is because the initial construction cost of such houses and the durability of the materials such as stones give a house an edge on the sales price.
It is evident that houses with masonry veneer type Stone, have higher sales price than the rest of the types. There is also a positive linear relationship between Masonry veneer area and sales price, and the correlation coefficient is 0.47.
From the above plot, we see that the external quality and condition of material is significant in explaining the sales price of houses. Houses with Excellent (Ex) external quality and material have higher sales prices than the rest, with houses having poor external quality and condition fetching low sales prices.
The type of foundation used for a particular house has a significant effect on its sales price. Based on the above plot, and the analysis of the variance test, we see that houses with Poured concrete PConc and stone foundations had significantly higher sales prices compared to the rest. We could argue that these houses are more durable due to the foundation material used, and therefore more valuable than the rest which used lesser durable foundation materials.
It can be shown that the sales prices are directly affected by basement quality and condition. Basements with Excellent quality and condition have higher sales prices than the rest of the basements. Poor quality and condition basements have the lowest qualities and conditions. The sales price is an increasing function of the quality and condition of basements
The garden level walls also have a significant effect on the houses’ sales price. Houses with Good exposure have higher prices than the rest of the houses, with the lowest house sales prices having No basement and No exposure.
There is an increase in sales prices with an increase in rating of basement finished area, with Good living quarters GLQ having higher sales prices than the rest and houses with no basements having the lowest sales prices.
The plot above shows that the houses with no basements have lower sales price than the rest of the basement types. There is a weak linear relationship between the finished area of basements and the sales price. This is also evident in Tukey’s test of significance which shows that only a few comparisons, especially relating to no basement types are significant while the rest are not significant. Therefore the sales price is not affected by basement type two categories and finished square feet.
As evident from this plot, there is a strong linear relationship between the total basement area and the sales price, with a correlation of 0.6453. With the total finished basement area, which we get by subtracting the unfinished basement area from the total basement area, there happens to be a quadratic relationship between the sales price and the finished basement area, with a correlation of0.4745. Thus in the modelling process, it may be wise to incorporate the quadratic term of the finished basement area.
From the above plots, we can see that homes with Gas related heating types such as GasW-Gas hot water and GasA-Gas forced warm air to have higher sales price than the rest of the houses. It is also evident that the quality and condition of the heating types affects the sales price significantly. The better the heating quality and condition the higher the sales price.
It is also evident that houses with central conditioning systems have relatively higher sales prices than the corresponding houses with no central conditioning. Hence the central conditioning variable is significant in determining sales prices.
The electrical systems installed in houses also have a significant effect on sales prices of homes with homes having the SBrkr - standard circuit breakers fetching relatively higher sales prices, and the homes with the Mixedelectrical systems having the least sales prices
Based on the above plots, it is evident that there is a strong positive linear relationship between the sales prices and the area size of the first floor of the homes, with the correlation being 0.6227. Generally, it is also evident from the first plot that houses with second floors tend to have higher sales prices than the rest. The relationship between the second-floor area and the sales prices of homes seems to be a quadratic relationship, with houses that don’t have a second floor being given a 0 sq. feet second-floor area. Thus in our modelling, it would be fit to include the quadratic term of the second-floor area.
It can be shown that houses with a higher low-quality finished area tend to have lower prices than the rest, with the highest sales price being witnessed in houses with 0 square feet of low-quality finished area.
To analyse this variable, I had to omit outliers based on houses with above grade living area of greater than 4000 square feet, since that is unreasonable for homes in this dataset From the plot above, there is a strong positive relationship between the sales price of homes and their above grade living area in square feet.
Based on this variable Bedrooms above basement level, we note that the sales prices across all homes are fairly equal. This can also be backed by Tukey’s Honesty significance distance which shows that majority of the levels in Bedrooms above basement levels have similar sales prices. Thus we could omit this variable during our modelling process as it is not significant.
From the above analysis, we can see that the number of kitchens above grade is not a significant factor in determining the sales prices of houses, since the probability values reported by Tukey’s Honesty significance test shows that the various numbers of Kitchens generally have equal sales prices. We could also attempt to analyze the interaction effect between the two variables, but due to the class imbalance for instance: there are no Fair quality kitchens in homes with 2 kitchens above grade. Such problems could affect the performance of our modelling process, and it will be wise to avoid the interaction term. It is important however to note that the kitchen quality is an important factor in determining the sales prices of homes, with excellent Kitchens fetching the highest prices, while Poor quality kitchens attracting very low sales prices. Thus in the modelling process, we would omit the Number of kitchens above the grade variable but include the Kitchen quality variable.
There is a positive linear relationship between the number of rooms above grade and the home sales prices, with higher sales prices being witnessed in homes with a higher number of rooms above grade.
There is no visible relationship between the home functionality rating and the sales price since the sales prices are generally the same for all categories of home functionality. Tukey’s test also shows that there are no significant differences in sales prices within the home functionality categories.
Based on the above analysis of the number and quality of fireplaces, we note that both the number of fireplaces and the quality are significant variables in explaining the variations in homes sales prices. Homes with excellent fireplaces have higher sales prices than the rest, and homes with poor quality fireplaces fetching lower sales prices. Homes with many kitchens have higher sales prices than the subsequent homes with lower kitchen numbers. The Tukey significance test shows that the number of fireplaces significantly affects sales prices. We also not that the fireplace quality affects the sales prices significantly as shown by Tukey’s test.
The garage location also has a significant relationship with the sales prices, with homes having No garage, Detached garages and Carports having lower sales prices than homes having Attached garages, Garages in basements and Built-in garages which have significantly higher sales prices.
I transformed the variable GarageYrBlt-The year the garage was built into a variable GarageAge, which I obtained by subtracting the GarageYrBlt Variable from the current year 2021, such that houses with no garages, which had missing values are given the value 0 years of Garage age. There is a significant negative relationship between the garage age and the sales prices of homes, with homes having newer garages fetching higher sales prices than homes having older garages. It is also evident that attached garages and built-in garages have higher sales prices overall than detached garages and no garages homes.
There is a significant relationship between the interior finish of the garage and the sales price, with finished garages having the highest prices than unfinished garages and homes with no garages.
From the plot above, we can see that the sales prices of homes are an increasing function of the car capacity of their garages, with homes having no garages, being given the value 0. It is surprising to note that homes with four garages have lower than expected sales prices… The garage size in square feet is also having a significant positive linear relationship with the homes sales price, with a correlation of 0.6439.
From the above plot, we see that the Garage quality is significant in determining the homes sales prices, with excellent quality and Good quality garages having higher sales prices than average, poor and no garages quality. It is also important to note that garage condition has a significant impact on sales prices, such that the sales price of homes is an increasing function of the garage conditions.
From the faceted plot above, it is observable that there is a positive linear relationship between the area of the wood deck and the open porch, with homes having no wood deck and open porch having a value of 0. As for the relationship between the enclosed porch area, three-season porch, screen porch and pool area, the relationship is slightly a weak and negative one since the majority of homes lack these porches especially the pools, thus it might be wise to omit such variables in our modelling process since they have no significant effect on the sales prices of homes.
The pool quality variable also has a small noticeable effect on sales prices, a glance at the Tukey Honesty significance test shows that there are no significant differences in sales prices across the different pool qualities, thus it is important that we should omit the pool area, and pool quality variables in our modelling process.
From the analysis of fence quality, the fence quality GdPrv-Good privacy fetch higher significant sales prices, although it is surprising to note that homes with no fences at all also have higher sales prices as can be noted by their median value which is slightly equal to the median sales prices of homes with Good privacy. This can be attributed to other factors and feature within these homes and not the fact that they lack fences. The Tukey test reveals that sales prices do not differ within the majority of fence qualities.
From the above plot, we note that the sales prices for homes with Tennis courts are significantly higher than the rest, though the homes are very few. The majority of homes have no None miscellaneous features, though their sales prices are still higher probably due to other factors. From the scatter plot above, we see that homes with shed also have higher significant sales prices.
From this plot, it is evident that the year sold is not significantly related to the sales prices. Thus it would be wise not to include it in our modelling process. It is also evident that across all years, higher sales prices are witnessed mid-month.
There is a higher sales price in sales types Con and New as compared to the rest of sales types, with the lowest sales types being in sales-type Other.
The partial sales condition seems to be associated with a higher sales price. This is because the Partial sale condition represents relatively new homes.
In the modelling process, we aim to use the training data to train a linear model, which will be used to predict the sales prices in the observations found in the testing dataset. Since we made a few transformations on the training set concerning certain variables, it would also be important to include those same transformations in the test set, for example, transformations that involved converting years to ages and filling in missing values in the test dataset.
We will create a linear regression model, including variables that we deemed significant from our exploratory data analysis section, and omit the insignificant variables from the model. It would be wise although to implement the full model (a model with all variables, whether significant or not), to benchmark.
The first model is the full model, which contains all the variables present in the training dataset, including the variables which we saw were insignificant. This model also contains all terms linear, no quadratic or square root transformations are in the model.
The second model omits all the insignificant variables, the variables which we concluded were insignificant based on the analysis of variance tests.
The third model includes all the significant variables alone, and some terms such as quadratic fits which we saw had a stronger correlation to the sales prices than their subsequent linear fits.
The fourth model is an extension of the third model while omitting all the insignificant variables, - here the insignificant variables are those which were reported as having higher probability values, above 0.05 level.
The models above have higher adjusted R squared performance, with the highest performance witnessed in the full model Mod0, followed by the third model Mod2, and lastly the fourth model Mod3. The higher performance in the full model can be attributed to the fact that we are using many variables to fit the model, although most of them are insignificant. This violates the good rule of thumb in fitting linear models, in that we must always ensure that our models are parsimonious(having as few variables as possible, while still explaining the majority of the variation in the response variable sales price). Thus we will favour the fourth model, which has a higher R^2 performance while excluding many insignificant variables.
Here we will use the Rooted mean square metric to compare the two desired models Mod2 and Mod3, as follows:
The RMSE for model 3 is a bit lower than the RMSE for model 2, thus this entails that model 3 is a better model fit than model 2.
We can gain further insight into the fit of model 3 by looking at various tests, such as the Shapiro Wilk test for the normality of the model residuals, and the Breusch pagan test for testing the heteroscedasticity assumption of linear regression models.
From the look of the histogram and the Normal Quantile-Quantile plot, we see that the residuals are slightly normal especially in the histogram, however, the Shapiro Wilk test for normality shows that the Normality assumption is violated since the probability value associated with the test is less than the 0.05 level. This could be that there are some outliers in the model.
The low probability value also associated with the Breusch Pagan test, indicates that heteroscedasticity is an issue or a problem.
The probability values associated with the RESET test, indicate that there is a problem associated with the functional form of the model, such as some explanatory variables are missing, and such.
To improve this model, we could try removing outliers and high leverage points as shown below:
We can again, test for the Rooted mean square of the model with no outliers and leverages, to see if there is an improvement.
From the look of the analysis, there is an increasing RMSE for the leverage removed model, which tells us to stick with the outlier removed model only for our predictions, probably because removing the high leverage points does not increase the fit of the model on the data.
In this analysis, and modelling that I have conducted, the findings are that the sales prices are affected by far fewer variables than the ones we were supplied within the training dataset and testing set. the variables affecting the sales prices significantly are shown below:
To answer the analysis questions in the Problem identification section.
The best rooted mean square achieved was with the outlier removed model 3 and the rmse was 107364.9. I obtained this RMSE, by:
I can admit that a linear model may not be the best model for determining the best factors affecting the sales price of the homes, this is because even my best model failed the tests for all the assumptions of fitting linear models such as:
This could be because some explanatory variables are missing that could aid in understanding the sales price better such as proximity to hospitals nearby, the general security level of the area, proximity to industrial sites, schools and such. It could also be that the relationships between the sales prices and all other factors are not linear, but could be non-linear, thus linear regression fails. Non-linear relationships could be learnt using models such as neural networks and support vector machines.
Aggarwal, R., & Ranganathan, P. (2017). Common pitfalls in statistical analysis: Linear regression analysis. Perspectives in clinical research, 8(2), 100.
Austin, P. C., & Steyerberg, E. W. (2015). The number of subjects per variable required in linear regression analyses. Journal of clinical epidemiology, 68(6), 627-636.
Bertsimas, D., & King, A. (2016). OR forum—an algorithmic approach to linear regression. Operations Research, 64(1), 2-16.
Darlington, R. B., & Hayes, A. F. (2017). Regression analysis and linear models. New York, NY: Guilford.
Fox, R., & Tulip, P. (2014). Is housing overvalued?. Available at SSRN 2498294.
Hayes, A. F., & Montoya, A. K. (2017). A tutorial on testing, visualizing and probing an interaction involving a multi categorical variable in linear regression analysis. Communication Methods and Measures, 11(1), 1-30.
Phan, T. D. (2018, December). Housing price prediction using machine learning algorithms: The case of Melbourne city, Australia. In 2018 International Conference on Machine Learning and Data Engineering (iCMLDE) (pp. 35-42). IEEE.
Zhang, Y., & Dong, R. (2018). Impacts of street-visible greenery on housing prices: Evidence from a hedonic price model and a massive street view image dataset in Beijing. ISPRS International Journal of Geo-Information, 7(3), 104.
Do you wish you could receive law assignment help from top lawyers in the country? Now you can hire only the best industry professionals on Myassignmenthelp.co.uk. In addition, our experts provide you with in-depth dissertation help to overcome the usual hassles of writing assignments.
Additionally, you can find managers from reputable companies going out of their way to provide reliable management assignment help. Their guidance is crucial when preparing for examinations and writing research papers. Hence, whenever you wonder, "Can't someone do my assignment in the UK?” don’t settle for anything less than the best academic stalwarts on Myassignmenthelp.co.uk.
Upload your Assignment and improve Your Grade
Boost Grades