Analysis of Homes in Detroit Using Detroit Open Data Portal and FOIA

Download Solution Order New Solution

Assignment Task

You have been tasked with undertaking a multi-part analysis of homes in Detroit, Michigan. You are provided with a database to facilitate this analysis. This database was constructed from the Detroit Open Data portal and numerous FOIA requests. More information is included in the database section below. Note that the database must be downloaded.

Assignment 

  • Section A: Conduct an exploratory data analysis of homes in Offer an overview of relevant trends in the data and data quality issues. Contextualize your analysis with key literature on properties in Detroit.
  • Section B: Use to conduct a sales ratio study across the relevant time period. Note that is designed to produce Rmarkdown reports but use thedocumentation and insert relevant graphs/figures into your report. Look to make this reproducible since you’ll need these methods to analyze your assessment model later on. Detroit has many sales which are not arm’s length (sold at fair market value) so some sales should be excluded, but which ones?
  • Section C: Explore trends and relationships with property sales using simple regressions
  • Section D: Explore trends and relationships with foreclosures using simple regressions

Part 2

Objective: Now that you have a decent understanding of the landscape in Detroit, create a new file (part_2.Rmd) which builds upon part_1 in the html report Rmarkdown style.

Part A

Create an ‘introduction’ to your report. Generally, only include stylized output (do not use base R print). This could mean using stargazer to show regressions, DT::datatable to show data.frames, and adding titles/labels to plots. Your introduction should include:

  1. Brief background (2-3 sentences) on issues in the Detroit assessment space

  2. 3 to 4 graphs with descriptive captions which include information on sale price, assessment accuracy, foreclosures, and outliers. Generally focus on single family homes and arm’s length transactions. While it is notable that so many properties are sold for small amounts, we typically only want to look at properties which are class 401, taxable (e.g. assessed over 2000 or so), and sell above $4,000.

Part B

We have two separate (but very related) problems we want to model. First, we want to find a way to identify if a home is likely to be overassessed in a given year. We will analyze homes and assessments from 2016. We will use tidymodels to create a workflow.

  • Create your workflow
  • Add to your workflow a classification model
  • Add to your workflow a recipe of preprocessing steps. Use 2016 sales and assessments with the parcels property characteristics (note that we only know if a home was overassessed if it sold). Create a classification metric of overassessment based on properties which sold and use this as your dependent Explain how you decided to construct this metric and how many classes it has.
  • Create testing/training data and evaluate your model using the classification metrics from tables 3 and 8.4 from the textbook and the classification probability metric ROC curves.

Part C

Second, building off of the workflow from part B. Create a second model to create your own 2019 assessments. (Note that I am choosing this year to avoid impacts from the pandemic and data quality issues. You may, if you’d like, create 2022 assessments. Limited sales data is released here .)

  • Create your workflow
  • Add to your workflow a model
  • Add to your workflow a recipe of preprocessing Use sales and assessments from before 2019 with the parcels property characteristics.
  • Create testing/training data and evaluate your model using numeric metrics RMSE and MAPE.

This IT Computer Science has been solved by our PhD Experts at My Uni Paper.

Get It Done! Today

Country
Applicable Time Zone is AEST [Sydney, NSW] (GMT+11)
+

Every Assignment. Every Solution. Instantly. Deadline Ahead? Grab Your Sample Now.