Highlights
Task
What is the correlation between weather and rental count?
Which activity is the most happening?
What is the pattern/trends?
Does burning of coal during winter affect the PM value?
Did the sensors in the living room reflect the relevant activities?
What information can we gather from the temperature readings at different rooms?
Which day of the week shows highest ….?
Are there any common patterns of activities happening in a day?
Background of the data
Krakow is placed at the 8th position out of 575 cities for PM2.5 level and 145th position out of 1100 cities for PM10 level in a study completed by the World Health Organisation.
A factor that might be causing the severe air quality could be due to usage of fossil fuel such as coal for heating of households, which worsens during months that are colder. After researching on the causes of air pollution in Krakow, another major factor is the high usage of personal transportation instead of utilising the public transportation to commute in and out of the city, cars produce high amount of PM2.5 and PM10 particles.
In recent years, to combat the issue of air quality, The city has partnered with Airly to install 56 low-cost air quality sensors that is installed behind the windows of households across Krakow to boost the current 8 monitoring stations that is operated by the State. This implementation thus allows the local government to pinpoint the issue to certain regions more accurately which allows them to combat the issue more accurately.
In this project, I have been tasked to use data from the month of June to August 2017 to perform data cleaning and transformation, so as to perform data visualisation and lastly, to run a linear regression model to predict the PM2.5 level based on the data provided.
Observations and variables of the data
I have received a file that has 12 months of data from 2017 and I have utilised the files from the month of June to August 2017 and the sensor location file to perform this project.
In each file, data are collected across 56 stations, on an hourly basis, with 6 variables being measured by each station are: temperature, humidity, pressure, PM1, PM10 & PM25. Another variable that is taken down is the timing of the measurement.
In total, there are 337 variables in total in each of the files. In June, there should be a total of 720 observations (30 days x 24 hours) and in July and August, there should be a total of 744 observations (31 days x 24 hours).
Data types
Variable \ Stage Prior to conversion After conversion
[Station ID] PM1 Integer No Change
[Station ID] PM2.5 Integer No Change
[Station ID] PM10 Integer No Change
[Station ID] Temperature Integer No Change
[Station ID] Humidity Integer No Change
[Station ID] Pressure Integer No Change
UTC Time String Converted to Date & Time
Data Cleaning & Transformation
Missing values
Variable \ Month June 2017 July 2017 August 2017
Station 169 PM1 4 1 0
Station 169 PM2.5 4 0 0
Station 169 PM10 4 0 0
Station 169 Temperature 4 0 0
Station 169 Humidity 4 0 0
Station 169 Pressure 4 0 0
Variable \ Month June 2017 July 2017 August 2017
Station 204 PM1 4 0 0
Station 204 PM2.5 4 0 0
Station 204 PM10 4 0 0
Station 204 Temperature 4 0 0
Station 204 Humidity 4 0 0
Station 204 Pressure 4 0 0
Variable \ Month June 2017 July 2017 August 2017
Station 222 PM1 4 6 3
Station 222 PM2.5 4 5 3
Station 222 PM10 4 5 3
Station 222 Temperature 4 5 3
Station 222 Humidity 4 5 3
Station 222 Pressure 4 5 3
I have connected the Statistics function to File Reader, and I execute and open view to reach the above part to identify stations with the least missing values so that I am able to meet the criteria of a minimum of 700 observations for each station I choose.Cleaning of dataTo remove other columns of data quickly, I used column filter to filter out the other columns so that only the UTC Time and the measurements from the 3 stations remain behind. I did the same steps for all 3 sets of data.To remove rows with missing values efficiently, I used missing values to weed out rows with missing value by applying the option of remove row for both string and integer. I repeated the same steps for the other 2 sets of data.Identification and converting of wrong data typeMost of the data types are correct (Integer) except for UTC Time which the data type is string, thus we have to use string to date&time to convert UTC Time to the correct data type. I selected date&time at the new type, and selected guess data type and format to get the date format. I repeated the same procedure for all 3 sets of data as the UTC Time is string for all 3 sets.
This IT Assignment has been solved by our IT Experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+Students in Australia, UK & US by helping them to score HD in their academics. Our Experts are well trained to follow all marking rubrics & referencing style.
Be it a used or new solution, the quality of the work submitted by our assignment Experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.
© Copyright 2026 My Uni Papers – Student Hustle Made Hassle Free. All rights reserved.