Introduction

This is an R Markdown document for the report for the Course Project in the Practical Machine Learning course on Coursera.

Data

The data for this assignment, as is common with Machine Learning, was split into testing and training sets and can be downloaded from the course web site:

The data for this assignment was generously provided from this source (Velloso et al. 2013)

The purpose of this assignment is to predict what type of exercise (the “classe” variable in the dataset) was being performed based on activity tracker information.

Loading and pre-processing the data

The training and testing data was loaded from the provided CSV files.

After looking at the training set, the X columns was removed as it is just the index, the user_name column was removed for possible bias in the model, all timestamp columns were removed to remove time as an influence on the model, and the new_window and num_window columns were also removed to eliminate possible bias from those columns on the models.

The training and test data sets provided also contain extensive amounts of NA values. However, some columns were found to have blanks ("") instead of NA values. A further analysis of these columns revealed that some (like kurtosis_picth_forearm) had very few valid values and were assessed to have limited predictive value when they did contain valid values. Columns that were significantly NA or blank were removed. The same pre-processing was applied to the testing set.

As a result the following columns were kept for further analysis:

##  [1] "roll_belt"            "pitch_belt"           "yaw_belt"            
##  [4] "total_accel_belt"     "gyros_belt_x"         "gyros_belt_y"        
##  [7] "gyros_belt_z"         "accel_belt_x"         "accel_belt_y"        
## [10] "accel_belt_z"         "magnet_belt_x"        "magnet_belt_y"       
## [13] "magnet_belt_z"        "roll_arm"             "pitch_arm"           
## [16] "yaw_arm"              "total_accel_arm"      "gyros_arm_x"         
## [19] "gyros_arm_y"          "gyros_arm_z"          "accel_arm_x"         
## [22] "accel_arm_y"          "accel_arm_z"          "magnet_arm_x"        
## [25] "magnet_arm_y"         "magnet_arm_z"         "roll_dumbbell"       
## [28] "pitch_dumbbell"       "yaw_dumbbell"         "total_accel_dumbbell"
## [31] "gyros_dumbbell_x"     "gyros_dumbbell_y"     "gyros_dumbbell_z"    
## [34] "accel_dumbbell_x"     "accel_dumbbell_y"     "accel_dumbbell_z"    
## [37] "magnet_dumbbell_x"    "magnet_dumbbell_y"    "magnet_dumbbell_z"   
## [40] "roll_forearm"         "pitch_forearm"        "yaw_forearm"         
## [43] "total_accel_forearm"  "gyros_forearm_x"      "gyros_forearm_y"     
## [46] "gyros_forearm_z"      "accel_forearm_x"      "accel_forearm_y"     
## [49] "accel_forearm_z"      "magnet_forearm_x"     "magnet_forearm_y"    
## [52] "magnet_forearm_z"     "classe"

Setting up cross validation

The training set is very large (19,622 rows) while the final testing set is very small (20 rows). As a result, predicting the out of sample error would best be done by using cross-validation to perform testing on a subset of the original “training” set.

K-fold cross-validation was chosen with k=5. A known seed was set before setting the trainControl. In addition, the “training” set was partitioned into a “training” set and a “validation” set, so that the “validation” set could be used to validate the model before applying to the “testing” set.

Model Execution

A random forest was selected for the analysis method in the hopes that it would provide the best combination of processing time and effectivness in prediction. A known seed was set before running the model.

As hoped, the model’s OOB estimate error rate is a mere 0.49%.

Call:
 randomForest(x = x, y = y, mtry = param$mtry) 
               Type of random forest: classification
                     Number of trees: 500
No. of variables tried at each split: 2

        OOB estimate of  error rate: 0.49%
Confusion matrix:
     A    B    C    D    E  class.error
A 5020    2    0    0    0 0.0003982477
B   11 3401    6    0    0 0.0049736688
C    0   19 3058    3    0 0.0071428571
D    0    0   39 2854    2 0.0141623489
E    0    0    0    4 3243 0.0012319064

And predicting against the validation set confirms that the predictions are accurate. In fact, not a single prediction is missed! Note: The validation set appears to range from 0 to nearly 20,000 because of the way that the validation set was created as part of the training set. The original row.names values from the training set carries over, such that the row.names values range from 8 to 19620, even though there are only 1765 rows in the set.

##      
## pred2   A   B   C   D   E
##     A 517   0   0   0   0
##     B   0 322   0   0   0
##     C   0   0 317   0   0
##     D   0   0   0 276   0
##     E   0   0   0   0 333

Summary

By assessing the data in advance, recognizing that certain columns had limited to no predictive value, and pre-processing the training dataset in order to remove the identified columns, the resulting random forest model has an extremely high accuracy rate with an OOB expected error of only 0.49%! Further, by splitting the provided training set into a training set and a validation set, the author was able to sanity check and confirm that the error rate was indeed low, as no errors were found among the validation set.

As well, this pre-processing and elimination of these unnecessary columns were essential in not only reducing the processing time of the random forest model, but also of preventing the author’s computer from becoming unusuable during the model’s training.

References

Velloso, E., A. Bulling, H. Gellersen, W. Ugulino, and H. Fuks. 2013. “Qualitative Activity Recognition of Weight Lifting Exercises.” Proceedings of 4th International Conference in Cooperation with SIGCHI (Augmented Human ’13).