What is a Training Algorithm?
After reducing the dimensionality of the data, we will be ready to choose how to train the machine. The process of teaching the computer how to solve a specific step is what we call Machine Learning. In this project we will offer you 4 different ways to teach the computer how to diagnose a new patient when they come in: Support Vector Machine, Random Forest, Neural Network and Generalized Linear Model. As we did with the second step, we will provide an explanation that will go over the workings of every one of these algorithms. This step is much more diversified than the one before, since each algorithm does the training in its own way. The objective, though, is the same for all of them: to get the most precision and accuracy possible in our machine. And they all do it in very similar ways. The data you upload, after going through Radiomics and the Dimensionality Reduction, will be ready to train a machine, so we will separate it in two different sets of data.
The first and bigger one is what we call Training Data, and it will be used by the machine to learn. It will find patterns within the patients, ways to classify them in groups, separate them according to one thing or another, and overall teach itself to diagnose them. Keep in mind, all the data that you upload will already have a diagnostic in place, so the machine will learn based on what the diagnosis for the Training Data are. If someone was to upload for training some magnetic resonance imagings with wrong diagnostics, the machine wouldn't learn properly.
The second set of data is the Test Data. Its job is to find out how well the machine was trained. When the Training Data have all gone through the algorithm that you chose, the machine will be able to predict a new patient's diagnostic based on their magnetic resonance imaging, but we will not know how precisely. Here is where the Test Data come in: we will make it go through our machine, hiding the diagnostic from it. The machine will make a prediction based on what it has learnt, and we will compare. If the machine predicts "the patient is at risk of having a heart attack" and the original patient really had a heart attack, it will mean that the machine correctly predicted the result of the diagnose, and therefore its accuracy rate will be 1 for 1. When this is done with the entirety of our Test Data, we will just need to count: If the machine has correctly diagnosed 38 out of 40 patients from our Test Data, then we know that it is correct 95% of the time. This will conclude the second step.