NN: More Information

Neural Network

The Neural Network algorithm is a little different from the other ones we've worked on. It is a computer program that tries to copy how a human brain works. It's composed of 3 layers: the input layer, used to introduce data, the output layer, where our result comes from and the hidden layer. This last one is where all the computing is done, and we'll extensively explain it later, since it can contain several layers itself. As a general rule of thumb, the more layers our hidden layer has, the better our neural network will be.

Every layer has nodes, which are, in our comparison of a neural network with a brain, neurons. Inside each of them there's a piece of information. And just as neurons need synapse to communicate this information to each other, our nodes will have channels, acting as a bridge that will be used to convey information.

However, not all the pieces of information stored in the nodes are equally important to our problem. Therefore, we will have to assign each node a value based on how important it is. These values are known as weights, and the more weight a node has, the more the result will be influenced by its information.

Every node from our input layer is connected to some nodes of our first hidden layer, and we transmit the information to them. However, we must be aware that we are working with weights, so we will have to calculate the weighted value of the node, as seen in the following image.

When the nodes in the hidden layer receive their calculated "score", we have to check if this score is high enough to keep the information going into deeper layers. This is why we need an "activation function", which is only an equation that tells us if the information in the node must continue deeper into the neural network. Some nodes will activate, and some will not. The ones that don't activate will automatically be weighted with a value of 0, and so will not affect our further calculations in the algorithm. The ones that activate, however, will repeat the same process that we did before with the input layer nodes: They are linked to the next layer of our network, each with a weight, and we will again calculate the value of the receiving nodes, based on the values we have in our current layer.

To summarize, we could say that the working of a neural network comes from having a layer full of nodes with values, that we use to calculate the values of the nodes of the next layer. This process is repeated for deeper layers until we reach the output layer, and every node in the output layer receives a value between 0 and 1. This value is a probability.

Like we did with the random forest, we need to check if our neural network is working correctly, so we have to check with some examples if our predictions are right. If they're not, that means that the weights of some nodes are off, and we must try a new value on them. We keep looking for new values until we reach a high rate of accuracy in our network. The process of training a neural network is long, and the more layers and nodes we add to it, the longer it will take. This means that a bigger neural network will result in higher accuracy and better predictions, but it will also need more time to be fully trained. We must find a balance where the highest accuracy possible is achieved without having so many layers that we lack processing power to train it.