MarzleyTech Learn

Home / Learn / Artificial intelligence (AI): from zero to pro / Level 4: Neural networks explained (and one you can train)

Level 4: Neural networks explained (and one you can train)

Neural networks power modern AI: speech recognition, image recognition, translation and chatbots. They're loosely inspired by brain cells, but at heart they're just lots of simple maths units connected together, trained with the gradient descent you saw earlier.

One artificial neuron

A neuron takes inputs, multiplies each by a weight (how important it is), adds a bias, and passes the total through an activation function that decides how strongly it "fires".

inputs      weights
 x1 ──── w1 ──┐
 x2 ──── w2 ──┼──► sum = w1·x1 + w2·x2 + bias ──► activation ──► output
 x3 ──── w3 ──┘

A common activation is the sigmoid, which squashes any number into 0 to 1 (like a probability).

Train a neuron to learn a rule

Let's teach one neuron the rule "approve a delivery request only if the customer has paid AND the address is confirmed" (logical AND), purely from examples.

Python · runs live in the interactive lesson
import math
import random

random.seed(1)

def sigmoid(z):
    return 1 / (1 + math.exp(-z))

# (paid, address_confirmed) -> approve?
examples = [((0, 0), 0), ((0, 1), 0), ((1, 0), 0), ((1, 1), 1)]

w1, w2, bias = random.uniform(-1, 1), random.uniform(-1, 1), 0.0
lr = 1.0

for epoch in range(3001):
    total_error = 0
    for (x1, x2), target in examples:
        out = sigmoid(w1 * x1 + w2 * x2 + bias)          # forward: make a prediction
        error = out - target
        total_error += error ** 2
        grad = error * out * (1 - out)                   # backward: how to change the weights
        w1 -= lr * grad * x1
        w2 -= lr * grad * x2
        bias -= lr * grad
    if epoch % 1000 == 0:
        print(f"epoch {epoch}: error {total_error:.4f}")

print(f"weights {w1:.2f}, {w2:.2f}, bias {bias:.2f}")
for (x1, x2), target in examples:
    print((x1, x2), "->", round(sigmoid(w1 * x1 + w2 * x2 + bias), 3), "expected", target)

The neuron starts with random weights and learns the AND rule: outputs near 1 only when both inputs are 1.

Why we need layers: the XOR problem

One neuron can only draw a straight line between categories. Some patterns need more. XOR ("exactly one of the two") can't be separated by one straight line. Add a hidden layer of neurons, and the network can learn it:

Python · runs live in the interactive lesson
import math
import random

random.seed(3)
sig = lambda z: 1 / (1 + math.exp(-z))

data = [((0, 0), 0), ((0, 1), 1), ((1, 0), 1), ((1, 1), 0)]   # XOR
H = 3                                                          # hidden neurons

w_in = [[random.uniform(-1, 1) for _ in range(2)] for _ in range(H)]
b_in = [0.0] * H
w_out = [random.uniform(-1, 1) for _ in range(H)]
b_out = 0.0
lr = 0.8

for epoch in range(10001):
    for (x1, x2), target in data:
        # forward pass
        hidden = [sig(w_in[j][0] * x1 + w_in[j][1] * x2 + b_in[j]) for j in range(H)]
        out = sig(sum(w_out[j] * hidden[j] for j in range(H)) + b_out)
        # backward pass (backpropagation)
        d_out = (out - target) * out * (1 - out)
        for j in range(H):
            d_hidden = d_out * w_out[j] * hidden[j] * (1 - hidden[j])
            w_out[j] -= lr * d_out * hidden[j]
            w_in[j][0] -= lr * d_hidden * x1
            w_in[j][1] -= lr * d_hidden * x2
            b_in[j] -= lr * d_hidden
        b_out -= lr * d_out

for (x1, x2), target in data:
    hidden = [sig(w_in[j][0] * x1 + w_in[j][1] * x2 + b_in[j]) for j in range(H)]
    out = sig(sum(w_out[j] * hidden[j] for j in range(H)) + b_out)
    print((x1, x2), "->", round(out, 3), "expected", target)

This is a real (tiny) multi-layer neural network, trained with backpropagation: errors flow backwards from the output, telling every weight how to change.

From tiny to huge

Our XOR networkA large language model
Parameters (weights)13Billions to trillions
Training examples4Trillions of words
Training timeUnder a secondMonths on thousands of specialised chips (GPUs)
Same core idea?Forward pass, measure error, backpropagation, gradient descentYes

Types of neural networks

NetworkGood forExample
Feed-forward (MLP)Tables of numbersCredit scoring
Convolutional (CNN)ImagesDetecting crop diseases from leaf photos, reading number plates
Recurrent (RNN, LSTM)Sequences (older approach)Early speech recognition
TransformerLanguage, and now images, audio and videoChatGPT, Claude, Gemini, translation

The transformer (introduced in 2017) is the architecture behind today's language models. Its key idea, attention, lets the model look at every word in the text and decide which words matter most for understanding each other word. That's the subject of the next lesson.

Tools professionals use

PyTorch and TensorFlow/Keras build and train neural networks, usually on GPUs (free GPU time is available in Google Colab and Kaggle notebooks). Hugging Face hosts thousands of ready-made models you can download and fine-tune.

Check yourself

  1. What does each input of a neuron get multiplied by?

    Show answer

    a weight

  2. Which activation function squashes numbers into the range 0 to 1?

    Show answer

    sigmoid

  3. Which simple pattern can't be learned by a single neuron, needing a hidden layer?

    Show answer

    XOR

  4. What is the algorithm called that sends errors backwards to update every weight?

    Show answer

    backpropagation

  5. Which neural network type is best known for images?

    Show answer

    CNN

  6. Which architecture is behind ChatGPT and Claude?

    Show answer

    transformer

Lesson 13 of 17 in Artificial intelligence (AI): from zero to pro · Printable course notes