Small Starts for Big Changes
Today's July 5th, 2026 at 5:40 PM. I just woke up from a mid-day nap and cracked a soymilk drink. I'll finally stop procrastinating and start Andrej Karpathy's Neural Networks: Zero to Hero YouTube series.

Building Micrograd
7:24 PM draw_dot errors. Typos and restarting notebook kernel
7:32 PM Supposedly learned:
- build out mathematical expressions
- they are scalard value
- Forward pass
- Multiple inputs going into mathematical expression that results into a single output
Next:
- Backpropagation: start at the end, reverse intermediate values, calculate the gradient
- every single value, calculate the derivative with respect to L and so on
- Loss function with respect to the weights of a neural network
- Need to know how weights are impacting the loss function
- leaf nodes will be weights to neural net
- weights iterated on with the gradient information

Yeah uh.. What are derivatives?
7:50 PM Starting 3Blue1Brown's video: The paradox of the derivative
After Yeah I still don't understand any of it.. Fancy Google search with Gemini: So calculus-wise, with the information from 3Blue1Brown, a derivative is the rate of change over a very small amount of time, a the best constant approximation around a point Neural network-wise, it's used for caluclating the loss function -- how much a tiny tweak of a weight will affect the final output, which allows for other functions or finessing like gradient descent. Hopefully that's right.
PS: trust actual professionals more than me
Back to Karpathy
8:13 PM, 34:42 / 2:25:51 of building micrograd
Trying to understand what he is doing. Bunch of values and variables.
Snack and Break Intermission -- Karpathy Restart
9:20 PM 42:55 / 2:25:51 of building micrograd
Manual Backpropagation
9:41 PM, 50:30 / 2:25:51 of building micrograde
What?

Session End
9:50 PM, 51:52 / 2:25:51 of building micrograd
Manual backpropagation example 1 complete. Why does DL have to be so complex..