Introduction to Tinygrad
Tinygrad is hard to read, even if you already live in autograd. Karpathy has said as much. This is my first pass at it: documentation, livestreams, and the parts of the source that actually matter, in one place.
Why Tinygrad?
It does not look like PyTorch, JAX, or TensorFlow, and that is the point. The stated mission is to democratize the petaflop: a compiler small enough that a handful of people can hold the whole thing in their heads.
What I like:
- Fully open source — including the hiring process, which is a story of its own
- No dependencies
- Multiple backends
- A genuinely small codebase
What I still don’t:
- Steep, and not kind to beginners
- The source is dense
- Docs lag; some examples are broken
- The API is not stable yet
That is enough throat-clearing. Here is the model.
Installation
Tinygrad is meant to be hacked. Install it editable so a change in the clone is a change in the import:
1 2 3 | |
The Tensor
The surface looks like PyTorch on purpose:
1 2 3 4 5 6 7 8 9 10 11 12 | |
Familiar API, different machine underneath.
The Devices
A Device is where a tensor lives and where kernels run: CPU, CUDA, METAL, and others. On Apple Silicon the default is METAL:
1 2 3 | |
Tinygrad picks the fastest backend it can see. You can override it:
1 2 | |
The same Python then runs on whatever you pointed it at.
Lazy Evaluation
Nothing has been computed yet:
1 2 3 4 5 6 7 8 9 10 11 | |
Tinygrad is lazy. A Tensor is a specification of work — a chain of UOps — not a buffer of results. Execution waits until you ask.
The UOP
UOps (micro-operations) are the IR. They are immutable and form a DAG. Each node has an op, a dtype, an argument, and sources.
1 | |
You get something like:
1 2 3 4 5 | |
That tree is a COPY from a BUFFER on the PYTHON device (the list [1, 2, 3, 4]) onto CPU. The UNIQUE id pins that buffer so nothing else can be confused with it.
flowchart LR
A["COPY<br/>dtypes.int<br/>arg=None"]
B["BUFFER<br/>dtypes.int<br/>arg=4"]
C["UNIQUE<br/>dtypes.void<br/>arg=0"]
D["DEVICE<br/>dtypes.void<br/>arg='PYTHON'"]
E["DEVICE<br/>dtypes.void<br/>arg='CPU'"]
A --> B
A --> E
B --> C
B --> D
Realization
realize() runs the spec:
1 2 | |
After that, the COPY is gone. You are looking at a BUFFER:
1 2 3 | |
flowchart LR
A["BUFFER<br/>dtypes.int<br/>arg=4"]
B["UNIQUE<br/>dtypes.void<br/>arg=1"]
C["DEVICE<br/>dtypes.void<br/>arg='CPU'"]
A --> B
A --> C
UNIQUE moved from arg=0 to arg=1: the data now exists as a new object on the target device.
Operations build computation graphs
Arithmetic is more graph:
1 2 | |
Scalar multiply becomes an explicit broadcast:
flowchart LR
A["MUL<br/>dtypes.int<br/>arg=None"]
B["BUFFER<br/>dtypes.int<br/>arg=4"]
C["UNIQUE<br/>dtypes.void<br/>arg=1"]
D["DEVICE (x2)<br/>dtypes.void<br/>arg='CPU'"]
E["EXPAND<br/>dtypes.int<br/>arg=(4,)"]
F["RESHAPE<br/>dtypes.int<br/>arg=(1,)"]
G["CONST<br/>dtypes.int<br/>arg=2"]
H["VIEW<br/>dtypes.void<br/>arg=ShapeTracker"]
A --> B
A --> E
B --> C
B --> D
E --> F
F --> G
G --> H
H --> D
The 2 is RESHAPEd and EXPANDed to (4,). Broadcasting is not a hidden NumPy trick here; it is a node the rewrite engine can see and fold.
1 | |
Smart Deduplication
UOps are immutable and globally unique, so identical specs are the same object:
1 2 3 4 5 6 7 8 | |
Realize one, and the other is already done:
1 2 3 4 5 6 | |
Automatic Differentiation
Gradients are the same machinery. For $y = x^2 + 3x + 1$:
1 2 3 4 5 6 | |
$$\frac{dy}{dx} = 2x + 3$$
At $x = 2$:
$$\frac{dy}{dx}\bigg|_{x=2} = 2(2) + 3 = 7$$
1 2 | |
Chain Rule
Tinygrad applies the chain rule without being asked. For $z = \log(x^2 + 1)$:
1 2 3 4 5 6 | |
- Let $z = \log(x^{2}+1)$. Set $g(x)=x^{2}+1$, so $z = \log(g(x))$.
- $\frac{dz}{dx} = \frac{1}{g(x)} \cdot g'(x) = \frac{2x}{x^2+1}$
- At $x=2$: $\frac{dz}{dx} = \frac{4}{5} = 0.8$
1 2 | |
Graph Optimization and Kernel Generation
The interesting part is the rewrite. Start with t + 3 + 4:
1 2 3 4 | |
flowchart LR
A["ADD (outer)<br/>dtypes.int<br/>arg=None"]
B["ADD (inner)<br/>dtypes.int<br/>arg=None"]
C["COPY<br/>dtypes.int<br/>arg=None"]
D["BUFFER<br/>dtypes.int<br/>arg=4"]
E["UNIQUE<br/>dtypes.void<br/>arg=8"]
F["DEVICE<br/>dtypes.void<br/>arg='PYTHON'"]
G["DEVICE (x5)<br/>dtypes.void<br/>arg='CPU'"]
H["EXPAND<br/>dtypes.int<br/>arg=(4,)"]
I["RESHAPE<br/>dtypes.int<br/>arg=(1,)"]
J["CONST<br/>dtypes.int<br/>arg=3"]
K["VIEW (x9)<br/>dtypes.void<br/>arg=ShapeTracker"]
L["EXPAND<br/>dtypes.int<br/>arg=(4,)"]
M["RESHAPE<br/>dtypes.int<br/>arg=(1,)"]
N["CONST<br/>dtypes.int<br/>arg=4"]
A --> B
A --> L
B --> C
B --> H
C --> D
C --> G
D --> E
D --> F
H --> I
I --> J
J --> K
K --> G
L --> M
M --> N
N --> K
Two ADD nodes. Then kernelize:
1 2 | |
Constant folding turns + 3 + 4 into + 7:
flowchart LR
A["ASSIGN (outer)<br/>dtypes.int<br/>arg=None"]
B["BUFFER (x0)<br/>dtypes.int<br/>arg=4"]
C["UNIQUE<br/>dtypes.void<br/>arg=10"]
D["DEVICE (x2)<br/>dtypes.void<br/>arg='CPU'"]
E["KERNEL 12<br/>dtypes.void<br/>arg=<SINK __add__>"]
F["ASSIGN (inner)<br/>dtypes.int<br/>arg=None"]
G["BUFFER (x5)<br/>dtypes.int<br/>arg=4"]
H["UNIQUE<br/>dtypes.void<br/>arg=9"]
I["KERNEL 5<br/>dtypes.void<br/>arg=<COPY>"]
J["BUFFER<br/>dtypes.int<br/>arg=4"]
K["UNIQUE<br/>dtypes.void<br/>arg=8"]
L["DEVICE<br/>dtypes.void<br/>arg='PYTHON'"]
A --> B
A --> E
B --> C
B --> D
E --> B
E --> F
F --> G
F --> I
G --> H
G --> D
I --> G
I --> J
J --> K
J --> L
The kernel AST after rewrite:
1 2 3 4 5 | |
flowchart LR
A["SINK<br/>dtypes.void<br/>arg=None"]
B["STORE<br/>dtypes.void<br/>arg=None"]
C["INDEX (output)<br/>dtypes.int.ptr(4)<br/>arg=None"]
D["DEFINE_GLOBAL<br/>dtypes.int.ptr(4)<br/>arg=0"]
E["SPECIAL (x3)<br/>dtypes.int<br/>arg=('gidx0', 4)"]
F["ADD<br/>dtypes.int<br/>arg=None"]
G["LOAD<br/>dtypes.int<br/>arg=None"]
H["INDEX (input)<br/>dtypes.int.ptr(4)<br/>arg=None"]
I["DEFINE_GLOBAL<br/>dtypes.int.ptr(4)<br/>arg=1"]
J["CONST<br/>dtypes.int<br/>arg=7"]
A --> B
B --> C
B --> F
C --> D
C --> E
F --> G
F --> J
G --> H
H --> I
H --> E
That graph is what gets compiled to device code.
Generated Code
DEBUG=4 prints the kernel. After optimization:
1 2 3 4 5 6 7 8 9 10 | |
DEBUG=2 paints the 4 yellow because it was upcasted. NOOPT=1 keeps the loop:
1 2 3 4 5 6 | |
Development and Debugging Tools
Useful flags:
DEBUG=2— data movement and kernel executionDEBUG=4— generated kernel codeVIZ=1— web graph-rewrite explorerNOOPT=1— skip optimizations
They compose in the shell: DEBUG=2 CPU=1 python docs/ramp.py.
Debug colours:
- Blue: global ops (
ADD) - Light blue: local ops (
COPY) - Red: reductions (
SUM) - Yellow: upcast (
EXPAND) - Purple: unroll (
UNROLL) - Green: groups (
GROUP)
VIZ=1 python ramp.py opens the rewrite explorer so you can watch UOps change instead of staring at dumps.
Advanced Topics
Low-Level UOP Construction
You can build UOps by hand:
1 2 3 4 5 6 7 8 9 10 11 12 13 | |
1 2 3 | |
Pattern Matching and Graph Rewriting
Rewrites are pattern matchers:
1 2 3 4 5 6 7 8 9 10 11 12 | |
There is sugar for the same idea:
1 2 3 4 5 | |
Performance and Training
With BEAM=2, Tinygrad is competitive today, and often ahead of PyTorch on unoptimized work, on AMD, and in training — about 20% faster than PyTorch on AMD in the HLBC implementation, if you believe their numbers.
There is no trainer.fit(). examples/beautiful_mnist.py is a complete MNIST trainer: everything you need, nothing you don’t.
Closing
You now have the spine: UOps, lazy tensors, realize, rewrite, kernel. The charm is not that each line is simple. It is that a small group of careful engineers can still own the whole stack — and you can be one of them.
Glossary
| Term | Meaning |
|---|---|
| Device | Hardware backend where tensors live and kernels run (e.g. CPU, CUDA, METAL). |
| Lazy / Realize | Tinygrad records operations lazily; realize() forces execution. |
| UOp | Immutable micro-operation node in the computation graph. |
| DAG | Directed acyclic graph composed of UOps. |
| Kernelize | Transformation that groups UOps into executable kernels. |
| ShapeTracker | Tracks tensor views/shapes without copying data. |
| Upcast | Optimisation that widens ops to work on multiple elements at once. |
Comments