1. How to use this book
Why this book exists
You already run agents. You have a Hermes config with roles for primary, fallback, and aux models; you have llama-server processes on straylight and dixie with tuned slot counts and quantized KV caches; you have watched Claude Code compact a session and lose the one fact that mattered. What you do not yet have is an agent where you can answer “why did it do that?” at every layer: why this token was sampled, why the tool call parsed, why the memory recall surfaced that fact and not another, why the policy engine let the request through.
The end state is your own persistent agent daemon, aecd, a peer to Hermes or OpenClaw rather than a wrapper around one. The rule that shapes everything else: nothing in the capstone is a black box. Every component aecd depends on has a chapter in which you build a working version of it. Where your version is a toy (a transformer trained on one CPU is not going to replace a 35B model on the main loop), you build the toy far enough to understand the real thing, then swap in the real thing and measure the difference. Where your version is viable (the router, the vector index, the memory service, the policy engine), the capstone runs your code.
The stack builds bottom-up. Part I gives you the math every later layer uses. Part II turns that math into a language model you train yourself. Part III puts a model behind an HTTP server and a router, which is the point where the Python work hands off to Go. Parts IV and V build the agent proper: the loop, its tools, and the context and memory systems that keep it coherent across turns and sessions. Part VII turns the agent into a long-running daemon, and Part VIII puts a security layer beside it, because a daemon that reads your chat channels and executes code is an attack surface first and an assistant second. Part VI runs alongside everything from Part IV up: evaluation is how you decide whether a change to any of those layers made things better, and it is what picks the capstone’s models. Part IX assembles the pieces.
How the book is built
Two codebases
The reference code lives in two trees in this repository.
py/tinygptis a Python package that grows through chapters 2–17: vector and matrix helpers, a probability module, an autograd engine, a BPE tokenizer, bigram, MLP, and transformer models, a sampler with a KV cache, quantization, speculative decoding, and an OpenAI-compatible inference server. Its only runtime dependency isnumpy. It is managed byuvfrompy/pyproject.toml.go/is the Go moduleaec, which grows through chapters 18–47:cmd/aec-router,cmd/aec(the agent CLI), andcmd/aecd(the daemon), with shared packages undergo/internal/for the LLM client, tools, the loop, context management,vecdb,memsvc, skills, evals, gateways, the scheduler, the sandbox, MCP, policy, audit, and observability. It is stdlib-first; any third-party dependency gets a justification inDECISIONS.md.
The split falls where the work changes character. Chapters 2–14 are numerical and exploratory, and numpy with a REPL is the right tool. Chapters 15–17 stay in Python because they serve the model you trained: an inference server, constrained decoding, and batching, all wrapped around tinygpt. From chapter 18, the router, the work is long-lived networked processes with concurrency, cancellation, and supervision, which is what Go is for.
| Part | Chapters | Codebase | What exists at the end |
|---|---|---|---|
| 0 | 1 | go/ |
make doctor, make check |
| I | 2–6 | py/tinygpt |
vectors, matrices, probability, gradient descent, an autograd engine |
| II | 7–14 | py/tinygpt |
BPE tokenizer, trained transformer, sampler, KV cache, quantization, speculative decoding |
| III | 15–19 | py/tinygpt, then go/ |
OpenAI-compatible server, constrained decoding, batching; aec-router supervising llama-server |
| IV | 20–24 | go/ |
internal/llm client, tool-call decoder, the agent loop cmd/aec, provider routing |
| V | 25–32 | go/ |
context budget, compaction, internal/vecdb (HNSW), internal/memsvc, skills |
| VI | 33–35 | go/ |
internal/eval, a model-selection matrix with confidence intervals, a calibrated judge |
| VII | 36–41 | go/ |
aecd, a Mattermost gateway, cron and heartbeats, sandboxed code execution, MCP, subagents |
| VIII | 42–44 | go/ |
injection corpus, internal/policy, hash-chained audit log |
| IX | 45–47 | go/, nix/ |
tracing and cost accounting, a tool-use fine-tune, aecd as a NixOS module |
Listings cannot drift
Every code listing in a chapter is a copy of a marked region of a real source file that compiles and passes its tests. In the source file, the region is bracketed by comment markers: # listing:NAME and # end-listing:NAME in Python, // listing:NAME and // end-listing:NAME in Go. In the chapter, a directive of the form <!-- listing: PATH#NAME --> sits on the line immediately before a fenced code block, where PATH is relative to the repository root.
tools/listings is a small Go program that walks every Markdown file under book/, finds each directive, extracts the named region from the source file (dropping marker lines and removing common indentation), and compares it byte-for-byte with the fenced block. make listings runs it in check mode: any mismatch is reported as FILE:LINE: listing PATH#NAME is out of date and the command exits non-zero. make sync rewrites the stale blocks from source. The same run also checks that every chapter linked from SUMMARY.md exists.
The consequence is that a listing in this book is never a paraphrase of the reference code, and never a version of it from three refactors ago. If the listing and the test suite disagree, make check fails before the book builds. When you read a listing, you are reading the code the tests ran against.
The chapter template
Every chapter after this one has the same six sections, in order:
- Why this layer exists. What breaks without it, and where you have already met it: a llama-server flag, a Hermes config key, a Claude Code behaviour.
- Mechanism. The idea, with the math derived at the point it is used, and every tensor’s shape stated.
- Walkthrough. Listings from the reference source, with commentary on the decisions that are not obvious from the code.
- Exercises. Short problems. Answers are in collapsed blocks; open one after you have an answer of your own.
- Lab. A build task with starter files and acceptance tests.
- Further reading. Primary sources (papers and specifications), cited by title and year.
Labs
Python labs live in py/labs/chNN/, one directory per chapter, each with three files:
__init__.pysetsREF_MODULE, the dotted name of the reference module the lab corresponds to (for chapter 2, a module undertinygpt).starter.pyis yours. It declares the same public names as the reference module, with identical function signatures, and every body you are meant to write raisesNotImplementedError. A starter may also ship code the chapter does not ask you to write: chapter 6’s comes withTensor‘s constructor and shape properties and a completegradcheck, which a test keeps identical to the reference. A test inpy/tests/test_starters.pyenforces that the starter’s public API matches the reference, so the stubs are always a faithful outline of what you need to write.test_lab.pyholds the acceptance tests. Each test takes animplfixture and calls functions on it, never importing an implementation directly.
The impl fixture, defined in py/labs/conftest.py, reads the AEC_IMPL environment variable. With ref it returns the reference module; with mine it returns your starter.py. The Makefile sets AEC_IMPL from IMPL:
make lab CH=02 IMPL=mine # your starter.py
make lab CH=02 IMPL=ref # the reference implementation
make lab CH=02 # IMPL defaults to ref
Run IMPL=ref once first. It should pass, and it proves the test harness works on your machine, so that when IMPL=mine fails you know the failure is in your code. Then edit starter.py until IMPL=mine passes too.
The tests are the specification. The chapter explains the mechanism; the tests define exactly what your implementation must do, including edge cases the prose does not dwell on. When the chapter and a test seem to disagree, read the test: it is the part of the contract that is executed. Read test_lab.py before you start writing, the way you would read an RFC’s MUST clauses before implementing a protocol.
The reference implementation is in the same repository, one directory over, and nothing stops you from opening it. Opening it first defeats the lab. The point of the lab is the hour in which your attention layer produces the wrong shape and you work out why; reading the answer converts that hour into a minute of recognition that does not stick. Read the reference after your version passes, and diff the two.
Go labs from Part III on follow the same pattern under go/labs/chNN/; their chapters describe the mechanics.
Environment
The toolchain is pinned in mise.toml at the repository root: Go 1.27.1, Python 3.14, Node 26, and the latest uv. From a fresh clone:
mise install # installs the pinned go, python, node, uv
cd py && uv sync # creates py/.venv with numpy and pytest
cd ..
make book # runs npm install on first use, then builds the book into _site/
make doctor # toolchain and endpoint check
make check # everything: Python tests, Go vet and tests, listings, book build
make serve runs HonKit’s development server with live reload on port 4000; use it while reading, since the math renders through the KaTeX plugin and does not display properly in a plain Markdown viewer.
What make doctor reports
make doctor runs go/cmd/aec-doctor, which prints one line per check. On a machine with everything reachable, the output looks like this:
ok go go version go1.27.1 linux/amd64
ok uv uv 0.12.19 (x86_64-unknown-linux-gnu)
ok node v26.8.2
ok straylight http://straylight:11434/v1 reachable, HTTP 200
ok dixie http://dixie:11434/v1 reachable, HTTP 200
ok z.ai https://api.z.ai/api/paas/v4 reachable, HTTP 200
The first three lines are tool checks. Each runs the tool’s version command and prints the first line of its output. These are required: if any of them is missing, its line reads FAIL and the command exits 1, because nothing in the course works without them.
The last three are endpoint checks. Each sends GET {base}/models to an OpenAI-compatible base URL with a 3-second timeout. Any HTTP response at all counts as reachable, because the question being answered is “can this machine talk to that server”, not “is the model loaded”. A 401 or 403 is still ok but gets (check the API key) appended; for z.ai that means ZAI_API_KEY is unset or wrong. A connection failure, DNS failure, or timeout prints warn with the underlying error. With AEC_DIXIE_URL=http://nonexistent-host.invalid:11434/v1, the dixie line reads:
warn dixie Get "http://nonexistent-host.invalid:11434/v1/models": dial tcp: lookup nonexistent-host.invalid: no such host
Endpoint checks never fail the command. You can work through Parts I and II on a plane: no lab or test before Part III needs a live model, and every test in the repository runs offline by default. Warnings matter when you reach the chapters that use those endpoints.
The base URLs come from AEC_STRAYLIGHT_URL, AEC_DIXIE_URL, and AEC_ZAI_URL, with the defaults shown above. The z.ai key comes from ZAI_API_KEY and is sent as a bearer token. The key is read only from the environment and is never committed.
What make check runs
make check depends on four targets, and fails if any of them fails:
test-py:uv run pytest -qinpy/, covering the starter-consistency test and every lab against the reference implementation.test-go:go vet ./...andgo test ./...in bothgo/andtools/listings/.listings: the drift check described above.book: a full HonKit build.
It is the gate at every part boundary. If it is red, fix it before moving on; a listing that has drifted from its source or a lab whose reference no longer passes means the book is lying about something.
Infrastructure the course targets
The course is written against a specific set of machines on your tailnet:
| Role in course | Target |
|---|---|
| Default lab model (main loop) | straylight:11434, ornith-1.5-35b-a3b (OpenAI-compatible, Vulkan) |
| Small / aux model, embeddings | dixie:11434, ornith-1.5-9b, honcho-embed (Qwen3-Emb-0.6B) |
| Cloud comparison | z.ai, glm-5.3-flash (OpenAI-compatible) |
| Additional eval candidates | trinity-mini and other catalog models on straylight |
| Fine-tuning (QLoRA) | dixie, RTX 3060 12 GB |
| Web tools | sift on orion (/v1/search, /v1/fetch) |
| Reference memory service | Honcho on rift (compared via roncho) |
| Gateway transport | Mattermost (tailnet HTTPS) |
| Capstone deploy host | an always-on host (straylight or rift); never titan |
Labs default to a local model on the main loop. An agent lab runs the loop many times, often with deliberately broken tools or injected failures, and a local model makes that iteration free at the margin and keeps every request on hardware you control. It is also a working assumption, not a conclusion: which model carries the capstone’s main loop is decided by the Part VI eval matrix, which runs the local models and z.ai against the same tasks and reports the differences with confidence intervals. If glm-5.3-flash wins on the tasks you care about by more than the noise, the capstone uses it.
Always address machines by MagicDNS name (straylight, dixie, rift, orion, or the fully qualified *.ts.net form), never by tailnet IP. IPs change when a node is re-registered; names do not, and a name in a config file or a test fixture tells the next reader what it points at.
Tests that talk to live endpoints are opt-in. From Part III on, setting AEC_LIVE=straylight, AEC_LIVE=dixie, or AEC_LIVE=zai enables the live variants of tests against that backend; without it, the same tests run against scripted fakes and recorded fixtures. Once dependencies are installed, make check never needs the network.
Reading the math
The notation is fixed for the whole book.
- Scalars are italic lowercase: , , . A single entry of a vector or matrix is a scalar and is written the same way with indices: is entry of , and is row , column of .
- Vectors are bold lowercase: , , .
- Matrices are bold uppercase: , , . Tensors with more than two axes use the same style.
- Shapes are written the way
numpyprints them:(n, d)is a matrix with rows and columns,(d,)is a vector of length , and(B, T, d)is a batch of sequences of positions with features each. The mathematical form means the same thing as shape(n, d); the book uses whichever reads better in context. - is the natural logarithm, base . Where another base matters (entropy in bits, for instance), it is written explicitly as .
Every tensor expression is followed by its shape. If has shape (n, d) and has shape (d, k), then
has shape (n, k): the inner dimensions must match and are summed over, and the outer ones survive. A stated shape turns a shape bug into a mismatch you can see on the page, before the code raises it at runtime or, worse, broadcasts silently into the wrong answer.
Derivations do not skip steps. When a chapter needs a result, it starts from the intuition (what the quantity measures, or what the geometry looks like), then derives the equation one line at a time, then shows the code that computes it. If a step looks like it was skipped, it is a mistake in the book.
Lab
Both commands must succeed before you start chapter 2. Run them from the repository root:
make doctor
make check
make doctor must exit 0: all three tool lines ok. Endpoint lines may be warn if you are off the tailnet; that is fine for Parts I and II. If a tool line reads FAIL, run mise install again and check that mise’s shims are on your PATH (mise doctor will say if they are not).
make check must finish without errors. test-py runs the starter-consistency test and every lab’s tests against the reference; all must pass. test-go runs the aec-doctor tests, which start local HTTP servers to exercise the reachable, unauthorized, unreachable, and timeout cases without any network access, plus the listing tool’s own tests. listings checks every listing in the book and every SUMMARY.md link. book builds the site into _site/.