The Artificer's Primer

A Propædeutic Enchiridion for the Making of Thinking Engines, Built Upward from the Matrix

The Artificer's Primer

This book ends with you running your own persistent agent daemon — something in the shape of Hermes or OpenClaw — on your own hardware, talking to your own inference backend, with a memory service, a vector database, and a policy engine standing between it and the outside world. Every layer between "linear algebra" and "systemd unit" is built by hand before you are allowed to lean on a library for it: tokenizer, attention, KV cache, OpenAI-compatible server, a router that supervises llama-server processes, an HNSW index, and the capability-based policy engine wrapped around all of it. Nothing here is disconnected from the destination: the vector index you build in Part V is the one your memory service is built on and your daemon queries in Part VII; the router you build in Part III is the one your daemon calls in Part IX.

The book runs on two codebases. tinygpt (Python, Parts I–III) is a from-scratch language model — autograd engine, tokenizer, transformer, sampler, and inference server — built on nothing but numpy. aec (Go, Parts III–IX) is the agent stack: router, tool-calling loop, context and memory subsystems, evaluation harness, gateway, and the security layer that makes running an autonomous daemon on your own machine defensible. The name aec is a placeholder; rename it once it's yours.

This is not an introduction to programming. You are assumed to already be fluent in at least one systems language, comfortable with Unix, and capable of reading a paper. What is not assumed is graduate-level math: every derivation starts from the geometric or probabilistic intuition and builds up to the equation actually used in the code, with no steps skipped. If your linear algebra and probability are rusty, Part I rebuilds them from scratch; if they're solid, it will move fast and still serve as notational reference for everything after.

Each chapter follows the same shape: why the layer exists (what breaks without it, and where you've already met it inside Claude Code, Hermes, or llama.cpp), the mechanism with math derived where it's used, a walkthrough of the actual reference implementation, exercises with worked answers, and a lab — a build task with starter files and acceptance tests you run against your own code. The book hands you a spec and a test suite and gets out of the way.

The ten parts:

  • Part 0 — Orientation. How to use this book, and getting your toolchain and inference endpoints working.
  • Part I — Mathematical foundations. Vectors, matrices, probability, calculus, and autodiff, in Python with only numpy.
  • Part II — Language models from scratch. Tokenization through inference engines: BPE, bigram and MLP models, attention, the transformer block, sampling, instruction tuning, quantization, and speculative decoding.
  • Part III — Serving models. An OpenAI-compatible inference server, constrained decoding, continuous batching, and the Go router that supervises llama-server processes.
  • Part IV — The model as a component. Wire protocols, tool calling at the protocol level, the agent loop itself, tool failure semantics, and provider routing.
  • Part V — Context and memory. The context window as a budget, prompt engineering, session memory and compaction, embeddings, a hand-built vector database, a memory service, and agent-authored skills.
  • Part VI — Evaluation. Eval design, the statistics of evals, and LLM-as-judge.
  • Part VII — Becoming a daemon. Runtime architecture, channel gateways, autonomy (cron and heartbeats), code execution sandboxes, MCP, and subagents.
  • Part VIII — Security. Threat modeling for agents, a capability-based policy engine, and secrets, audit, and supply-chain hygiene.
  • Part IX — Operations and capstone. Observability and cost accounting, fine-tuning for tool use, and the capstone: your own aecd packaged as a NixOS module for an always-on host.

Read chapter 1. If you want to know where this book came from — it was generated from a single prompt and a short steering conversation — read How this book was made.

results matching ""

    No results matching ""