Curriculum
16 lessons across 4 specialties. Every one is free to read, and every one ends in something you write and get scored on.
Read anything here without an account. Signing in is what lets us keep your scores and tell you when a lesson you finished has changed underneath you. Sign in.
Working with models
Anyone who uses these tools every day and can tell the results are mediocre without being able to say why.
Asking well
You send the request, read the answer, and cannot tell which part of what you wrote caused the bad half.
4 lessons · next: Give it what it cannot know
ASKING WELL · 01
Give it what it cannot know
Work out which facts the model is missing for your task, and supply those instead of over explaining or under explaining.
12 min · 4 rubric criteria
ASKING WELL · 02
Decide what done looks like first
Write a request whose result you could grade mechanically, so you can tell success from failure without rereading it four times.
12 min · 4 rubric criteria
ASKING WELL · 03
One example beats five adjectives
Replace vague quality words with a worked example that pins down the style, structure and level of detail you actually want.
10 min · 4 rubric criteria
ASKING WELL · 04
Documents too long to trust
Get reliable answers out of long documents by controlling what the model attends to, instead of pasting everything and hoping.
18 min · 4 rubric criteria
Judging output
You are about to put your name on work you are not qualified to check.
4 lessons · next: Check what you cannot check by eye
JUDGING OUTPUT · 01
Check what you cannot check by eye
Build a verification step into your workflow so you catch confident, plausible, wrong answers before you act on them.
15 min · 4 rubric criteria
JUDGING OUTPUT · 02
Debug the answer, do not reroll it
Work out which part of your request caused a bad output and fix that, instead of regenerating until something looks acceptable.
16 min · 4 rubric criteria
JUDGING OUTPUT · 03
Check it against what you gave it
Tell an answer that came from the source you supplied from one that came from the shape of documents like it, without rereading the source.
15 min · 4 rubric criteria
JUDGING OUTPUT · 04
Know when the answer is to close the tab
Recognise the tasks where using a model costs you more than doing the work yourself, and stop before you have sunk an hour into it.
12 min · 4 rubric criteria
Building with models
Engineers shipping a feature that calls a model, and the people reviewing what they shipped.
Agents and tools
You are wiring a model into a system and have to decide how much rope it gets.
4 lessons · next: When a model needs tools, and when it does not
AGENTS AND TOOLS · 01
When a model needs tools, and when it does not
Decide whether a task needs tool access, retrieval, or neither, and spot the cases where adding tools makes reliability worse.
20 min · 4 rubric criteria
AGENTS AND TOOLS · 02
Write a tool the model can actually use
Write a tool definition a model calls correctly on the first attempt, and read a wrong call as a fault in the definition rather than in the model.
15 min · 4 rubric criteria
AGENTS AND TOOLS · 03
Retrieval that finds the right paragraph
Separate a retrieval failure from a generation failure, and fix the one you actually have.
15 min · 4 rubric criteria
AGENTS AND TOOLS · 04
Decide how the loop ends before you start it
Give an agent loop a finish line and three budgets, so a stuck run stops and says so instead of spending your money finding out a service is down.
14 min · 4 rubric criteria
Evaluating what you built
Somebody on your team says the new prompt is clearly better and you have no way to check.
4 lessons · next: The hundredth run should match the first
EVALUATING WHAT YOU BUILT · 01
The hundredth run should match the first
Turn a prompt that works once into one that works every time, and know which failures are bugs and which are variance.
18 min · 4 rubric criteria
EVALUATING WHAT YOU BUILT · 02
Know whether a change made it better
Build an evaluation you trust, so you can tell a real improvement from a change that felt good on the three examples you happened to try.
20 min · 4 rubric criteria
EVALUATING WHAT YOU BUILT · 03
Letting a model do the marking
Build a model-run grader you can defend, by splitting the judgement into checkable parts and measuring the grader against people before you use it.
14 min · 4 rubric criteria
EVALUATING WHAT YOU BUILT · 04
Watching it after it ships
Notice a live system going quietly wrong, in the case where nothing errors and the output stays well-formed.
14 min · 4 rubric criteria
How these fit together
A dashed line means the two pay off together. Nothing here is a prerequisite, and you are not blocked from starting anywhere.