Why local, and for whom
The three honest reasons to run a model on your own machine, the ceiling nobody tells you about, and a test for whether you are the person this course is for.
Every other course in this school hits the same wall in its first lesson. The teacher cannot paste the student's writing. The manager cannot paste the review. The admin cannot paste the log. The answer is always some version of "take the identifying parts out," and that answer is correct, and it is also a tax you pay on every single task forever.
A model running on your own computer removes the tax. Not by being more trustworthy than a hosted service — by being a program on your disk that has no network connection to anything. There is no policy to read, no setting to verify, no tenant to share. The text goes from your file to your processor and back. That is the whole pitch, and it is a good one.
It is also not free of cost. This lesson is about the trade you are actually making, so you decide it once, deliberately, instead of discovering it in month two. Nothing gets installed today. Today you find out whether the thing is worth an afternoon of your life, and you produce the one document the other five lessons are built on.
Before you start
- The computer you actually do work on. Not the spare one in a drawer. A model on a machine you open twice a month is a model you will never use.
- Twenty minutes with your real week in front of you. Email, files, notes, tickets — whatever your work is actually made of. The exercise at the end is an audit, and an audit from memory produces a list of things you wish you did rather than things you do.
- A blank document you are willing to keep. Plain text or markdown, saved where you will find it in three weeks — call it
local-ai.md. Every lesson adds a line: your memory budget, your model, your routing rule, your daily job. By lesson 6 it is the whole setup on one page. - An assistant you already use, if you have one. You will use it once, near the end, on text that contains nothing about anybody. Without one, the exercise still works on paper.
- Nothing downloaded. People who install first spend the afternoon comparing tools and never write the list, and the list is what decides everything you build later.
One thing to know before you read on: you are allowed to finish this lesson and stop. That is a real result, not a failure.
The three reasons that hold up
Privacy by construction. This is the big one. When the machine is the boundary, you stop making a judgment call per paste. The things you have been carefully not asking about — the personnel note, the medical paperwork, the half-finished contract, the journal — become ordinary inputs. For a lot of people this single change is worth the entire setup afternoon.
Cost that stops moving. After the download, running the model costs you electricity and the time your fan spends spinning. There is no per-token meter, no monthly seat, no anxiety about whether the summarize-everything script you just wrote is going to be expensive. Jobs that are obviously not worth paying for — rename these four hundred files, draft a one-line summary of every note in this folder — become obviously worth doing.
It works when nothing else does. No internet on the plane, no connectivity at the site, the provider having a bad afternoon. A local model is a local program. It does not have outages, it does not deprecate the version you liked, and it does not change its pricing in the middle of a project.
The ceiling, stated plainly
A model you can run on a personal computer is smaller than the hosted ones by a wide margin, and you will feel it. Specifically:
- Hard reasoning degrades first. Multi-step logic, math with several dependent steps, code that has to be right across a whole file. Small models do the first two steps well and then quietly go sideways, in fluent, confident prose.
- Long documents are harder. They can take a lot of text in, but attention across all of it thins out. Ask about page one of forty and you will often get page one; ask what connects page three to page thirty-one and you may get an invention.
- It knows less. Fewer facts about the world, more gaps, and the gaps get filled with plausible-sounding guesses rather than an admission. Treat every factual claim as unverified.
- It is slower on your hardware. Often fast enough to read along with, sometimes not. You will measure that yourself in lesson 3 rather than take a number for it.
Anyone who tells you a model that fits on a laptop matches the largest hosted systems is selling something. What is true, and more interesting, is that the floor has risen enough that a large share of ordinary work now sits comfortably below the ceiling. Rewriting. Summarizing. Extracting fields from messy text. Reformatting. Drafting the boring version so you can fix it. Explaining something you half-understand. Asking questions about a document you paste in. That is most of what people actually use assistants for.
So who is this for
You are the right reader if any of these is true:
- There is a category of your work you have simply never handed to an AI tool because of what is in it.
- You have repetitive text work in volume, and the meter is what stops you.
- You want a setup that keeps working regardless of subscriptions, pricing changes, or connectivity.
- You want to actually understand these systems, and running one where you can see every moving part is the fastest way in.
You are the wrong reader if you need frontier-level reasoning for your main job, if you do not want to maintain anything ever, or if you are hoping local will be cheaper and better. It is cheaper and private. It is not better.
And you do not have to choose. The most useful arrangement, and the one this course builds toward in lesson 5, is both: a local model handling the private and the repetitive, a hosted one handling the hard and the public, and a rule you can state in one sentence about which gets what.
What "local" actually means here
One clarification, because the word gets stretched. Local means the model weights are files on your disk and the program doing the computation is running on your processor. It does not mean "a private cloud." It does not mean "a provider who promises not to train on it." Those can be fine choices; they are different choices. When you finish lesson 3 you will unplug your network and watch the thing keep answering, which is the only demonstration that counts.
Three phrases to treat as different from local. "We do not train on your data" is a promise about use, not location. "Private endpoint" means isolated infrastructure you do not physically control. "On-device" in a phone feature is genuinely local, but it is a fixed model doing a fixed job, not something you point at your own folder. All three can be reasonable. None is what this course builds.
A worked example
Everything in this section is invented. Marisol Fenwick-Okada is not a real person, she does not work anywhere, and the output below was written by me to show the shape — it was not captured from a real tool.
Marisol is a self-employed bookkeeper who uses a hosted assistant and likes it. She also has a mental list of things she has never once pasted into it, and had never written that list down. She set a timer and went through her last two weeks the boring way: sent mail, the folder she saves client paperwork into, her notes app, her invoicing tool.
Her rule while writing was to describe the shape of the task, never the content: not a name, not a business, not an amount. That rule is what makes the list safe to hand to anything, and she wanted that, because she intended to use her existing assistant to help sort it. Four lines from what she ended up with:
- rewrite a blunt payment-chasing email into something neutral
- pull every date and amount out of a scanned statement, into a table
- summarize a handwritten meeting note into three bullets
- turn two pages of messy notes into a tidy handover document
Then she ran this, in a fresh conversation, with the whole list pasted in:
Below is a list of task shapes. For each line, tell me which
single thing is stopping me from handing it to an AI tool:
SENSITIVITY — the text itself contains a person, a body,
money, or something under an agreement.
VOLUME — nothing sensitive, but I would do this many times
and paying per use would stop me bothering.
CAPABILITY — the task needs hard multi-step reasoning,
current facts about the world, or correctness across a long
document.
Output only the original line, then a dash, then one word.
Do not rewrite my lines. Do not add lines. If a line could be
two of these, pick the one that would stop me first.
- rewrite a blunt payment-chasing email into something neutral
- pull every date and amount out of a scanned statement, into a table
- summarize a handwritten meeting note into three bullets
- turn two pages of messy notes into a tidy handover document
What came back (illustrative — I wrote this, it is not real output):
- rewrite a blunt payment-chasing email into something neutral — SENSITIVITY
- pull every date and amount out of a scanned statement, into a table — SENSITIVITY
- summarize a handwritten meeting note into three bullets — SENSITIVITY
- turn two pages of messy notes into a tidy handover document — SENSITIVITY
She changed two things after reading that. First, she stopped treating the sort as the answer and started treating it as a draft, because a column that says the same word four times is not telling her anything — it is telling her that she only wrote down the sensitive ones. So she went back and added the jobs she does in bulk and had never thought of as AI work at all: renaming scanned files, stripping headers out of exports, turning a folder of receipts into one list. Those came back VOLUME, and they are the ones she is most likely to actually automate.
Second, she rewrote three lines to be more specific. "Summarize a handwritten meeting note" became "summarize a handwritten note into three bullets, under forty words, keeping every date." Vague lines produce vague prompts later. A line specific enough to argue with is a line you can build against in lesson 6.
Do this now
Twenty minutes, with your actual week open in front of you.
- Open the blank document and title it Things I could not paste. Save it where you will find it. This is the document every remaining lesson adds to.
- Walk your last two weeks in whatever holds your work — sent mail, your files, your notes, your ticket queue. Not from memory; from the record.
- Write a line for every task you did not hand to an AI tool because of what the text contained, or handed over only after sanitizing it. Describe the shape, never the content: not "HR stuff" but "rewrite three paragraphs of a performance note into neutral language."
- Add the bulk jobs too — the repetitive text work you have never considered handing over because doing it four hundred times would obviously cost money. Those belong on the list even though nothing about them is private.
- Mark each line S, V, or C — sensitivity, volume, or capability — for the single thing that stops you. Do it by hand, or paste your list into the sort prompt above and correct what it gets wrong. Either is fine; the lines contain nothing about anybody by construction.
- Rewrite your three vaguest lines so each one names the input, the output, and the constraint.
You are done when your document holds at least five specific task lines, each marked S, V, or C, and at least one line in the V column.
If the list comes out genuinely empty — no S lines, no V lines — you have learned something useful: you do not need this course, and you should spend the afternoon on something else.
If it goes wrong
- The list comes out as four vague categories. "Client stuff." "Personal." That is a list of folders, not of tasks. Fix it by going one level down: open the folder, find the last thing you actually did in it, and write that single action down instead. One real line beats six categories.
- Every line is marked S and nothing is marked V. Almost everyone does this the first time through: sensitivity is what you notice, cost is what you have quietly accepted. Fix it with a different question — what text work do you repeat constantly and would never dream of paying per use for? Those V lines are usually where the first working habit comes from.
- You write down tasks you wish you did rather than tasks you do. The tell is that the list is aspirational — "analyze my industry," "plan my quarter." Fix it with the record: if you cannot point to the email, file, or note that the line came from, delete the line.
- You catch yourself pasting real content into the sort prompt. Stop, delete the conversation, and rewrite the lines as shapes. The sort is a convenience, not the point. From lesson 3 onward you will have somewhere safe to do it; today, describing rather than pasting is itself the lesson.
- You decide the ceiling means this is not worth it. Check which column you are reasoning from. If your V lines are empty and your S lines are all hard judgment work, you are right, and you should stop. If your V lines are full, the ceiling is irrelevant to them: nothing on that list needs reasoning at all.
Resources
- Large language model — Wikipedia — plain background on what these models are and how they are built; read it if "model" still feels like a black box before you size one in lesson 2.
- Open weights — Wikipedia — what it means for a model's weights to be downloadable, and how that differs from open source; useful before you pick a family.
- Hallucination (artificial intelligence) — Wikipedia — the background on why a confident wrong answer is the default failure, which is the ceiling described above in one page.
- AI Risk Management Framework — NIST — a US federal framework for reasoning about AI risk; borrow its vocabulary when you have to explain your boundary to someone who wants it in writing.
- Ollama — the tool lesson 3 installs; look at it today only to see what you are heading toward, and do not download it yet.
This lesson is free to read. An account keeps your progress and opens the discussion.
Discussion
This lessonMembers post here. The subscription is $15 a month.
Nothing here yet.