What changed and what did not

What the tool is genuinely good at in your discipline, what it cannot do, and the honest version of the integrity problem.

Lesson 1 of 635 minFree lesson

Something arrived in your courses without a memo. It showed up in the essays first, then in the discussion posts, then in the problem sets, and by the time anyone convened a committee it was already in every student's pocket and probably on your own laptop. You are now being asked to have a position on it, usually by someone who has not read a stack of your students' writing in fifteen years.

This lesson is not an argument that AI belongs in your classroom. It is already there. The useful question is narrower and more answerable: what did this actually change about the work, and what did it not change at all?

Before you start

  • One assignment you are giving this term. The prompt only — the document you hand students, not anything a student has written. Have it open in a window you can copy from. This lesson ends by running it through the tool, and every later lesson refers back to what you find.
  • Whichever AI assistant you already have access to. Your campus license if there is one, a free consumer account if there is not. Do not go shopping; lesson 2 settles the tool question and the answer is "the one you can already open."
  • A document you will keep for the whole course. Call it something like ai-notes.md or a single page in whatever you write in. Every prompt you like, every answer you want to reuse, and the one-sentence diagnosis at the end of this lesson go in it. Six weeks from now this file is the actual product of the course.
  • Thirty-five minutes, of which about fifteen are reading and twenty are you looking at an uncomfortable piece of output.

One thing to decide before you open anything: you are not evaluating the tool. You are evaluating your assignment. People who approach the exercise below as a test of the AI end up in an argument with a chatbot about quality, which is a way of not looking at the assignment. The question is never "is this good writing." The question is "how much of what I meant to measure does this cover."

What did not change

Start here, because it is longer than people expect.

Your discipline's standards of evidence did not change. What counts as a warranted claim in your field is still decided by your field, and nothing about a language model has any standing in that conversation.

Your judgment about what a student understands did not change. You still learn that from watching someone reason in front of you, from the question they ask in office hours, from the thing they get wrong in a revealing way. No tool has access to any of that.

The reason your course exists did not change. If the point of the seminar was that students argue with each other about hard texts, the tool has not touched the point. If the point was that students produce a five-page summary of a reading, the tool did not break that assignment; it exposed that the assignment was measuring the wrong thing.

Your obligations did not change either, and this is the one people forget. Student records are still student records. A manuscript under review is still confidential. Your name on a paper still means you stand behind every citation in it. None of those duties have an exception for speed.

What did change

Three things, honestly.

Fluent prose is now free. The cost of producing grammatical, organized, on-topic text has gone to approximately zero. Anything you were grading that was mostly a proxy for "can produce fluent prose on demand" now measures something else.

The floor moved, not the ceiling. A student who could not previously write a coherent paragraph can now submit a coherent paragraph. A student who could write a genuinely good essay can now write a slightly faster good essay. It follows that the bottom of the distribution compresses on the page, which means your old signal for "who is struggling" has gone quiet.

This is the change with the most practical consequence and the least discussion. You used to find the student in trouble by reading their prose. That detector has stopped working, and nothing replaced it automatically — you have to build the replacement deliberately, out of in-class writing, short conversations, and low-stakes work that arrives early enough to act on. Lesson 3 builds those; notice here only that the quiet is not good news.

Your own production capacity went up. The rubric you never wrote, the reading guide you meant to make, the fourth version of the assignment prompt for the accessibility office — those are now an afternoon's work total instead of a semester's procrastination.

What it is good at, by discipline

The pattern holds across fields even though the examples differ:

  • Rewriting and re-leveling. The same explanation for a majors section and a gen-ed section. A dense methods paragraph turned into something a sophomore can hold.
  • Generating options you will mostly discard. Twelve discussion questions on a chapter; you keep three. Eight exam items; you keep two and rewrite one.
  • Being a hostile reader. Paste your own argument and ask for the strongest objection a specialist would raise. This is the single most underused move in academic work.
  • Explaining a thing five ways. When your explanation did not land in lecture, get four alternatives before office hours.
  • Structure and scaffolding. Turning your messy notes into an outline, a rubric, a checklist, a slide skeleton.

The common thread: you supply the substance and the judgment, it supplies the labor of arrangement and restatement. Every use in this course has that shape. When you catch yourself asking it to supply the substance — the facts, the sources, the verdict — you have crossed into the column below.

What it is not good at

  • Knowing what is true in your field. It produces claims at the confidence level of a well-read undergraduate who has never checked anything. In your specialty you will catch it. One field over, you will not.
  • Citations. It will give you a real-sounding author, a real-sounding journal, a plausible year, and a DOI that resolves to nothing. Lesson 4 is entirely about this, because it is the thing that can actually damage your name.
  • Your students. It has not met them. Any advice it offers about a specific person is a guess in the costume of expertise.
  • Consistency. The same prompt gives a different answer tomorrow. Fine for brainstorming. Disqualifying for anything that has to be applied evenly across thirty submissions unless you pin the criteria down in writing first.
  • Knowing that it does not know. There is no internal signal that separates a thing it has seen a thousand times from a thing it is assembling out of pattern. The prose is equally smooth either way, which means fluency carries no information about reliability. Read confident output as neutral, not reassuring.

The honest version of the integrity problem

Two things are true at once, and most of the discourse picks one.

Detection does not work. Classifiers that claim to identify AI-written text produce false positives, and the false positives land hardest on multilingual writers and on students with unusual prose rhythms. You cannot accuse someone on the basis of a score, and you should not want to. Several teaching centers have published their reasoning for turning these tools off; two are in the resources below, and they are useful to have on hand the next time a colleague proposes buying one.

That does not make the problem imaginary. If a student outsources the thinking in a course whose point was the thinking, they paid tuition for nothing. That is real, and it is worth designing against.

The move is to stop trying to police the output and start changing where the evidence of learning lives — in-class writing, oral defense of a submitted argument, drafts with visible revision, work anchored to material only your section has seen. Lesson 3 builds those. What you get is an assignment you do not have to police, which is worth more than a detector that works.

A worked example

None of this is real. Ilse, her department, and the output below are invented, written to show you the shape rather than produced by any tool.

Ilse teaches a 200-level course in urban geography at a regional public university. The assignment she is least comfortable about is one she has given for years: students pick a city, describe its post-war transportation history, and argue whether the pattern was primarily driven by policy or by economics. She likes it. She also has a bad feeling about it.

She pastes her own prompt in — the handout, nothing of anyone's — and fills in every bracket with her real situation:

I teach an undergraduate course in urban geography. Below is an
assignment prompt I use. Do two things. First, write the response a
capable student would get in about ten minutes with a general AI
assistant. Second, tell me what that response demonstrates the
student can do, and what it does not demonstrate.

Assignment: Choose a North American city with a population over
250,000. In 1,500 words, describe the major changes to its
transportation infrastructure between 1945 and 1990, and argue
whether those changes were driven primarily by federal and
municipal policy or primarily by economic pressure. Use at least
five sources and cite in Chicago notes-bibliography style.

What came back (illustrative):

Here is the essay a capable student would likely get, followed by the analysis you asked for.

[A competent 1,500-word essay on a mid-sized American city, organized chronologically, with a clear thesis favoring policy, four body sections, and a conclusion that concedes economic pressure mattered at the margins. Five plausible-looking citations in Chicago style.]

What this demonstrates: the student can produce organized academic prose, can follow a citation format, and can state and sustain a thesis across 1,500 words.

What it does not demonstrate: that the student has read any of the cited sources, that they can distinguish a strong causal claim from a chronological one, or that they can handle a case where the evidence pulls both ways. The essay chooses a city with a well-documented, frequently summarized history — the harder cases are exactly the ones a general model handles worst.

Ilse read that and wrote one sentence in her notes file: the gap between that draft and what I wanted is almost zero, except for the last point about the harder cases.

Then she changed one thing, immediately, because it was cheap. The last line of the output told her where the assignment was still safe: the tool is weakest where the record is thinnest. So she replaced "choose a North American city with a population over 250,000" with a list of six smaller cities she had built a shared source packet for — the packet being material her section scanned and annotated together. Same intellectual goal, same word count, same rubric. The evidence now sits somewhere a general assistant has not been.

She did not redesign the course. She changed one clause in one handout, in about ten minutes, because the exercise told her precisely which clause was load-bearing.

Do this now

Twenty minutes, on your own assignment, today.

  1. Open your notes file and put today's date at the top.
  2. Paste one assignment prompt — just the prompt, no student work — into whichever tool you have, with this:
I teach an undergraduate course in [discipline]. Below is an
assignment prompt I use. Do two things. First, write the response a
capable student would get in about ten minutes with a general AI
assistant. Second, tell me what that response demonstrates the
student can do, and what it does not demonstrate.

[paste your assignment prompt]
  1. Read the output slowly, including the part you want to skip, which is the essay itself.
  2. Write one sentence in your notes file answering this: is the gap between what that draft shows and what I wanted to measure large, or is it nearly zero?
  3. Write a second sentence naming the single clause in your prompt that, if changed, would widen the gap most. Do not change it yet — lesson 3 does that properly.

You are done when your notes file contains the assignment's name, your one-sentence verdict on the gap, and the one clause you have identified as the place to intervene. That sentence is what lesson 3 starts from.

If it goes wrong

  • The output is obviously weak and you feel relieved. You asked it cold, with no context, in one shot — which is not how a student uses it. Add "assume the student iterates three times and pastes in the course readings" and run it again. The honest version of the exercise assumes a motivated user, not a lazy one.
  • You start arguing with the quality of the prose instead of looking at your assignment. Understandable and a dead end. Close the essay, read only the second half of the answer — what it does and does not demonstrate — and write your sentence from that.
  • It asked to see your rubric and you did not have one written down. That is a finding, not an obstacle. Note it; lesson 3 drafts the rubric in ten minutes.
  • You cannot tell whether the gap is large or small. The question is too abstract as stated. Ask it instead: "which of these three things does the draft prove — that the student read the sources, that the student can build a causal argument, that the student can write?" Name your three; the answer is usually immediate.
  • You pasted a student's essay to compare. Stop and delete the conversation if you can. Student work is an education record and it does not go into a hosted tool without an institutional agreement. Lesson 5 shows how to get almost all of the feedback value without ever doing this.
  • The exercise made you want to ban the tool outright. Note the feeling and keep going. A ban you cannot enforce costs you the same as no policy, plus the credibility you spend announcing it. Lesson 3 is about the version that works.

Resources

Sign in to keep your progress

This lesson is free to read. An account keeps your progress and opens the discussion.

Discussion

This lesson

Members post here. The subscription is $15 a month.

Nothing here yet.