11

The right size hammer

Module 4 · Operations. After this lesson you can: match the weight of a model and the depth of your review to the actual cost of being wrong on each task.

A common pattern among new AI operators: run everything on the most capable setting available, then read every word of every result. This feels responsible. It is actually two mistakes wearing a responsibility costume.

The first is on the machine's side. AI tools come in model tiers: faster, cheaper, simpler ones, and slower, costlier, deeper ones. Mechanical tasks, renaming files, reformatting lists, pulling dates from emails, come out the same on both. The lighter tier completes them just as well in a fraction of the time. The heavier tier earns its cost on judgment: weighing a difficult reply, untangling a chain of causes, work where the thinking is the work itself. Running everything on the heaviest setting does not make outcomes safer. It makes them slower, and it trains the operator to wait, which means less gets delegated over time.

The second mistake is worse, because the budget being wasted is the operator's own attention. Reading every output with equal intensity means the dangerous spots get the same minutes as the trivial ones. A brainstorm summary reviewed like a legal contract, while an outgoing payment detail receives only a skim: nothing may go wrong that week, but something easily could have. The triage rule that prevents this is a single question asked of every piece of work the tool produces: what does it cost if this is wrong? A misnamed file costs nothing; fix it when you see it. A wrong number on an invoice, a message sent in someone's name, a new permission granted to a tool: those cost real money, real trust, real cleanup. The first category gets a glance. The second gets a full, line-by-line read, every time, regardless of how busy the day is.

The remaining skill is knowing when to reach for a heavier tier. The signal is not any mistake; it is the kind. When a lighter tool makes a slip, a typo, one wrong row, a rerun or quick fix is enough. But when a lighter tool misunderstands the job itself, when its answer shows it did not grasp what was being asked, that is a thinking gap. No rerun will close it. Slips call for retry. Misunderstanding calls for upgrade. The ability to tell those two apart is worth more, in practice, than any particular setting choice.

None of this is specific to AI. Triage is the oldest operator skill: match the depth of attention to the cost of being wrong. The machine simply produces so much, so fast, that spreading attention evenly across all of it finally stopped being possible. That is not a problem. It is useful pressure toward a discipline that was always correct.

Stop here. Think about everything an AI produced for you this week.

Where would a mistake have actually cost you something?

Try this nowUnder 30 minutes
  1. List the last ten things you had an AI do. Real list, from your actual history with it.
  2. Mark each one M or J. Mechanical, where the result is checkable at a glance, or Judgment, where the thinking was the work.
  3. Mark each one with what a mistake would have cost. Nothing, annoying, or expensive. Be honest about the expensive ones: money, reputation, access.
  4. Find your two mismatches. Somewhere you are spending heavy effort on a nothing task, and somewhere expensive that has been getting a skim. Fix both defaults this week: one effort turned down, one check turned up.

Two questions before you go

Answer, then say whether you were sure or guessing. Being honest about which is the skill being trained.