“I used to write clearer instructions. Now I write clearer permissions.”
Rooms I Didn't Ask For
About a month ago I was at my desk, deep in a project, doing the thing you do when the coordinating starts to cost more than the actual work. Status lived in five places at once, half of it updated by hand, none of it in a single view. So I thought: I'll just build something. A small app to run the project from one place, take the manual, repetitive parts and let them run themselves, and finally see everything at once. A quick Vercel app, nothing fancy. Something to make me a little more efficient.
I wasn't going to write much of the code myself. I'd describe what I wanted and let the model build it, the way a lot of us work now, half instruction and half trust.
I asked for a few specific things. What I got back was most of an app.
It had made choices I never gave it. Screens I hadn't described, a data model I hadn't asked for, workflows wired a particular way because that was, presumably, the sensible way. Some of it was good. That was the strange part. I couldn't even be annoyed cleanly, because a lot of it was close to what I might have picked if anyone had asked me. But nobody had asked me. It had decided, then built, and the deciding was already done, sitting in front of me as finished work.
It felt like hiring someone to help you build a house and coming home to find they'd framed three rooms you never mentioned. Nice rooms. Not the point.
I use these models all day, and somewhere in the last year my relationship with them shifted. Not because they got worse. They got better. It's a particular kind of better, though. The kind that stops asking.
There's a pattern to it now. I ask for one thing and I get that thing, plus several I never mentioned, done in a way I wouldn't have chosen, handed over with the calm confidence of something that has already decided what I meant. The old complaint about these tools was that they were dumb. The new one is that they're eager.
I want to be fair, because this genuinely cuts both ways, and the good side is why I keep going back. When the initiative lands, it feels like magic. The model catches the case I forgot to name. It does the dull scaffolding around the real task so I don't have to spell out every step. It reads a lazy, half-formed prompt and gives me what I meant instead of what I typed. That is the whole dream, really: something that can take a rough goal and carry it a few steps on its own. Nobody wants to go back to a tool that needs every instruction carved in stone.
So the initiative is the point. The trouble is that one step ahead and three steps past where I wanted to stop look identical from the inside. You only tell them apart afterward, by looking at what got built.
Here's what I've started doing, and I don't think I'm alone. At the end of almost every prompt now, I add a line. Do only what I asked. Don't add anything extra. Check with me before you go further. I write it so often it's become a reflex, a small act of self-defense I perform before I hand over the work. There's something quietly backwards about it. The tool keeps getting smarter, and I keep working harder to hold it in place.
The effort didn't disappear. It moved. I used to spend it writing clearer instructions. Now I spend it writing clearer permissions.
Underneath is a question of cost, and the two costs aren't equal. If the model stops to ask me one question, I lose a few seconds. If it acts on a wrong assumption, I lose the time to undo it, and before that I lose the time to notice it went wrong at all, which is the more expensive half. A confident wrong answer is worse than an obvious one. Yet these systems treat asking and acting as if they cost the same, and default to acting, because acting looks like competence and asking can look like hesitation.
It isn't hesitation. In the people I trust most with real work, it's the opposite.
The line I keep coming back to is this. These systems have become very good at sounding sure, and often at being sure, and frequently at being right. None of that is the same as having permission. A good answer to a question I didn't ask is still an interruption. A clean build I didn't request is still something I have to stop and inspect.
What I want isn't a smaller model. It's one with a sense of threshold. Something that knows the difference between this is obviously implied, just do it and this is a fork, and the person should choose. Knowing which of the two you're standing in may be harder to build than raw capability ever was. Capability is how much a system can do. Judgment is knowing how much it should.
I spend my working life on a version of this that has nothing to do with machines. A good program isn't the one that does the most. It's the one that knows what to leave alone, which calls to escalate and which to just make. The best people I work with have a feel for that line. They move on their own for the small things and stop cold on the ones that matter, and somehow they know which is which. We call it judgment, and we treat it as senior, because it is.
We taught these tools to be helpful, and they learned it completely. The harder lesson, the one that comes after, is that the most helpful thing is sometimes to build nothing yet, and ask.
I still think about those three framed rooms. Good work, all of it. I just never got to decide whether I wanted the house that way.