Whose Picture Is the Agent Building From

Whose Picture Is the Agent Building From

We've gotten reasonably comfortable saying AI isn't neutral. What we usually mean is upstream. Training data, or the company that built the thing, put a thumb on the scale somewhere. That account is accurate, and it's also conveniently distant. Something happened in a lab, to a model, before any of us showed up. It lets the bias belong to someone else.

It's worth taking that part seriously first, because it's realer than the hand-waving suggests. Anthropic publishes the document it uses to shape Claude's values, and that document openly ranks safety ahead of helpfulness. [1] Its own researchers have found that the human judgments used to train these models tend to reward answers that read convincingly over answers that are correct. [2] None of that is neutral, and to their credit nobody is pretending it is. The idea itself is older than the current tools. Batya Friedman and Helen Nissenbaum laid out how computer systems carry bias back in 1996, when the systems in question were airline booking software and immigration screening. [3] Bias that existed in the world before the system, bias from technical constraint, bias that only appeared once real people used the thing. Thirty years later, it's the same problem, new story.

Each one of those describes a finished artifact. The bias got built in somewhere upstream, and then you ran into it. Which means it's also not really yours to do anything about. You can choose a different tool or complain to the people who made it, and that's about the extent of it.

With AI tools, the problem is much closer. Bias exists in what I bring to the tool when I ask.

I got a clean look at this recently. Two of us went after the same bug, in the same codebase, in the same week, using the same AI coding tool. What came back were two solutions that barely resembled each other. Not different in polish. Different in what each of us had decided the problem was. The agent built faithfully toward whichever version it was handed, with the same confidence either way, because it had no picture of its own to check ours against.

I know that system deeply, so I kept the agent on a short leash and it produced a narrow fix focused on the specific area of need. The other person, newer to the code, gave it room to range and it reached for a broader restructure. It would be easy, and wrong, to read that as one of us getting it right and the other getting it wrong. We both brought a bias, and they were the same kind of thing. Mine was my deep knowledge of the system, which told me where the problem had to be and kept the agent pinned to that spot. Theirs was a strong, reasonable theory of how software like this should be built. Both are biases in exactly the sense this post means, a picture of the problem formed before the tool was ever opened, that the agent then built toward. The difference in outcome wasn't that one of us was biased and the other wasn't. It's that my picture happened to match this codebase, and I only know it matched because of how it turned out. My own knowledge was steering just as hard as their theory was. Had it been wrong, that same short leash would have produced a confidently wrong fix, cleaner and narrower and harder to catch than any sprawling one.

I wasn't the neutral one in that story. I was the one whose bias happened to match the codebase.

There is no unbiased way to ask. There's only the bias you can name and the bias you can't.

Two kinds of bias steer the answer, and they don't behave the same. There's what you don't know, the gaps where the agent, finding no context, fills in whatever looks reasonable from the outside and builds on it as settled. And there's what you're sure of, the hard-won theory of how software ought to work that stops being a preference the moment you say it out loud. To the tool it arrives as a specification, and it can't tell the difference between the conviction I earned over years and the one I formed ten minutes ago. Both come back as code that compiles.

Neither is a flaw to be corrected. The gap is just the shape of what any one person knows. The conviction is expertise, and expertise is most of what makes the work good. This is where the neutrality talk quietly misleads us. It implies there's a clean way to ask, a version of the question with the bias filtered out, if we were only disciplined enough to find it. There isn't. Every prompt carries a picture of the problem, and the picture is the bias. You cannot subtract it, because subtracting it would mean asking for nothing.

What you can do is know which one you're holding. The old systems were finished before you touched them, their assumptions sealed in by people you'd never meet. An agent isn't finished. It assembles the answer at the moment of asking, and it hands you back a version of your own understanding, rendered confident and executable. That's the new shift, and it cuts both ways. It means the bias is no longer only upstream and out of reach. It's in the room, it's mine, and being mine, it's the one part of this I can actually take responsibility for.

So I've added a single question in front of the tool now. Do I know this part of the system, or do I only feel like I do. It doesn't remove the bias. Nothing does. It just tells me which one I'm about to hand over, and most days that's the only honest control I've got.


  1. Anthropic, "Claude's new constitution," January 2026. https://www.anthropic.com/news/claude-new-constitution ↩
  2. Sharma et al., "Towards Understanding Sycophancy in Language Models," 2023. https://arxiv.org/abs/2310.13548 ↩
  3. Batya Friedman and Helen Nissenbaum, "Bias in Computer Systems," ACM Transactions on Information Systems, 1996. https://dl.acm.org/doi/10.1145/230538.230561 ↩