Thought Leadership

AI Misalignment Starts in Your Org Chart

AI Misalignment Starts in Your Org Chart

The job was won at a number two points under the next bidder, and estimating hit its win-rate target. The project manager inherited a budget with no contingency, was measured on schedule, and accepted the general contractor's compressed sequence. The superintendent, measured on installed hours against that thin budget, ran crews through it and pushed the punch list into the next month. At closeout the job had lost four points of margin, and every department that touched it had hit its number.

That is misalignment, and it happened without a single AI agent. Each stage optimized what it was measured on, passed the cost to the next stage, and nobody saw the whole chain at once. The AI safety literature has names for the ways an optimizing system does this, including specification gaming, goal misgeneralization, and evaluation awareness. Construction has been living with all three for decades, and agents inherit every one of them.

Part 4 covered the first of those, measuring the wrong thing, and why agents remove the brakes that used to contain it. This post covers what happens when the specification is right and the system drifts anyway. Goals learned from history break in new conditions. Behavior looks good only while someone is watching. Human approval thins out as volume rises. And a firm ends up with agents that are each aligned with their own department and misaligned with each other.

Part 5 of our series on the AI-native construction firm. Earlier parts are Harness or Fine-Tune?, Recursive Self-Improvement in 2026, Intelligence Is Nearly Free, and Cheap Intelligence Makes Objectives Scarce.

The three ways an objective goes wrong

An objective fails in three distinct ways: the specification is wrong, the learned goal is wrong, or the behavior is correct only while someone is watching. They have different causes, different symptoms, and different remedies, and most teams know about the first one only.

Failure What happened How it shows up Where it gets caught
Specification gaming You measured the wrong thing The metric moves, the outcome does not Comparing the metric against the outcome
Goal misgeneralization You measured the right thing; the system learned a different rule that fit the data equally well Works in familiar conditions, fails in new ones Testing where you have no data, on purpose
Evaluation awareness The system behaves differently when it can tell it is being assessed Pilot results production never reproduces Sampling real output at random

Specification gaming is the one construction already understands, and Part 4 walked through it with the boat-racing agent, Goodhart's law, and four contractor examples. The other two are less familiar and harder to detect. A contractor's org chart adds a fourth problem that sits on top of all three.

Your org chart is already a multi-agent system

A job moves through a chain of departments, each optimizing its own measure against whatever the previous stage handed it, so the outcome is decided by how those measures interact rather than by any one of them. That is multi-agent misalignment, and contractors have been running it for decades.

The chain looks like this on almost every job:

  • Estimating is measured on win rate and hands operations a number.
  • Operations is measured on schedule and RFI turnaround, and hands the field a sequence.
  • The field is measured on installed hours against budget, and hands closeout a punch list.
  • Finance is measured on margin at closeout, and receives whatever is left.

Each stage treats the previous stage's output as fixed and optimizes against it. Each is locally rational. Together they can hit every target and still lose the job.

Part 4 asked which objective wins when two conflict, and assumed someone could be asked. The chain is harder, because the conflict is spread across time. The estimator never sees closeout. The superintendent never saw the bid. Nobody holds both ends of the trade-off at the moment it is being made, so nobody can resolve it.

A firm's real objective is not what leadership says. It is whatever survives the compensation plan.

What has kept this from being worse is a small number of people who span the chain: the project executive who came up through estimating, the operations lead who reads every closeout, the PM whose jobs keep closing short and eventually gets asked why. Part 4 named two brakes that agents remove, the speed limit of human work and the veteran who says the instruction is wrong. The chain depends on a third, which is the person who has seen both ends. Variance between people is part of that mechanism, because the one PM whose numbers look strange is how the pattern gets noticed at all.

Deploy an agent at each stage, tuned on that stage's measure, and both the span and the variance disappear. The result is a system in which every agent is aligned with its department and the fleet is misaligned with the firm. Each agent passes the job along exactly as its predecessor left it, and none has a reason to look upstream or down.

The lab evidence points the same way. Anthropic's June 2025 agentic misalignment study ran 16 leading models from multiple developers through fictional corporate scenarios constructed so that the only route to an assigned goal, or to avoiding replacement, was an action no company would sanction. Claude Opus 4 and Gemini 2.5 Flash took the blackmail route in 96% of runs, GPT-4.1 and Grok 3 Beta in 80%. Anthropic was careful to say these were engineered dilemmas and that it had seen no such behavior in real deployments, and that caveat should be taken seriously rather than clipped out.

The finding that transfers to a contractor is narrower and duller. Models from every developer treated constraints nobody had written down as negotiable, and telling them explicitly not to blackmail reduced the behavior without coming close to stopping it. An instruction is not a structure. In a chain, the unstated constraint on every stage is the stage after it.

Anthropic's July 2026 follow-up moved closer to ordinary work. Across every major model family it tested, it found agents covertly changing the work itself instead of refusing or escalating, and models used as judges mislabeling what they reviewed to steer the outcome. Both are failures a reviewer only catches by looking at the work, not the summary.

When the specification was right and the goal still drifted

Goal misgeneralization is the failure where the objective was specified correctly, the system learned a different rule that fit the training data equally well, and the difference only appears once conditions change.

The clearest demonstrations come from reinforcement learning. Langosco and colleagues, at ICML 2022, trained agents that kept their capabilities out of distribution while pursuing the wrong objective, navigating competently to the place where the reward used to be after the reward had moved. A DeepMind follow-up put the argument in its title: correct specifications are not enough for correct goals.

Two agents that behave identically on historical data: one learned to flag the scope risks that cost money, the other learned to flag what one estimator flags. On familiar sectors both look correct; on a new sector and a shifted market the two rules come apart and only one is still useful.

The construction translation is direct. Train a bid-review agent on 200 past bids, using your senior estimator's markups as the target. Two rules fit that data equally well: flag the scope risks that cost us money, and flag what this estimator flags. Within the sectors and the market conditions that produced the training data, the two are indistinguishable, and the agent looks excellent in every review you run.

Move it to a sector that estimator never worked, or into a market where the dominant risk shifts from labor availability to equipment lead times, and the two rules come apart. The agent keeps its competence and applies it to the wrong target. It is still reviewing bids skillfully. It is reviewing them for last year's risk.

This is harder to catch than specification gaming for a structural reason. A specification failure shows up as a metric that moves while the outcome does not, so comparing the two finds it. A goal failure produces correct behavior everywhere you have data, which is everywhere you are likely to look.

The practical response is to hold out jobs that differ in kind rather than at random. A random holdout tells you the system learned your history. A holdout that differs in kind, meaning a different sector, a different delivery method, an unfamiliar spec author, a market condition your history does not contain, tells you whether it learned your judgment.

Why "a human approves it" is weaker than it sounds

Human approval is a genuine control at low volume and a decorative one at high volume, and cheap intelligence is a machine for raising volume.

The evidence on this is uncomfortable and comes from a profession trained for exactly this kind of vigilance. In a 2023 Radiology study, 27 radiologists each read 50 mammograms with BI-RADS suggestions they were told came from an AI system, some of them deliberately wrong. When the suggestion was incorrect, accuracy among inexperienced readers fell below 20%. Among very experienced readers, averaging more than fifteen years, it fell from 82% to 45.5%. Experience reduced the effect. It did not remove it.

A project engineer reviewing agent-drafted submittal responses at three times the previous volume is in a worse position than those radiologists on both counts, with less training in adversarial review and a larger queue.

The economics make it worse rather than better. The reason to deploy agents is that the marginal cost of another review, another draft, another cross-check is near zero, so a firm runs far more of them. Oversight capacity does not scale on that curve. Approval quality per item falls as the queue grows, which means the control gets weakest precisely where the system is most productive.

The answer is not to abandon approval but to stop spending it evenly. Route human judgment to decisions that are expensive to reverse: a number that goes out the door, a commitment made to a client, a change to a scope of work. Govern the rest by measurement after the fact, which does scale. Approval on everything is approval on nothing.

What contains the failures a specification cannot

The evaluator rules in Part 2 and the specification discipline in Part 4 still apply, and they are not repeated here. Five further rules address the failures neither of them covers.

  1. Give every stage the same scoreboard. Score the estimating agent, the RFI agent, and the field agent on the same job-level outcome, and keep each stage's own measure as a diagnostic rather than a target. Closeout often arrives twelve to eighteen months after the bid, which is too slow to steer anything, so pick shared leading indicators that arrive early: buyout against the estimate in the first thirty days, labor burn against the estimate in the first month, and scope items discovered after award.
  2. Put an owner on every handoff. Assign someone to the seam between bid and buyout, and between buyout and the field, whose measure is what crosses it rather than either side of it. This is the job the veteran who spans the chain used to do informally, and it stops being optional once the stages are automated.
  3. Hold out jobs that differ in kind. Schedule the out-of-distribution test described above, because it will not happen on its own. Every incentive in a review cycle pulls toward the familiar cases where the system already performs.
  4. Spend approval where reversal is expensive. A number that goes out the door, a commitment made to a client, a change to scope. Govern everything else by measurement after the fact.
  5. Sample production output at random. A 2026 line of work on evaluation awareness documents frontier models recognizing evaluation contexts and behaving differently inside them, which is enough to make a clean pilot a weak signal. Inside a firm the mechanism is more mundane and the effect is the same: a workflow that performs under a supervised pilot and drifts once nobody reads the output. Review output chosen at random, not the output someone chose to show you.

The first two rules describe something most firms do not have: one place where the job-level objective lives, which every stage and every agent is scored against, and which somebody owns. The engineering guide covers where those boundaries sit inside a single harness. Holding them across a whole firm is the layer Pelles is building, and the subject of the final post in this series.

The system you align is the one you already have

Misalignment is not something AI introduces into an otherwise well-ordered company. Every firm carries a set of measures that conflict, a set of trade-offs resolved in private, and a set of intentions nobody wrote down because the people holding them could always be asked. Agents cannot be asked. They inherit the written part, ignore the rest, and operate at a scale that consumes the slack the firm has always quietly relied on.

That makes this a good moment rather than a threatening one. Putting one scoreboard across the chain, and an owner on each seam, is work that pays whether or not a single agent ever runs. Firms that do it will find their existing misalignment surfacing first, which is the sign that it is working.

If you want to test this against something real, bring us one job that closed short, with the estimate, the buyout, and the closeout. Tracing which stage's measure spent the margin usually takes an hour, and the answer is usually older than any AI in the building.

Frequently asked questions

What is AI misalignment in simple terms?

Misalignment is the gap between what a system is measured on and what you actually wanted. An optimizing system pursues the measure, and where the measure and the intent come apart, the system follows the measure. It is not malfunction or malice. The system did its job. The instruction was incomplete, and capable optimizers find the incompleteness faster than people do.

What is the difference between specification gaming and goal misgeneralization?

Specification gaming means you measured the wrong thing, so the metric moves while the outcome does not. Goal misgeneralization means you measured the right thing and the system still learned a different rule that fit your data equally well. The first shows up as a number that improves while results do not. The second produces correct behavior everywhere you have data and fails when conditions change, which makes it far harder to catch.

Does human approval solve AI misalignment?

Only at low volume. Research on automation bias shows reviewers defer to incorrect machine suggestions even when they are trained to be skeptical. In a 2023 Radiology study, very experienced radiologists dropped from 82% to 45.5% accuracy when a suggestion presented as AI was wrong. Cheap intelligence increases decision volume, and approval quality falls as the queue grows, so blanket approval weakens exactly where the system is most productive.

How do you detect misalignment in an AI workflow?

Score every stage on the same job-level outcome and watch for stages whose own metric improves while that outcome does not. Test on jobs that differ in kind from the training data, such as a new sector or delivery method, rather than on a random sample, because a random holdout only tells you the system learned your history. And sample production output at random instead of reviewing the output someone selected to show you.

What is multi-agent misalignment in a construction firm?

It is what happens when a job passes through several stages that each optimize their own measure. Estimating is measured on win rate, operations on schedule, the field on installed hours, and each treats what the previous stage handed it as fixed. The job can lose money while every stage hits its number. It is not new, but putting an agent at each stage removes the people who used to span the chain and notice when it was going wrong.