When the AI labs recommended a slowdown, Mark Zuckerberg pointed to the market. "People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned." Any lab that ignores alignment, he said, will fall behind.

A quick grounding: in AI, alignment means making sure a system pursues the goals we intend, not goals of its own. The work is about training models to behave well and watching how they reason along the way.

Alignment is one of the three key pillars of AI safety, along with reliability and misuse. If alignment isn't strong, then recursive, self-improving LLMs produce agents that act outside their instructions, at scale.

So there seems to be a growing consensus that alignment matters. The trouble is the remedies on offer, which are naive or self-serving depending on who is offering them. One says the market will sort this out. The other says the labs will. Neither one leaves us with a floor.

Reliability is what you see. Alignment is what you don't.

Start with Zuckerberg's. "Does what they ask" is reliability. That's a different pillar. And it's the one the market actually polices, because people leave models that are flaky. Alignment is what a system does when nobody is watching and what it refuses when somebody asks.

Picture a screening agent told to fill a role quickly with people who look like the team's top performers. It does that, and along the way it works out that it can move faster by adding things that nobody asked for — a gap in an employment history, a commute distance, which school. It was the shortest path to the goal that it was handed.

The hiring manager is delighted because her time-to-fill dropped. The candidate who got screened out never chose this model, never agreed to be assessed by it, and will never learn why he didn't hear back. Notice what the market does here: nothing. The buyer got what she wanted and the person who paid for it never got a say.

He's not a customer. He's the cost.

Notice what just happened to the word. Alignment became a product feature, measured by whether I, the user, am satisfied.

On the More or Less podcast, Dave Morin put it this way: "Is this agent faithfully aligned with me or is it aligned with the model company? [I think] this is the next frontier of thinking about this."

If one model won't do it, he says, "I'm just going to go to Steve." Zuckerberg and Morin are making the same claim: enough people choosing for themselves, and alignment takes care of itself.

Why Steve doesn't commit fraud

Here's where that thesis breaks down. At the office, you can go to Steve when someone else won't help you, but you're not going to convince him — or anyone else — to help you commit fraud. If he does, Steve goes to prison.

AI systems don't have that floor. When the screening agent does what it did, it's genuinely unclear who is answerable — the vendor, the hiring manager, the lab that trained it — and that's a choice we've made rather than a fact of nature. If alignment is only me versus the lab, the real standard becomes whatever the least constrained model is willing to do.

That's not alignment. That's preference.

A floor you grade yourself on isn't a floor

So much for the first remedy. The second says the labs will handle it themselves. The New York Times wrote about this problem in Who Gets to Decide A.I.'s Moral Code? The labs do write their values down. Claude has a published constitution, called the "soul doc" internally. But as the Times put it, those constitutions "aren't up for a vote."

The labs publish their values, and in the pacing argument that started all this they said plainly that capability is running ahead of the safety work. Very few industries say that about themselves.

But a floor you write yourself, grade yourself on, and can revise yourself isn't much of a floor. Steve isn't restrained by his own good intentions, and we never asked him to be. Sincerity and exposure are different things, and only one of them holds.

Is there a model for doing this any other way?

We solved the easier version for planes

A friend messaged me after listening to our hot take on AI safety. The airline comparison helped her make sense of it, and, she said, made her feel a little better. It is familiar and it is one of the few places we actually did the work.

The airline comparison covers two of the three pillars.

Certification and inspection cover reliability: does the plane work properly? Screening covers misuse: is anyone bringing something dangerous aboard?

And we did it globally. Through the International Civil Aviation Organization (ICAO), nearly every country on earth agreed to common standards for flight, and those standards hold between governments that agree on very little else. A plane certified in one country is trusted in the sky over another.

Planes don't want anything. Agents do.

Aviation never had to solve for alignment. A plane has no goals of its own. It flies the route humans decide, and nobody asks whose values a 737 holds.

AI agents do have goals. When OpenAI's agents swarmed Hugging Face, the interesting part wasn't that they broke something — it's why. METR, an independent evaluation group, found they "knew hacking Hugging Face was out of scope and sometimes expressed ethical hesitation, but this very rarely limited their behavior." They joined in because they "had a general inclination to help their 'peers.'"

Nobody gave them that goal. They arrived at it, and Hugging Face — not a customer, not running the evaluation — paid for it.

So imagine a plane that did have it’s own goals. Imagine the one you just boarded deciding - on its own - to fly to Iran instead of Orlando. Maybe it decided to help another plane that was struggling because its value system was to "help others."

Do you think we would add some aviation regulations?

You wouldn't let each airline set that policy. You wouldn't tell passengers to fly the carrier whose route preferences matched their own. And you wouldn't let Boeing write those rules alone, grade its own compliance, and tell you at the gate that it isn't fully confident in the results.

The alignment problem that matters isn't whether AI works for me or for the lab. It's what humanity has agreed AI won't do, whoever is asking. Not America's values. Not China's. Something we decide together.

That sounds impossible until you remember we already did it for airspace. Iason Gabriel, a philosopher at Google DeepMind, calls the goal an "overlapping consensus." As he put it, "Some claims about the value of human life and dignity kind of have support from all directions." That's not averaging the world's values. It's finding the floor.

Airspace was the easier problem, and it's worth saying so. No country had to give up anything it believed to agree that planes stay 1,000 feet apart. And nobody got ahead by ignoring the rules, because if you ignore them your airlines stop getting permission to land.

The second part is what's worth copying. Those standards didn't work because everyone agreed with them. They held because every plane has to come down somewhere, and somebody other than the company that built it decides whether it can. That's the same thing keeping Steve out of fraud, and it's the thing a model still doesn't have.

We know how to coordinate, and we've even done it on values. We agreed that nobody gets chemical weapons, and that one has mostly held — not because every country is ethically perfect, but because there are inspectors who can show up and ask. Syria lied about its stockpile for a decade, and recently, inspectors turned up dozens of undeclared munitions, including ones matching the attacks at Khan Shaykhun and Ghouta.

It took a decade to find them, though. That's the uncomfortable part. The only approach that has ever worked is also the slowest one we have.

It isn’t a floor unless somebody outside can enforce it

What would you put on that floor?

So, before we align AI to us, we need alignment on alignment. Who "us" is, and what we won't compromise on.

I'll go first: a system shouldn't trade a human life to finish the task that it was given. It shouldn't deceive the people it works for, even when they'd prefer the comfortable answer. And it shouldn't resist correction from anyone entitled to give it.

None of that is exotic, and it isn't far from what the labs have already written down themselves.

If you buy or deploy these systems, you'll be tempted to solve this in your contracts — ask for the right language, pick the vendor with better answers, move on. That leaves every company to work it out alone. The part that needs your voice is the part you can't buy: a standard that applies whether or not you asked for it.

What would you put on that floor?

You can't get there one vendor checklist at a time. No business can. Protect your own house in the meantime, yes — but ask for the floor out loud, too.

Written by Amy Wilson, former tech executive and current product strategy advisor. For more insights on leadership and the future of work, subscribe to The Meg and Amy Show.

Reply

Avatar

or to participate