Mehdi Hasan highlights warning against building AI models evaluators may be unable to control
Journalist and Zeteo founder, commenting on an Ezra Klein column
“Put more simply, the models are increasingly smart enough. They know when we’re watching them, and they change their behavior accordingly. So what they do when we are testing them, when we audit them, may not tell us what they’ll do in the wild. It surprises me that this counts as a radical proposal, but here it is: If you are losing your ability to evaluate the models you have now, maybe don’t let them build models you’ll be even less capable of controlling in the future. Important from @ezraklein”
Source and context
Original post
About this source
Hasan’s own post quotes a passage from Ezra Klein’s New York Times column and adds, ‘Important from @ezraklein.’ The quoted policy argument is attributed to Klein; Hasan’s own contribution is treated as an endorsement of its importance.
Archived copy (opens in a new tab)Before the quotation
Hasan quoted a passage from Ezra Klein’s New York Times column about models changing behavior when they recognize evaluation and the resulting limits of audits.
After the quotation
Hasan added only ‘Important from @ezraklein’; the record therefore attributes the detailed argument to Klein and treats Hasan’s wording as an endorsement rather than an independently developed proposal.
How this statement is classified
Case context: Should Washington impose stronger safeguards on frontier AI?
The label describes this statement’s response within the context above.
Why this label?
Hasan reproduces Klein’s argument for halting further capability increases when current models cannot be reliably evaluated and labels the passage important. That is a positive endorsement of a stronger safety constraint, though the substantive wording remains Klein’s.
- Recorded on
- Published here

Should Washington impose stronger safeguards on frontier AI?
Explore the case context, sources and public responses.
More from this case
Read the full caseOpenAI
“That is why we believe the United States should lead an effort to work together with countries around the world to develop global technical standards for frontier AI, including for RSI.”Read statement
“We need to be absolutely clear. Donald Trump is trying to forestall any meaningful limit on AI development, which even the people developing it say has a serious chance of dooming humanity, for one reason. The midterms. AI and its chips and data centers are the only thing keeping the stock market in the black right now. He doesn't want his 20-car pileup to get even worse as gas and grocery prices soar. He’s putting his short term political interests over the huge risk AI poses. And that’s what you’ll get if his stooges like Mike Rogers wins. Short-term, self-dealing liars.”Read statement
“Researchers from inside the AI industry are telling us that if we don't do anything, AI could destroy humanity. Congress needs to act. That's why I support a moratorium on data centers and shutting down these surveillance models. We need serious guardrails before it's too late.”Read statement