Evan Hubinger on Artificial Intelligence
Anthropic alignment researcher; personal public assessment, not a formal corporate response.
“I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
Source and context
Original post
About this source
Primary response by Anthropic’s Alignment Science Lead. Hubinger agrees with Coxon about the severity of potential future risk, credits Anthropic’s effort, and explicitly concedes the lack of a demonstrated alignment solution. His follow-up distinguishes present-model risk from future superintelligence risk.
Before the quotation
Hubinger replied directly to Coxon’s thread about the risk of future AI systems.
After the quotation
Hubinger's follow-up says present-model risk is low and his concern is future recursively self-improving superintelligence. His statement that Anthropic is trying its best is consistent with Coxon's fuller WIRED interview, not an independent rebuttal of the structural criticism.
How this statement is classified
He directly qualifies Coxon’s criticism rather than rejecting the underlying danger or unresolved alignment problem.
- Recorded on
- Published here
- People and groups discussed
- Anthropic
More from this case
OpenAI
“OpenAI declined to comment when approached by CNBC”Read statement
Anthropic
“both enormous benefits and unprecedented risks”Read statement
Anthropic
“Anthropic declined to comment on the post.”Read statement