Evan Hubinger on Artificial Intelligence

Anthropic alignment researcher; personal public assessment, not a formal corporate response.

I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.

Source and context

Original post

Evan Hubinger responds to Jacob Coxon on superintelligence risk and Anthropic’s efforts (opens in a new tab)X response and follow-up

About this source

Primary response by Anthropic’s Alignment Science Lead. Hubinger agrees with Coxon about the severity of potential future risk, credits Anthropic’s effort, and explicitly concedes the lack of a demonstrated alignment solution. His follow-up distinguishes present-model risk from future superintelligence risk.

Before the quotation

Hubinger replied directly to Coxon’s thread about the risk of future AI systems.

After the quotation

Hubinger's follow-up says present-model risk is low and his concern is future recursively self-improving superintelligence. His statement that Anthropic is trying its best is consistent with Coxon's fuller WIRED interview, not an independent rebuttal of the structural criticism.

Mixed or conditional
How this statement is classified

He directly qualifies Coxon’s criticism rather than rejecting the underlying danger or unresolved alignment problem.

Recorded on
Published here
People and groups discussed
Anthropic

More from this case

OpenAI

OpenAI declined to comment when approached by CNBC
Read statement

Anthropic

both enormous benefits and unprecedented risks
Read statement

Anthropic

Anthropic declined to comment on the post.
Read statement
Read the case