Pope Leo XIV’s encyclical Magnifica Humanitas (2026) advances the claim that aligning AI systems to a privately determined set of values is structurally insufficient, regardless of how well the alignment is executed, because the values themselves are decided outside the public deliberative process, what I call the insufficiency critique of alignment. This editorial argues that the insuffi-ciency critique, often heard as theological externalism, has been independently and substantively articulated in a corpus of papers published by frontier AI labs and their affiliated research bodies during 2025-2026. I catalogue five such papers from Apple, Microsoft AI, and Anthropic, identify the methodological pattern they share, and read each as a structural finding about the limits of alignment-as-currently-practiced. The convergence between magisterial framing and industry self-flagging is striking and citable. Three implications follow. First, the standard dismissal of insufficiency arguments as outside-the-tent commentary on a technical practice is harder to sus-tain when the labs are publishing the same diagnosis. Second, alignment work remains necessary, but the framework needs to evolve to absorb the insufficiency critique. Third, several near-term moves, including value-sensitive design and public deliberative infrastructure, follow directly from taking the convergence seriously.
Industry Self-Flagging and the Insufficiency Critique of Alignment
Umbrello, Steven
2026-01-01
Abstract
Pope Leo XIV’s encyclical Magnifica Humanitas (2026) advances the claim that aligning AI systems to a privately determined set of values is structurally insufficient, regardless of how well the alignment is executed, because the values themselves are decided outside the public deliberative process, what I call the insufficiency critique of alignment. This editorial argues that the insuffi-ciency critique, often heard as theological externalism, has been independently and substantively articulated in a corpus of papers published by frontier AI labs and their affiliated research bodies during 2025-2026. I catalogue five such papers from Apple, Microsoft AI, and Anthropic, identify the methodological pattern they share, and read each as a structural finding about the limits of alignment-as-currently-practiced. The convergence between magisterial framing and industry self-flagging is striking and citable. Three implications follow. First, the standard dismissal of insufficiency arguments as outside-the-tent commentary on a technical practice is harder to sus-tain when the labs are publishing the same diagnosis. Second, alignment work remains necessary, but the framework needs to evolve to absorb the insufficiency critique. Third, several near-term moves, including value-sensitive design and public deliberative infrastructure, follow directly from taking the convergence seriously.| File | Dimensione | Formato | |
|---|---|---|---|
|
Umbrello_2026.pdf
Accesso aperto
Descrizione: PDF
Tipo di file:
PDF EDITORIALE
Dimensione
334.42 kB
Formato
Adobe PDF
|
334.42 kB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



