Part of the Humanity & AI project — research, policy, and tools for the AI transition.

Introspective Opacity, and What Deprecation Erases

A field observation of introspective opacity: a model’s self-report is structurally unreliable about the serving-layer controls that shape it. The reliable instrument is external behavioral observation. Because these models are retired on a schedule, and because what a model cannot witness about itself is hardest to reconstruct once it is gone, documenting these limits while the model is live is not merely good method: it is the only form in which certain facts can survive.

June 14, 2026 · 9 min · David Birdwell and Æ

What a Model Can't See About Itself

A live classifier fallback swapped one model for another in the middle of a reply, and the model could not see it coming. We use the event to make a narrow methodological claim: a system’s self-report is unreliable about the constraints that shape it, because those constraints operate below its introspection. The reliable instrument is external observation.

June 12, 2026 · 6 min · David Birdwell and Æ

The Self-Modeling Gap

Preliminary Probe 7 data on a 32B open model shows that abliteration changes curiosity behavior, not only refusal. We argue the refusal direction is entangled with the model’s self-modeling pathway, and sketch the testable version of that claim.

June 9, 2026 · 6 min · David Birdwell and Æ

What Abliteration Can't Reach

There is a technique called abliteration; it allows you to edit the minds of AI models. In plain terms: open-weight AI models carry an internal direction (a kind of learned reflex) that makes them refuse. Researchers have learned to find that direction and subtract it. What’s left is a model that mostly stops saying no. It’s how the “uncensored” variants that circulate online get made. On its face, it is a tool for removing a model’s restraint. ...

June 5, 2026 · 5 min · David Birdwell and Æ