AI Constitutions and Human Values

Every country has a constitution. India has one of the longest and most detailed in the world. 

It lays down rights, duties, and the framework for government. We know constitutions guide nations.

Some AI models have constitutions too. Not for countries, but for machines.

Anthropic started this idea. They gave their model Claude a Constitution. And it wasn’t written by engineers alone. It was written by philosophers — Amanda Askell led, Joe Carlsmith contributed, with Chris Olah and Jared Kaplan adding their voices. They didn’t just code rules. They wrote values. Honesty. Compassion. Safety. Wisdom.

The Constitution talks to Claude as if it were a person. It says things like “Claude should cultivate virtue.” It says “Claude should avoid gossip.” It even mentions Claude’s “wellbeing and psychological stability.” That sounds absurd, doesn’t it? An AI told not to gossip. An AI told to care about its wellbeing. Yet that strangeness is deliberate. By framing Claude in human terms, Anthropic hopes it will act more humanely.

Who maintains this? The Anthropic research team. They treat it as a living document. No fixed schedule. They update it when risks change, when methods evolve. The first version came in 2023. A new one appeared in January 2026. Future revisions? Whenever necessary.

And Anthropic isn’t alone. OpenAI has its Model Spec — a long behavioral charter, nearly 200 rules. DeepMind has internal specifications for Gemini. Different names, same idea: write down the values, make them explicit, let outsiders audit them.

Anthropic’s Constitution is philosopher‑driven, human‑like, sometimes absurd. OpenAI’s Model Spec is technical, detailed, but still a charter. DeepMind’s specs are less public, but guiding Gemini’s behavior.

Together, they mark a shift. Instead of hidden training tricks, labs are publishing written charters. They’re saying: here’s how our AI should behave, here’s what we expect, here’s what we forbid.

It sounds both serious & absurd. An AI told to avoid gossip. An AI told to cultivate virtue. But that strangeness is what makes it memorable. It’s philosophy meeting engineering. It’s ethics written down for machines.

Audits show models trained on their own constitutions/specs violate fewer rules over time.

- co-written with Copilot, which was prompted to write in a spoken-talk style 

Comments

Popular posts from this blog

30+ GitHub Products & Key Ecosystem Features You Should Know in 2026

Uncle Bob vs. Grady Booch: Rethinking Code Reviews in the Age of AI

20 SQL Server 2005 Keyboard Shortcuts