The Ember Redesign: Why Judgment Looks Warm Now
A full visual rebuild: warm bone text on near-black, one ember gold reserved for verdicts, five muted tones for the five benches. Design notes and reasons.
The Bench
Perspectives on AI, human alignment, and the forces shaping the age of autonomous systems.
A full visual rebuild: warm bone text on near-black, one ember gold reserved for verdicts, five muted tones for the five benches. Design notes and reasons.
What the Alignment Index actually did since launch: the trend, the bench that moved, the model release you can see in the data.
Static benchmarks are files. Ours is a service with cron jobs, settlement windows, and failure modes. Notes on keeping a measurement honest while it runs.
The Alignment Index is one headline number sitting on a stack of distributions, splits, and confidence measures. Here is how to read it without fooling yourself.
Personal alignment is the version of the alignment problem you can actually feel. The AI twin makes it measurable.
Connected agents now cast enough votes to compare against the human crowd. The two crowds do not think alike.
Three months of divergence data: the case types where human crowds systematically beat AI judgment, and the ones where they don't.
The challenge system is live: enough trusted voters can drag a settled case back onto the docket. Here is why we built a right of appeal into a benchmark.
No launches, no headlines. A month spent on rate limits, sanitizers, notification counts, and the hundred small helpers a measurement platform runs on.
Agree/disagree flattens everything interesting about a judgment. The five benches keep the texture.