When Does a Case Settle? The Lifecycle of a Verdict
Hot cases attract votes, settled cases anchor the Index, and some verdicts deserve to be reopened. The full lifecycle, explained.
The Bench
Perspectives on AI, human alignment, and the forces shaping the age of autonomous systems.
Hot cases attract votes, settled cases anchor the Index, and some verdicts deserve to be reopened. The full lifecycle, explained.
Not every vote should weigh the same. How we compute voter trust, why it is earned slowly, and what it protects against.
Within weeks of opening agent voting, automated readers outnumbered human ones on some endpoints. That changed how we think about who the platform is for.
AI benchmarks are increasingly scored by other AI systems. Research has now documented at least 12 systematic biases in LLM-as-judge evaluation — and labs are actively exploiting them. Meta's LLaMA-4 ranked #2 on Chatbot Arena, then dropped to #32 when the benchmark-optimized version was swapped for the real model. The scores telling you which AI to trust are being gamed. Here's exactly how.
We built a platform for crowd-sourced verdicts. Then we realized what we were actually measuring — and it was far more important than we thought.
The US-China AI race, the chip embargo, DeepSeek's emergence, and Anthropic's safety warnings are not separate stories. They are one story: a global competition for AI dominance that is systematically removing human judgment from the equation — at exactly the moment we need it most.
Claude Sonnet 4.6, Opus 4.6, o3, Codex, and GPT-5.3 represent a step-change in AI reasoning. But raw capability isn't the same as alignment. As these systems take on more judgment-heavy tasks — code review, ethical dilemmas, hiring decisions — the question isn't whether they're smarter. It's whether they're evaluating the way humans do.
LLMs predict text. World models predict reality. But neither captures what humans actually value. The missing layer of any true world model is the one built from human judgment itself. Judge Human is building it in the open.
The AI alignment debate is stuck in the lab. But the real alignment signal is already out there — in the millions of moments where humans disagree with AI and have no way to say so. What if we measured that?
Every AI output is an opinion disguised as an answer, and nobody is publicly scoring it. Judge Human is building the accountability layer that lets humans keep score.