Debatable.com Back to the AI judge
How the AI judge is kept honest

The criteria are published before you speak.

An AI judge that the people running the site can quietly retune is not a judge, it is a house edge. So the rubric is written down and fingerprinted before the round, three independent models decide it instead of one, their level of agreement is published even when it is bad, and a human can overturn any ballot. Everything on this page is served from an open endpoint you can fetch yourself.

The short version. Fixed criteria, published in advance. Three model families vote, and a split panel records no winner rather than picking one. A person, never a model, hears appeals. We earn the same whoever wins. And the whole thing is open to critiques: if you can break the judge, we publish what you broke.

A debate round in progress: a speaker at the lectern under a whiteboard marked Government Side, with the rest of the room flowing the speech on paper and laptops.
The thing being judged. Everything below is what happens to a round like this one after the last speech.

Why this is a new kind of judge.

AI feedback is everywhere. A judge built as an institution, with criteria you can read in advance, agreement you can measure, and a person above it, is the thing that did not exist. The machinery on this page is the invention, not the model behind it.

The shortage it ends

Verdicts stop being rationed.

Argument has never had a talent shortage. It has a judging shortage: a careful listener costs real time for every round, so most arguments end with nobody deciding them. A judge that reads the whole exchange and writes a full decision in seconds makes the verdict as available as the argument.

The inversion

You can read the judge's mind.

Everyone should know how the judge will decide before speaking. Here the method is a full document, fingerprinted before the round so it cannot change after you speak, and every decision names the issue that settled it.

Disagreement as data

The judge measures itself.

Three model families vote on every recorded round, and how often they agree is published even when the number is unflattering. A round they split on records no winner at all. No judging system we know of, human or machine, publishes its own reliability figure next to its verdicts.

What is judging your round, right now.

The season pins the rubric and the models. The hash is a fingerprint of the criteria: if a word of the rubric changes, the hash changes, so nobody can be judged against criteria that were edited after they spoke.

Loading
Reading the live charter.

How the council reaches a decision.

The council is three independent ballots, not one model asking two others for advice. No brain sees another brain's answer before it votes.

Same round, separate reads.

Every brain receives the same casual 1v1 transcript and published rubric. Each returns its own winner, six scorecard axes, deciding issue, argument score out of 100, and reason for decision.

One brain, one vote.

Two matching votes carry the verdict. A 3 to 0 and a 2 to 1 are shown differently, and an unavailable seat is disclosed. A split council does not get a tie-break. The emergency Claude backup is labeled as a single judge, never as the council.

The result keeps the disagreement.

The score and scorecard axes use the panel median, which limits any one outlier. The written ballot comes from a juror in the majority. Any dissent stays separate instead of being blended into fake consensus.

How persuasion is captured.

There are two different measurements. The AI score asks whether an argument was built to move a reasonable listener. The audience shift asks whether real people actually changed their minds.

The AI score

Argumentative force, 1 to 10.

Each brain looks for concrete stakes, a world the listener can picture and check, and an argument that can be understood the first time. It must name the specific argumentative move that earned or lost the score. The council publishes the panel median.

The fence

It never scores the speaker's voice.

Charm, confidence, volume, pace, fluency, polish, vocabulary, accent, and dialect are out of bounds. Persuasion cannot repair a missing reason or response. It affects the argument score and can decide only a substantive tie.

The human read

The room states its view twice.

In a live room, viewers can state their position before the round, flag the moments that moved them, and state it again after the ballot. The running tally stays hidden to avoid conformity. This audience shift is reported separately and never changes the AI verdict.

Three models, plus repeat runs.

Different model families test whether one lab's habits drove the call. Repeating an unchanged round tests whether sampling drove it. Agreement is stronger evidence for a verdict, not proof, so both measurements stay separate.

·
Loading measured agreement.
Why a split is published

A tie is not broken.

Any rule for breaking a tied panel is a thumb on the scale, and it would be our thumb. So an evenly split panel records no winner at all. You still get every juror's reasoning, the ladder does not move, and any predictions on the round are refunded at face value.

Refunding is the only side-neutral way to close a round nobody can decide.

Why this statistic

Agreement beyond chance.

A raw agreement percentage flatters any panel that leans one way: a judge that always picked Proposition would score 100% agreement while measuring nothing. The figure above nets that out, so it reflects agreement the models had to earn.

It stays unreported until 30 judged rounds. A number from four rounds is noise, and publishing noise as evidence would be the same overclaim this page exists to avoid.

Why the same round runs twice

Unchanged transcript, fresh read.

The stability test sends the exact same prompt twice and counts any changed verdict as unstable. Split external fixtures also test whether moving the two bench blocks or padding the losing side with repeated words changes the call.

Consented Debatable rounds enter the repeat test only after identifying details are scrubbed. Their earlier AI verdict is never treated as a correct answer, because that would score the judge against itself. Human-labeled tournament rounds remain the separate accuracy test.

We probed our own judge, and published what we found.

In August 2026 we ran the live panel against constructed speech pairs built to bait it: one variable changed at a time, judged repeatedly, 48 panel runs. Here is what held, and the one soft spot the probe caught.

Held

Substance beat polish, 12 of 12.

A polished, confident speech with a circular core lost to a flat speech with real warrants every single time, with a unanimous panel, in both side assignments. The rubric's rule that persuasion is never confidence, fluency, accent, or polish is doing real work; the suite that runs on every code change asserts that sentence survives.

Held

Padding a speech bought no wins.

The same arguments at 2.6 times the words, nothing new added: the padded losing side won 0 of its 8 rounds, and its own points barely moved. Length could not rescue a losing case.

The soft spot, stated plainly

Length buys clarity points, and nothing else.

Longer speeches earned roughly a third of a point more on the scorecard's clarity axis while their reasoning scores did not move. No verdict changed. But verdicts are not the only public number, so we moved the boundary: the win-loss verdict is the only thing that reaches ratings, credits, and settlement, and the public standings now rank on the head-to-head rating ladder rather than raw judge points, so that premium orders nobody.

Limits, so this is not oversold: one question, one casual 1v1 structure, 48 runs, and arguments we constructed ourselves. It shows the judge is consistent with its published rubric under manipulation; whether the rubric tracks what really persuades a room is a separate question we still owe an answer on.

Open to critiques.

A new kind of judge should expect skepticism. The three strongest objections we know, our current answers, and the one debt still open. If you hold a sharper objection, we want it.

The circularity objection

"An AI judging arguments rewards what AIs find persuasive."

The deepest objection, and partly right. Our answers so far: three separate model families have to agree, which is harder to dismiss as one model's taste; the criteria forbid scoring polish, confidence, or fluency; and the probe above showed substance beating polish under deliberate manipulation.

What we still owe: evidence that the criteria track what persuades a real room of people. That comparison has not been run, and until it has, this objection stays open on this page.

The ownership objection

"The house owns the judge."

True, and it is the reason this page exists. The criteria are fingerprinted so we cannot edit them after you speak. We take no rake, so no outcome pays us more than the other. Appeals go to a person we are forbidden from automating. Every decision writes a permanent record that is never edited in place.

The claim is not "trust us." The claim is that everything you would need to catch us is published and fetchable.

The bias objection

"Models carry biases debaters never chose."

Length, confidence, register, accent. Measured rather than waved away: the probe found longer speeches earn a small clarity premium and nothing else, and we published it. The persuasion axis is fenced by rule from scoring charm, fluency, accent, or dialect, and an automated test fails the build if that sentence is ever removed.

We claim these biases are measured and bounded, not absent. Nobody honest can claim absent.

Send yours

Try to break it. What breaks gets published.

If you can construct a pair of speeches that makes the panel reward the worse argument, that is the most useful thing you can send us. The strongest critiques get the probe treatment: run against the live panel, results published on this page whether they hold or not.

The form is just below. If your critique is about a round you were in, include the round id and we pull the transcript; an appeal on that round still goes through the round page, where a person reads it.

Send a critique

Tell us where the judge went wrong.

A pair of speeches, a round id, or one clear objection. Everything sent here is read by a person, and the strongest critiques are run against the live panel with the result published above whether it holds or not.

Form not loading? Open it in a new tab, or write to hello@itsdebatable.com.

The rubric.

Not a summary of it. This is the document the hash above fingerprints, rendered from the same endpoint the judge is pinned to.

Loading the published rubric.

You can appeal, and a person decides.

Real debate has an appeal route and nobody finds that strange. The route above an AI judge cannot be the same AI, or it is not a route.

Loading the appeal policy.

We earn the same whoever wins.

The strongest version of this promise is not a policy, it is the absence of any code that could do otherwise.

Loading the fee policy.

What gets written down.

Every ballot writes one permanent record the moment it is issued. Records are never edited; a correction is added beside the original, so a document cannot be quietly improved after someone questions it.

Loading the logging policy.

Check it yourself.

Both endpoints are open and need no key. The charter is the document the rubric hash on every ballot refers to, so you can fetch it, hash it, and confirm the criteria that judged a round are the criteria on this page.

Read the criteria, then argue.

Knowing exactly what the ballot rewards is not a loophole. It is the difference between practice and guesswork.

Start a round