Debatable.com Back to the AI judge
How the AI judge is kept honest

The criteria are published before you speak.

An AI judge that the people running the site can quietly retune is not a judge, it is a house edge. So the rubric is written down and fingerprinted before the round, three independent models decide it instead of one, their level of agreement is published even when it is bad, and a human can overturn any ballot. Everything on this page is served from an open endpoint you can fetch yourself.

The short version. Fixed criteria, published in advance. Three model families vote, and a split panel records no winner rather than picking one. A person, never a model, hears appeals. We earn the same whoever wins.

What is judging your round, right now.

The season pins the rubric and the models. The hash is a fingerprint of the criteria: if a word of the rubric changes, the hash changes, so nobody can be judged against criteria that were edited after they spoke.

Loading
Reading the live charter.

Three models, not one opinion.

Three calls to the same model measure its mood. Three different model families measure whether the round was actually decidable. When they agree, the verdict is a fact about your debate. When they split, it was a fact about the judge, and we say so.

·
Loading measured agreement.
Why a split is published

A tie is not broken.

Any rule for breaking a tied panel is a thumb on the scale, and it would be our thumb. So an evenly split panel records no winner at all. You still get every juror's reasoning, the ladder does not move, and any predictions on the round are refunded at face value.

Refunding is the only side-neutral way to close a round nobody can decide.

Why this statistic

Agreement beyond chance.

A raw agreement percentage flatters any panel that leans one way: a judge that always picked Proposition would score 100% agreement while measuring nothing. The figure above nets that out, so it reflects agreement the models had to earn.

It stays unreported until 30 judged rounds. A number from four rounds is noise, and publishing noise as evidence would be the same overclaim this page exists to avoid.

The rubric.

Not a summary of it. This is the document the hash above fingerprints, rendered from the same endpoint the judge is pinned to.

Loading the published rubric.

You can appeal, and a person decides.

Real debate has an appeal route and nobody finds that strange. The route above an AI judge cannot be the same AI, or it is not a route.

Loading the appeal policy.

We earn the same whoever wins.

The strongest version of this promise is not a policy, it is the absence of any code that could do otherwise.

Loading the fee policy.

What gets written down.

Every ballot writes one permanent record the moment it is issued. Records are never edited; a correction is added beside the original, so a document cannot be quietly improved after someone questions it.

Loading the logging policy.

Check it yourself.

Both endpoints are open and need no key. The charter is the document the rubric hash on every ballot refers to, so you can fetch it, hash it, and confirm the criteria that judged a round are the criteria on this page.

Read the criteria, then argue.

Knowing exactly what the ballot rewards is not a loophole. It is the difference between practice and guesswork.

Start a round