A growing library of real, judged arguments.
Every round on Debatable leaves a record: the motion, what each side said, how the judge scored it. That library is what our AI learns from. Research labs can license an anonymized slice, built only from rounds users chose to share.
Training data the open web cannot produce: two sides directly answering each other, with a judge's verdict attached.
Learn from the legends
The best rounds do not disappear. They become documents: transcripts, ballots, cases. The AI studies them, and every new debater inherits what the strongest ones figured out.
PRO, rebuttal: Con says connection. Fine. Ask who owns the connection. When the feed decides what a billion people see before breakfast, that is not a town square, that is an editor nobody elected.
CON, reply: Would you also ban the printing press for having owners?
PRO: The press never knew your pulse. The feed does, and it optimizes for the spike, not the truth.
Reason for decision: Pro proves a concrete harm and explains why it outweighs Con's benefit. Con's final response never answers the mechanism.
Scores: Pro 88 · Con 75
Claim: The icebreaker gap weakens Arctic deterrence.
The US operates two aging icebreakers against Russia's forty-plus; presence, not treaties, decides who writes the rules of new shipping lanes.
Sample pages showing the shape of a corpus row. Real rows carry the full turn-by-turn text.
Why this dataset is hard to replicate
Real clash, not monologues
Most text on the internet is one side talking. Every row here answers the other side in real time, under a clock, often mid-interruption. Scraped op-eds, podcasts, and forums do not contain that.
Every response has context
The corpus keeps the claim being answered, the prior turns, the clock, and the final comparison. That makes direct rebuttal measurable in a way isolated essays are not.
Every round comes with a verdict
Users rate rounds one to five. Judges write ballots with speaker points and reasons. Researchers get the arguments plus which one won, no extra labeling work.
What's in the licensable subset
Research sharing is selected by default, with the terms in privacy §7. No rounds contribute until the person confirms they are 18 or older and saves with sharing on. Existing opt-outs stay off. The question appears after repeat use and explains the research and licensing purpose. Only future eligible typed rounds and voice transcripts carry a contributable: true flag; audio and video are excluded.
Each row, after anonymization, is shaped roughly:
Anonymized means stripped of name, email, account id, IP, and any device fingerprints. What remains is the speech and its structural metadata. Voice audio is never stored; only the text transcript is eligible.
One structure, many questions
Public rounds share one casual 1v1 structure. The useful variation comes from motions, languages, argument quality, interruptions, and judge outcomes.
The growth curve, not the row count
Volume today is small. What's compounding is the architecture: a learning loop that's been writing every generation to the corpus since 2026-05-13, a consent layer that went live 2026-05-25, and a daily distillation pass built to re-shape the AI from rated outputs, which is waiting on rating volume before it writes its first pattern set. The licensable subset is just starting. The wedge is what the dataset becomes at scale, not what it is this week.
License inquiries
Open to conversations with AI research orgs, academic labs, and dataset aggregators. Happy to share a sample export under NDA and walk through the schema.
Open support Read the consent termsCite or cover the project
Use itsdebatable.com/research as the stable public URL for the corpus. The press and media kit has verified product facts, official logos, and the media contact path.