A growing library of real, judged arguments.
Every round on Debatable leaves a record: the motion, what each side said, how the judge scored it. That library teaches our AI something new every night. Research labs can license an anonymized slice, built only from rounds users chose to share.
Training data the open web cannot produce: two sides actually clashing, in real formats, with a judge's verdict attached.
Learn from the legends
The best rounds do not disappear. They become documents: transcripts, ballots, cases. The AI studies them nightly, and every new debater inherits what the strongest ones figured out.
PRO, rebuttal: Con says connection. Fine. Ask who owns the connection. When the feed decides what a billion people see before breakfast, that is not a town square, that is an editor nobody elected.
CON, POI: Would you also ban the printing press for having owners?
PRO: The press never knew your pulse. The feed does, and it optimizes for the spike, not the truth.
RFD: CG takes the round on the extension: quantified harm with a mechanism OG never provided. Opening Opposition's principle argument was strong but unweighed against practice.
Speaks: PM 77 · LO 78 · MG 81 · MO 76
Tag: Icebreaker gap collapses US Arctic deterrence.
The US operates two aging icebreakers against Russia's forty-plus; presence, not treaties, decides who writes the rules of new shipping lanes.
Sample pages showing the shape of a corpus row. Real rows carry the full turn-by-turn text.
Why this dataset is hard to replicate
Real clash, not monologues
Most text on the internet is one side talking. Every row here answers the other side in real time, under a clock, often mid-interruption. Scraped op-eds, podcasts, and forums do not contain that.
Every format speaks its own language
Policy reads evidence cards. APDA improvises. BP builds extensions. LD argues from a value framework. The corpus captures debaters switching styles between formats, a signal generic web text never shows.
Every round comes with a verdict
Users rate rounds one to five. Judges write ballots with speaker points and reasons. Researchers get the arguments plus which one won, no extra labeling work.
What's in the licensable subset
Opt-in only. The toggle lives in every user's profile, off by default, with the legal terms in privacy §6. When a user turns it on, future rounds (typed and voice) carry a contributable: true flag; everything else stays internal.
Each row, after anonymization, is shaped roughly:
Anonymized means stripped of name, email, account id, IP, and any device fingerprints. What remains is the speech and its structural metadata. Voice audio is never stored; only the text transcript is eligible.
Per-format internal counts
Snapshot from the last nightly aggregation. Includes all generations, not just the opt-in subset, so you can see where the volume is concentrated.
The growth curve, not the row count
Volume today is small. What's compounding is the architecture: a learning loop that's been writing every generation to the corpus since 2026-05-13, a consent layer that went live 2026-05-25, and a daily distillation pass that re-shapes the AI based on rated outputs. The licensable subset is just starting. The wedge is what the dataset becomes at scale, not what it is this week.
License inquiries
Open to conversations with AI research orgs, academic labs, and dataset aggregators. Happy to share a sample export under NDA and walk through the schema.
aidandavidhollinger@gmail.com Read the consent terms