Research corpus

A growing library of real, judged arguments.

Every round on Debatable leaves a record: the motion, what each side said, how the judge scored it. That library teaches our AI something new every night. Research labs can license an anonymized slice, built only from rounds users chose to share.

Training data the open web cannot produce: two sides actually clashing, in real formats, with a judge's verdict attached.

Internal corpus rounds
All rounds, used by the nightly learning loop
Voice-round transcripts
Real-time argumentation, no audio stored
Formats actively captured
APDA, BP, WSDC, Asian, PF, LD, Policy, Congress, MUN
Languages represented
Including Hindi, Spanish, Mandarin
Loading live counts…

Learn from the legends

The best rounds do not disappear. They become documents: transcripts, ballots, cases. The AI studies them nightly, and every new debater inherits what the strongest ones figured out.

Round transcript · THBT social media has done more harm than good

Quick Clash · PRO vs CON · judged

PRO, rebuttal: Con says connection. Fine. Ask who owns the connection. When the feed decides what a billion people see before breakfast, that is not a town square, that is an editor nobody elected.

CON, POI: Would you also ban the printing press for having owners?

PRO: The press never knew your pulse. The feed does, and it optimizes for the spike, not the truth.

Sample page · what a transcript row looks like
Judge ballot · BP final.docx

Closing Government wins · 3-2 panel

RFD: CG takes the round on the extension: quantified harm with a mechanism OG never provided. Opening Opposition's principle argument was strong but unweighed against practice.

Speaks: PM 77 · LO 78 · MG 81 · MO 76

Sample page · ballots carry the verdict and the why
1AC · Arctic security (Policy)

Tagged evidence · contention one

Tag: Icebreaker gap collapses US Arctic deterrence.

The US operates two aging icebreakers against Russia's forty-plus; presence, not treaties, decides who writes the rules of new shipping lanes.

Sample page · Policy keeps its cards, APDA stays impromptu

Sample pages showing the shape of a corpus row. Real rows carry the full turn-by-turn text.

Why this dataset is hard to replicate

Real clash, not monologues

Most text on the internet is one side talking. Every row here answers the other side in real time, under a clock, often mid-interruption. Scraped op-eds, podcasts, and forums do not contain that.

Every format speaks its own language

Policy reads evidence cards. APDA improvises. BP builds extensions. LD argues from a value framework. The corpus captures debaters switching styles between formats, a signal generic web text never shows.

Every round comes with a verdict

Users rate rounds one to five. Judges write ballots with speaker points and reasons. Researchers get the arguments plus which one won, no extra labeling work.

What's in the licensable subset

Opt-in only. The toggle lives in every user's profile, off by default, with the legal terms in privacy §6. When a user turns it on, future rounds (typed and voice) carry a contributable: true flag; everything else stays internal.

Each row, after anonymization, is shaped roughly:

{ motion: "THBT the means justify the ends", side: "GOV" | "OPP" | "PRO" | "CON" | "AFF" | "NEG" | "...", format: "apda" | "bp" | "worlds" | "pf" | "ld" | "policy" | "...", kind: "case" | "rebuttal" | "judge" | "voice_round" | "...", systemPrompt: "[format-aware system block fed to the model]", userPrompt: "[user-side text + prior turns]", output: "[the AI's reply, or the user-turn block for human rows]", durationMs: 12340, context: { language: "en", persona: "debater", ... }, rating: 4, // 1-5, when given saved: false, contributable: true, // stamped at write time createdAt: "2026-05-25T17:34:01Z" }

Anonymized means stripped of name, email, account id, IP, and any device fingerprints. What remains is the speech and its structural metadata. Voice audio is never stored; only the text transcript is eligible.

Per-format internal counts

Snapshot from the last nightly aggregation. Includes all generations, not just the opt-in subset, so you can see where the volume is concentrated.

Loading…

The growth curve, not the row count

Volume today is small. What's compounding is the architecture: a learning loop that's been writing every generation to the corpus since 2026-05-13, a consent layer that went live 2026-05-25, and a daily distillation pass that re-shapes the AI based on rated outputs. The licensable subset is just starting. The wedge is what the dataset becomes at scale, not what it is this week.

License inquiries

Open to conversations with AI research orgs, academic labs, and dataset aggregators. Happy to share a sample export under NDA and walk through the schema.

aidandavidhollinger@gmail.com Read the consent terms