Research corpus

A growing library of real, judged arguments.

Every round on Debatable leaves a record: the motion, what each side said, how the judge scored it. That library is what our AI learns from. Research labs can license an anonymized slice, built only from rounds users chose to share.

Training data the open web cannot produce: two sides directly answering each other, with a judge's verdict attached.

Internal corpus rounds
All rounds, the corpus the learning loop draws from
Voice-round transcripts
Real-time argumentation, no audio stored
1v1
Public round structure
One person on each side, one shared clock
Languages represented
Including Hindi, Spanish, Mandarin
Loading live counts…

Learn from the legends

The best rounds do not disappear. They become documents: transcripts, ballots, cases. The AI studies them, and every new debater inherits what the strongest ones figured out.

Round transcript · THBT social media has done more harm than good

Casual 1v1 · PRO vs CON · judged

PRO, rebuttal: Con says connection. Fine. Ask who owns the connection. When the feed decides what a billion people see before breakfast, that is not a town square, that is an editor nobody elected.

CON, reply: Would you also ban the printing press for having owners?

PRO: The press never knew your pulse. The feed does, and it optimizes for the spike, not the truth.

Sample page · what a transcript row looks like
Judge ballot · casual 1v1.docx

Pro wins · 2-1 panel

Reason for decision: Pro proves a concrete harm and explains why it outweighs Con's benefit. Con's final response never answers the mechanism.

Scores: Pro 88 · Con 75

Sample page · ballots carry the verdict and the why
Comparison note · Arctic security

Claim · reason · consequence

Claim: The icebreaker gap weakens Arctic deterrence.

The US operates two aging icebreakers against Russia's forty-plus; presence, not treaties, decides who writes the rules of new shipping lanes.

Sample page · one complete argument with evidence

Sample pages showing the shape of a corpus row. Real rows carry the full turn-by-turn text.

Why this dataset is hard to replicate

Real clash, not monologues

Most text on the internet is one side talking. Every row here answers the other side in real time, under a clock, often mid-interruption. Scraped op-eds, podcasts, and forums do not contain that.

Every response has context

The corpus keeps the claim being answered, the prior turns, the clock, and the final comparison. That makes direct rebuttal measurable in a way isolated essays are not.

Every round comes with a verdict

Users rate rounds one to five. Judges write ballots with speaker points and reasons. Researchers get the arguments plus which one won, no extra labeling work.

What's in the licensable subset

Research sharing is selected by default, with the terms in privacy §7. No rounds contribute until the person confirms they are 18 or older and saves with sharing on. Existing opt-outs stay off. The question appears after repeat use and explains the research and licensing purpose. Only future eligible typed rounds and voice transcripts carry a contributable: true flag; audio and video are excluded.

Each row, after anonymization, is shaped roughly:

{ motion: "THBT the means justify the ends", side: "PRO" | "CON", roundType: "casual_1v1", kind: "case" | "rebuttal" | "judge" | "voice_round" | "...", systemPrompt: "[casual 1v1 system block fed to the model]", userPrompt: "[user-side text + prior turns]", output: "[the AI's reply, or the user-turn block for human rows]", durationMs: 12340, context: { language: "en", persona: "debater", ... }, rating: 4, // 1-5, when given saved: false, contributable: true, // stamped at write time createdAt: "2026-05-25T17:34:01Z" }

Anonymized means stripped of name, email, account id, IP, and any device fingerprints. What remains is the speech and its structural metadata. Voice audio is never stored; only the text transcript is eligible.

One structure, many questions

Public rounds share one casual 1v1 structure. The useful variation comes from motions, languages, argument quality, interruptions, and judge outcomes.

The growth curve, not the row count

Volume today is small. What's compounding is the architecture: a learning loop that's been writing every generation to the corpus since 2026-05-13, a consent layer that went live 2026-05-25, and a daily distillation pass built to re-shape the AI from rated outputs, which is waiting on rating volume before it writes its first pattern set. The licensable subset is just starting. The wedge is what the dataset becomes at scale, not what it is this week.

License inquiries

Open to conversations with AI research orgs, academic labs, and dataset aggregators. Happy to share a sample export under NDA and walk through the schema.

Open support Read the consent terms

Cite or cover the project

Use itsdebatable.com/research as the stable public URL for the corpus. The press and media kit has verified product facts, official logos, and the media contact path.