Built for the Diploma Programme. Run by a swarm.
Swarm-grade feedback for IB DP essays — descriptor-pinned, evidence-cited, and honest about what an AI can and cannot judge.
DP assessment deserves better than a vibe score.
The IB’s published assessment criteria are admirably specific — and genuinely hard to self-apply. Students guess at what “sustained analysis” looks like in their subject. Supervisors drown in comment debt. And generic AI graders hand back a number with a pat on the head.
essaycriteria is built exclusively around the DP’s final written tasks: the Extended Essay in every subject group, the TOK essay on the current prescribed titles, and Group 1 Higher Level essays. No generic rubrics, no one-size-fits-all checklists — the correct official descriptor set loads for the exact task you submit, and every comment in your report is pinned to a criterion and a descriptor. If a finding can’t quote your own writing, it doesn’t ship.
For students
See exactly where your essay sits on every official descriptor — and what specifically moves it up a band. The roadmap tells you what to fix first, in your own next draft.
For supervisors
The swarm absorbs the first-pass marking load — the fifth draft of the fourth essay that still buries its thesis. You spend your comments where a human actually matters.
For schools & coordinators
Cohort-level band-gap heatmaps, bulk EE submissions, RPPF screening and draft-over-draft tracking — a department-wide view of where marks are being left on the table.
AI that coaches, never ghost-writes.
AI can now produce a passable essay in seconds — which is precisely why honest, rigorous feedback matters more than ever, and why a tool that writes for students is a tool that teaches nothing.
Our swarm only ever reads and reports. It diagnoses weak evidence, drifting registers and missing counter-claims the way a meticulous examiner would — then hands the pen back to you. Every report is built to be checked, not believed: every claim cites the line it refers to and the descriptor it is measured against, so you can argue with it, take it to your supervisor, and prove it wrong if you can.
The struggle of mastering the written word is the point. We exist to make sure the struggle is aimed at the right things.
Coach, not ghostwriter. essaycriteria never writes, rewrites or generates essay content — and its integrity scan flags citation anomalies and sudden register shifts, so problems surface before your supervisor finds them.
How the swarm reaches a verdict.
One AI grading alone drifts — between runs, between models, between moods of the same prompt. So we don’t rely on one. The swarm runs multiple frontier AI engines through multiple marking iterations, and only a cross-examined consensus survives into your report.
First read
Specialist agents — each running on the frontier engine best suited to its job — read your essay in parallel against the exact descriptor set for your task.
Cross-examination
Agents challenge each other’s findings over several iterations. The Devil’s Advocate hunts for counter-evidence; unsupported claims are struck.
Convergence
Band decisions must agree across engines and across rounds. Where they don’t, the disagreement is escalated — never averaged away.
Sign-off
The Chair, a chief-examiner agent, rules last. Nothing enters your report unless it cites your own text and names the descriptor it answers to.
Accuracy through consensus, not confidence. A single model asked twice can give two different bands. An ensemble that must argue its way to agreement — engine against engine, round after round — converges on the decision the descriptors actually support. That is the whole trick, and we built the entire product around it.