Why This Matters
"My summative is a test, and half the class did fine on the test but couldn't actually explain the statement of inquiry when I asked them out loud."
That gap is the sound of a misaligned summative. The task produced marks, but not evidence of the thing the unit was actually about. This matters more than almost any other single design choice in the unit, because the summative task is where your assessment criteria — the official rubrics the IB provides for each subject group, describing what achievement looks like at different levels — meet real student work. If the task doesn't genuinely require what the criteria describe, the marks you generate don't mean what you think they mean, and the whole unit's evidence becomes unreliable.
This article builds the chain from criteria to task, so that a student literally cannot do well on the summative without having engaged with the statement of inquiry.
Understand the Idea
MYP assessment runs on a chain of four linked ideas, each one narrower than the last. Objectives are the broad, subject-group-wide statements of what a course of study should develop in students — set at subject-group level, not written by individual teachers. Criteria are the four assessed categories within a subject (each subject group has its own set, and the number and names vary by subject — MyIBHub won't assert specific criteria names here, since they differ across subjects and have been revised across guide editions; check your current subject guide). Strands are the specific, numbered sub-statements within each criterion — the actual line-by-line descriptions of what a student needs to show (for example, a strand might ask a student to "analyse" a source, while another strand in the same criterion asks them to "evaluate" it — different verbs, different cognitive demand, different evidence required). A task is the actual piece of work you hand a student, designed so that completing it well requires satisfying the strands you've selected.
The chain only works if you move down it in order: objectives are fixed by the subject group, so your real design job starts at criteria — which criteria, out of the several available in your subject, will this unit generate evidence for? — then strands — which specific numbered statements within those criteria will the task require? — then, only then, the task itself.
On criteria coverage across the year: a common misconception among new teachers is that every unit must assess all four criteria, every time. That's not accurate, and MyIBHub wants to correct it plainly: criteria coverage is planned across the year as a whole, not necessarily crammed into every single unit. Exactly how many times each criterion needs to be assessed, and over what period, is set out in current subject-specific guidance that has been revised over time — confirm the current requirement with your subject guide and your programme coordinator rather than assuming a fixed rule. What MyIBHub can say with confidence is the design principle: pick the criteria that genuinely fit what this unit is trying to teach, not the criteria that are simply "due" this term.
Command terms are another piece of the chain worth naming explicitly. These are the specific verbs used in criteria strands and in task instructions — "analyse," "evaluate," "discuss," "justify" — each with a particular cognitive demand attached to it in official IB documentation. If your task asks students to "describe" but the strand you selected requires "evaluate," the task cannot generate the evidence the strand needs, no matter how well students complete it. MyIBHub recommends checking your subject guide's command term glossary before finalising task wording, since exact definitions are set at subject-group level and can vary slightly between subjects.
Two more ideas complete the picture. Authenticity means the task resembles something a person outside the exam hall would actually need to produce or do — a real genre, a real audience, a real purpose — rather than an artificial exercise that only exists to be marked. Access and equity means the task is designed so that a student's ability to demonstrate the criteria isn't blocked by something unrelated to the criteria themselves — reading load unrelated to the subject, cultural assumptions students haven't been taught, format demands (e.g., timed handwriting) that disadvantage some learners for reasons that have nothing to do with what's being assessed. Good task design builds access in from the start — different entry points, scaffolded stages, choice of format — rather than retrofitting accommodations afterward.
See It in Practice
Here's the traditional-test version of the Grade 8 I&S industrialisation unit — the one most of us would have written without thinking too hard about it — and the criterion-aligned performance task it became.
Before — the traditional test. Forty-five minutes, in class, closed book. Ten short-answer questions: "Name three changes in production during the Industrial Revolution." "Explain one cause of urbanisation in this period." "Give two effects of child labour." Students who'd memorised the textbook chapter did well. Nobody had to use a source. Nobody had to weigh one account against another. Nobody had to make an argument. The statement of inquiry — about groups competing over whose account of progress gets believed — never actually had to be engaged with to score full marks. A student could ace this test having forgotten the SOI existed.
After — the criterion-aligned performance task. Students are given three primary and secondary sources on industrialisation: an 1834 letter from a mill owner celebrating productivity gains, an extract from a British parliamentary inquiry into factory conditions featuring worker testimony, and a short passage from a modern historian arguing that "progress" in this period depended entirely on whose suffering was made visible and whose was not. The task: produce an evidence-based argument document — in the form of a briefing paper for a museum exhibit — that takes a position on whose account of "progress" during industrialisation deserves the most weight, using evidence from all three sources and explicitly addressing at least one perspective that complicates your own position.
This task is authentic — a museum briefing paper is a real genre with a real audience (curators deciding what the exhibit says), not an invented school-only format. It cannot be done well without engaging the SOI, because the task's actual demand — weighing competing accounts — is the SOI, operationalised. And it maps cleanly onto strands a teacher could select from her subject's criteria (again, MyIBHub is deliberately not asserting specific criterion letters or strand numbers here, since these vary by subject and guide version — check your current I&S subject guide).
Mapping table — which part of the task evidences which strand (illustrative, using generic strand descriptions rather than specific official numbering):
| Task component | What it evidences |
|---|---|
| Selecting and using evidence from all three sources | A strand requiring students to select and use relevant sources |
| Explicitly weighing the mill owner's account against the worker testimony | A strand requiring analysis of different perspectives on an issue |
| Taking and justifying a position on whose account deserves more weight | A strand requiring a supported argument or evaluation |
| Addressing a perspective that complicates the student's own position | A strand requiring consideration of counter-perspectives or limitations |
It's worth sitting with why the "before" version felt like a reasonable thing to write in the first place — because it probably would have been fine by the standards most of us were trained under outside the MYP. A closed-book recall test is efficient to write, quick to mark, and produces a tidy spread of scores. None of that is a criticism of the teacher who wrote it; it's a description of what most non-criterion-based assessment cultures optimise for. The shift the MYP asks for is a genuine one: optimise for evidence of the criteria strands instead, even when that evidence is messier to produce and slower to mark. The museum-briefing task takes longer to design and longer to grade than the ten-question test. It also tells you something the test never could — whether a student can actually weigh competing accounts, which is the one thing this entire unit was built to teach.
Apply It
Teacher Activity: Map your draft task to strands — then fix the gaps.
- Take your own unit's draft summative task (or sketch one now, using your SOI and inquiry questions from Articles 12 and 13 as the source of what it should require).
- List the criteria strands you intend to assess, in your own paraphrase, from your current subject guide.
- Go through your task instructions line by line and write the strand each line is meant to evidence, the way the table above does.
- Flag any strand on your list that has no task instruction generating evidence for it. This is a design fault, not a minor gap — a strand you've selected but can't actually assess isn't providing evidence, whatever the mark scheme implies.
- For each flagged strand, either add or adjust a task instruction so the strand is genuinely evidenced, or remove the strand from this unit's selection and plan to assess it elsewhere in the year.
Artefact to produce now: A short mapping table (even three or four rows, like the one above) linking your task instructions to your selected strands, plus a one-line note on any strand you removed and where in the year it'll be picked up instead.
Avoid This
The test that measures memory instead of the criteria. Short-answer recall tests are easy to mark but rarely evidence anything beyond factual recall strands. Fix: check whether your task requires analysis, evaluation, or argument — not just retrieval — if your selected strands demand those.
The task with no real audience or genre. "Write an essay about industrialisation" is authentic to nothing. Fix: give the task a real form (briefing paper, podcast script, exhibit label, letter to an editor) and a real intended reader.
Assessing every criterion in every unit "to be safe." This overloads the task, the marking, and the students, and it's based on a misconception. Fix: select the criteria that genuinely fit this unit's SOI and plan the rest across the year — confirm current coverage expectations with your subject guide rather than defaulting to "all four, every time."
No task-specific clarification. Handing students the generic subject-wide criteria language and expecting them to infer what it means for this task disadvantages students who haven't yet learned to translate rubric language. Fix: draft a short task-specific clarification — Article 23 covers this in depth.
Check Yourself
Pick one strand you've selected for your unit's summative. Point to the exact sentence in your task instructions that a student would need to complete in order to generate evidence for it. If you can't point to one, the strand isn't actually being assessed yet.
Continue Learning
→ Next: Article 15 assembles everything — concepts, SOI, questions, task and criteria — into one complete, teachable unit plan.