Why This Matters
"I gave the criterion a level 6 because the essay was really good."
If you've ever said something close to that, you're not alone — and you're also not quite doing what the criterion asked. A newer teacher usually reads a criterion the way you'd read a title on a folder: as a general vibe of quality. Strong essay, level 6. Weak essay, level 3. It feels reasonable. It also means two students with very different work can land on the same level for completely different reasons, and you won't be able to explain why to a student, a parent, or a moderator.
The fix isn't more instinct. It's reading one level deeper — into the strands — before you ever open a student's work.
Understand the Idea
In the MYP (Middle Years Programme, the IB's curriculum framework for students roughly aged 11–16), every subject group is assessed against a small set of criteria — broad categories of learning, each given a letter (A, B, C, D) and a name that describes the kind of thinking or skill it covers. The exact labels, the number of criteria, and their focus are set by subject group and are revised on IB's own cycle, so treat any specific label as something to verify, not memorize from this article. 📌 Official source: your subject guide, current for this school year, is the only place to confirm exact criterion wording.
Here's the part that gets skipped: a criterion is not one thing to judge. It's broken into strands — usually labelled with small roman numerals (i, ii, iii) — and each strand describes a distinct piece of what a student needs to do. One strand might ask a student to describe, another to explain, another to evaluate. These verbs are not decorative. They're command terms: defined instructional verbs that specify the depth of thinking required, and they typically get more demanding as achievement levels rise within a strand.
This matters because achievement levels — the numeric bands (often something like 1–8, but confirm this for your subject and year band) that describe quality within a criterion — are almost always described strand by strand inside level bands, not as one holistic sentence. A level descriptor is usually built from several clauses, each one answering to a different strand. When you assess "the whole criterion" in one instinctive sweep, you're averaging across strands in your head without writing anything down — which is exactly how two very different pieces of work end up with the same level for different, unexamined reasons.
One more layer: strand demand shifts across the programme. A strand that asks students to "identify" in an early year of the MYP may ask them to "analyse" or "justify" by a later year, even under the same criterion letter and similar wording. This year-band progression (informally, something like years 1, 3, and 5 of the five-year programme, though your school's exact year structure may differ) is set centrally and revised periodically — so again, check your current guide rather than relying on what a colleague remembers from two years ago.
It's worth pausing on where criteria actually come from, because that origin story explains why they're written the way they are. Each subject group has a small set of published subject-group objectives — statements of what students should be able to do by the end of the programme in that subject. The assessment criteria are the classroom-facing version of those objectives: broken into criteria, broken further into strands, and finally described at each achievement level so a teacher has language to match against real work. This is a chain, not a coincidence of naming. Objective → criterion → strand → level descriptor. If you ever find yourself unsure why a strand is worded the way it is, the objective it descends from is usually the fastest way to understand its intent — and that chain is exactly why strand wording isn't arbitrary or interchangeable between subjects, even when two subjects use superficially similar language like "analyse" or "evaluate."
It also explains why you shouldn't expect criteria to look identical across subjects, or even across the two or three criteria within one subject. A criterion built around investigating and a criterion built around communicating are testing different objectives, so their strands ask for different things even when a command term happens to repeat. Don't assume that because you've internalized one subject's criteria, a second subject's criteria — even one you also teach — will follow the same pattern. Always start from that subject's own guide.
See It in Practice
Let's build a fully invented, illustrative-only criterion so you can see the mechanics without me pretending to quote an official document. Call it Criterion X: Investigating, from a subject group I'm making up for this example.
Imagine it has three strands:
- Strand i: describe a process or issue (command term: describe)
- Strand ii: explain causes or connections (command term: explain)
- Strand iii: evaluate the significance or reliability of information (command term: evaluate)
This is not real wording from any subject guide — it's a scaffold to show you how to read one.
Now picture a Grade 8 Individuals & Societies task from the "change and industrialisation" unit you built back in Chapter 03: students investigate how a specific technological change affected a community.
Weak approach (criterion-label assessment): The teacher reads the essay, thinks "this is solid, clear structure, good detail," and assigns a level 6 straight across the criterion.
Strong approach (strand-level assessment): The teacher asks three separate questions.
- Strand i — Did the student actually describe the process of change accurately and with relevant detail? (Yes — clear, detailed account of the shift from hand-loom to power-loom weaving.)
- Strand ii — Did the student explain why it happened and what it connected to — causes, consequences, relationships? (Partially — mentions economic pressure but doesn't connect it to specific labour conditions.)
- Strand iii — Did the student evaluate the reliability or significance of the sources used? (Barely — one line, no real judgement of source reliability.)
Now the picture looks different: strong on description, developing on explanation, weak on evaluation. That's not a single level 6 — it's a best-fit judgement across three unequal strand performances, which is a genuinely harder and more useful conversation to have with that student than "you got a 6."
Now push the same habit one step further, because strand-level reading doesn't stop at "did they do it" — it also has to ask "how well, compared to the actual level language." Suppose two students both attempt strand iii, the evaluate-reliability strand.
Student A writes: "This source might be biased because the author worked in the factory." That's a gesture toward evaluation — it names a reason to question the source — but it stops there. It doesn't say what the bias might distort, or how that limits what the source can responsibly be used to claim.
Student B writes: "This source was written by a factory owner defending working conditions, so it likely understates hardship; it's still useful for showing what management believed was acceptable, just not for judging what workers actually experienced." That's the same command term — evaluate — attempted at a visibly greater depth: it names the direction of the likely bias and it draws a specific, limited conclusion about what the source is and isn't good evidence for.
If your level descriptors distinguish (illustratively) between "some evaluation of reliability" at a lower band and "consistent, well-reasoned evaluation of reliability and significance" at a higher band, Student A sits closer to the lower band and Student B closer to the higher one — on this strand alone. This is the granularity that a whole-essay, single-impression read simply cannot produce, because "the essay felt strong overall" doesn't distinguish between a student who evaluates well in one paragraph and poorly in another.
Apply It
Time to do this with your own material — not the invented example above.
- List every strand under that criterion, in order, exactly as your guide currently states them.
- Next to each strand, write down its command term(s) — the specific verb(s) that define what depth of thinking is required.
- For each strand, write one sentence: "A student meeting this strand well would show me ___ in their response to this task." Be concrete — name the kind of evidence, not just "good work."
- Check your task instructions (the actual prompt you'll hand students). Does the wording use the same command terms as the strands? If your task says "describe" but the strand demands "evaluate," your task won't generate the evidence you need to award a high level on that strand.
Artefact to produce: A one-page strand breakdown for your task — strand, command term, evidence sentence — that you'll keep open while marking. This becomes the seed of the task-specific clarification you'll build in Article 23.
Avoid This
"I know this criterion from last year, I don't need to reread it." Subject guides and their criteria are revised on IB's cycle. What you remember may be superseded. Fix: reread the current guide for this unit, every year, even ones you've taught before.
Treating strands as always equally weighted. Some teachers assume each strand contributes exactly a third of the level (with three strands) or a quarter (with four). Best-fit judgement doesn't work by arithmetic averaging — it's a holistic, evidence-based call across the level descriptors as a whole. Fix: read the level descriptors themselves rather than inventing a percentage split.
Copying last year's task-specific wording onto this year's task without checking the strand wording still matches. Fix: always re-verify against the current guide before reusing anything.
Assuming command terms mean the same thing they'd mean in everyday English. "Analyse" in an assessment context implies a specific kind of structured breakdown, not just "talk about in detail." Fix: use your Command Term Reference Card, and when in doubt, check your official subject guide's command term list.
Check Yourself
Without opening this article again: can you name the three questions you should ask of any piece of student work before assigning a level (one per strand, roughly: what did they do, how well, against what command term)? If you can only describe "read it and get a feeling for the level," go back and re-run the Apply It steps on one real piece of student work.
Continue Learning
Next step → Article 22 tackles the habit most new teachers bring with them: converting levels into percentages in your head. Downloadable resource: Criterion Unpacking template + Command Term Reference Card. 📌 Official source: always confirm criteria, strands, and command terms against your current subject guide before finalizing any task or rubric.