Methodology

How we evaluate, score, and classify every curriculum on the Science Ed Ledger.

How we evaluate

We read the curriculum - the student text, the labs, the teacher guides, the assessments - and extract every passage where a core science concept is taught. That includes passages about matter, atoms, chemical reactions, properties of materials, phase changes, energy, heat, sound, light, electricity, magnetism, cells, photosynthesis, rocks and minerals, stars, and other topics where the correct explanation involves atoms, molecules, electrons, or photons. We also extract observational passages (taxonomy, life cycles, anatomy, earth layers) to check for mapping errors and to assess subject-level anchoring. Every extracted passage is classified, the science is verified, and scores are computed. Sample data from our tables are available upon request. Our methodology is public. If you disagree with a score, the data is there to check.

Not every passage is scored the same way

Science curricula contain many different kinds of passages. Some make causal claims that require atomic-molecular grounding ("heat makes ice melt," "energy is stored in the battery"). Others describe and classify at the observation level ("frogs are amphibians," "the Earth has layers"). Our pipeline distinguishes between these because it would be unfair to penalize a K-1 lesson on animal classification for not mentioning atoms - that lesson doesn't need atoms to be correct.

Every passage is tagged as one of three types:

โš›
Mechanism - the passage makes a causal or explanatory claim where the correct explanation involves atoms, molecules, electrons, or photons. These passages are scored for the What-Why Gap. Examples: properties of materials, chemical reactions, phase changes, energy, heat, sound, cells, photosynthesis, rocks and minerals, stars.
๐Ÿ‘
Observational - the passage describes, classifies, or identifies at the organism, system, or phenomenon level. These passages are not scored for the gap - they don't need atoms to be correct. Examples: taxonomy, life cycles, habitats, anatomy, kinematics, moon phases, fossils, simple machines, "what is biology."
โ—
Boundary - the passage could be taught observationally at K-2 but requires atomic grounding at 3-5. These are resolved by grade band.

Observational passages can still contain mapping errors - a geology passage that says "heat makes the rock expand" is observational geology with a substance-model error. Every passage gets checked for mapping errors regardless of type. The passage type only affects the gap score.

What we measure

1. The What-Why Gap

For every mechanism passage, we ask: does this passage connect the concept to atoms, molecules, electrons, or photons? If the answer is no - if the curriculum asks students to explain why ice melts, why batteries die, why rocks are hard, why stars shine, without providing the atomic-level agents that would make a real explanation possible - that passage has a gap.

The gap score is the percentage of mechanism passages that lack atomic-molecular grounding. A 0% gap means every mechanism passage is grounded. A 100% gap means atoms are never introduced before explanations are expected. Observational passages are excluded from this calculation - they don't inflate the denominator.

We also check for scaffolding: if a passage doesn't name atoms in that specific sentence, but the same lesson establishes atomic grounding for the same concept elsewhere, we tag it as scaffolded and exclude it from the adjusted gap. The adjusted gap is the primary score - it measures what students actually experience, not whether every sentence repeats the word "atoms."

2. Mapping Errors

We check every science explanation for correct ontological categorization - is each concept placed in the right category? In science, some things are entities (atoms, molecules, electrons - things that exist and act), some things are processes (collisions, vibrations, rearrangements - things that entities do), and some things are emergent properties (heat, energy, force, sound, color - measurable outcomes that arise from what entities do). When a curriculum puts a concept in the wrong category, students build a mental model that is structurally wrong - not just imprecise, but categorically misaligned with how the science actually works.

For example: Is energy treated as a substance that can be stored and used up (entity), or as a measurable property that changes when atoms and molecules rearrange (emergent property)? Is force treated as something an object "has" (entity), or as the result of an interaction between objects (process)? Is heat treated as a fluid that flows from hot to cold (entity), or as the transfer of kinetic energy when faster-moving molecules collide with slower-moving molecules (process)?

These misplacements - what we call "mapping errors" - are not minor wording issues. Decades of research in science education has shown that when students categorize a process or emergent property as if it were a substance, the resulting misconceptions are uniquely resistant to correction. Chi, Slotta, and de Leeuw (1994) proposed that many of the most robust student misconceptions arise from placing concepts in the wrong ontological category - treating processes as things - and that correcting these misconceptions requires not just new information but an ontological category shift. Reiner, Slotta, Chi, and Resnick (2000) documented that students consistently reason about heat, light, and electric current as if they were material substances, and that this "substance-based commitment" persists even after direct instruction. Chi (2005) showed that misconceptions rooted in ontological miscategorization are more resistant to change than other types because the student's entire mental framework treats the concept as the wrong kind of thing - adding correct information to the wrong framework doesn't fix the framework.

This is why mapping errors matter more than they appear to. A curriculum that says "heat flows into the ice" is not just using informal language - it is placing heat in the substance category, and every student who reads that sentence adds one more piece of evidence to their mental model that heat is a stuff that moves around. That model will resist correction in high school physics because the student isn't holding a wrong fact - they're holding a wrong category.

We flag three types of mapping error:

โ—†
REI - Reification. An emergent property or process is treated as a substance or agent. Comes in two forms: REI-substance ("heat flows into the ice," "energy is stored in the battery" - the concept is given substance properties like flowing, being stored, being used up) and REI-agent ("inertia keeps the ball moving," "gravity pulls things down" - the concept is given causal agency, as if it acts on its own).
โ—†
LAE - Label as Explanation. A name is substituted for the mechanism. "This happens because of gravity" provides a label but no explanation of which entities interact and what they do. The student learns that knowing the name IS knowing the science.
โ—†
AEC - Alternative Entity Construction. The curriculum uses vague entities ("tiny particles," "building blocks") where specific entities (atoms, molecules) should be named. The student invents their own model - and those invented models inherit macroscopic properties (hot particles, heavy particles, colored particles).

The map error score is the percentage of student-facing passages containing at least one mapping error.

References: Chi, M. T. H., Slotta, J. D., & de Leeuw, N. (1994). From things to processes: A theory of conceptual change for learning science concepts. Learning and Instruction, 4(1), 27โ€“43. ยท Reiner, M., Slotta, J. D., Chi, M. T. H., & Resnick, L. B. (2000). Naive physics reasoning: A commitment to substance-based conceptions. Cognition and Instruction, 18(1), 1โ€“34. ยท Chi, M. T. H. (2005). Commonsense conceptions of emergent processes: Why some misconceptions are robust. Journal of the Learning Sciences, 14(2), 161โ€“199. ยท Slotta, J. D., & Chi, M. T. H. (2006). Helping students understand challenging topics in science through ontology training. Cognition and Instruction, 24(2), 261โ€“289.

3. Factual Errors (FSE)

Separately from framing, we verify whether the science itself is correct. A factual error is a statement that contradicts established science - not imprecise framing, but wrong. "T. rex evolved 170 million years ago" (it lived 68-66 mya). "Human cells have 3 billion base pairs on 46 chromosomes" (actually 6 billion on 46). An unbalanced chemical equation presented as balanced.

Factual errors are a disqualifier. Any curriculum with one or more confirmed factual errors receives an automatic โœ• Not Recommended verdict regardless of all other scores. A mapping error can be compensated for by a good teacher. A factual error teaches students something that is wrong.

4. Practices (informational)

We assess whether students engage in genuine scientific practices. Each practice is scored as Present, Partial, or Absent. The practices score is reported but does not affect the verdict. A program with excellent practices and wrong science still gets Not Recommended.

We evaluate seven practices adapted from the NGSS Science and Engineering Practices:

1. Asking QuestionsDo students ask or respond to investigable questions that drive the lesson - not just answer teacher prompts?
2. Developing and Using ModelsDo students build, use, or evaluate models (physical, visual, or conceptual) to explain phenomena - not just look at illustrations?
3. Planning and Carrying Out InvestigationsDo students help design investigations with controlled variables, or only follow pre-written procedures?
4. Analyzing and Interpreting DataDo students organize observations, identify patterns, or draw conclusions from data they collected?
5. Using MathematicsDo students measure, count, calculate, or use quantitative reasoning - not just qualitative description?
6. Constructing Explanations and Designing SolutionsDo students construct their own explanations of how or why something happens, grounded in evidence and mechanism? Engineering design is scored here, as the designing-solutions half of this practice, rather than as a separate eighth practice.
7. Engaging in Argument from EvidenceDo students debate competing explanations, evaluate claims against evidence, or defend their reasoning?

Each practice scored Present earns full credit (14.3%), Partial earns half credit (7.1%), and Absent earns zero. The total practices percentage and a color tier (green โ‰ฅ 70%, amber 30โ€“69%, red < 30%) are reported on each detail card.

Subject anchors

Beyond the passage-level scoring, we check whether each subject area in the curriculum establishes at least one explicit connection to atoms and molecules. A biology unit that says "living things are made of atoms and molecules arranged into cells" is anchored - even if most of its passages are observational. A biology unit that teaches accurate descriptive biology with zero mention of atoms anywhere is unanchored - the "clean but empty" pattern.

Anchored subjects get credit for connecting to the atomic foundation. Unanchored subjects flag a structural disconnect - students learn accurate facts that float free of the explanatory framework they'll need later.

The Foundation Score

The Foundation Score starts at 100 and deducts for problems found:

Each percentage point of gap costs 1 point. A curriculum with a 12% adjusted gap loses 12 points. The gap captures the proportion of mechanism passages where atomic grounding is missing - the wider the gap, the more the student is navigating without a starting point.

Each mapping error costs 3 points. A curriculum with 9 mapping errors loses 27 points. Map errors are weighted more heavily because they don't just fail to teach the right thing - they actively install wrong mental models. Research shows these ontological miscategorizations are uniquely resistant to later correction (Chi, 2005).

Any confirmed factual error = Not Recommended. FSE overrides the foundation score entirely. A curriculum that teaches wrong science cannot be recommended regardless of what else it gets right.

The formula: Foundation = 100 โˆ’ gap% โˆ’ (map error count ร— 3), capped at 0.

Verdict tiers

โ—
Deep (above 85%) โ†’ โœ“ Recommended. Foundations are sound. The science is right. Students are building accurate mental models from the ground up.
โ—
Mid-Deep (50โ€“85%) โ†’ โš  Conditional. Most of the science is right, but identifiable gaps or language errors need teacher attention. A knowledgeable teacher can compensate.
โ—
Mid-Shallow (25โ€“50%) โ†’ โš  Request Rewrite. Significant structural problems that go beyond what a teacher can easily patch. The publisher should be asked to revise.
โ—
Shallow (below 25%) โ†’ โœ• Not Recommended. The foundational problems are pervasive. Students will build misconceptions as a direct result of the curriculum's design.
โŠ˜
Any score + factual errors โ†’ โœ• Not Recommended. A unit that contains factual errors is automatically Not Recommended regardless of its foundation score. A curriculum that teaches something scientifically wrong cannot be recommended at any tier.

Depth classification

Every scored curriculum receives a depth classification on its detail page. The classification explains why the curriculum scores the way it does - not just what we found, but what produced it.

โ– 
Deep by design. Atoms-first is the foundational architecture. The curriculum was built around the principle that atomic-molecular theory comes first, and every subject is grounded in it. Removing the atomic foundation would gut the curriculum. Near-zero gap, near-zero mapping errors, anchored across all subjects.
โ– 
Shallow by default. The curriculum follows conventional sequencing - atoms deferred because that's how science has traditionally been taught, not because anyone argued it should be. No design documents defend the approach. The shallowness is fixable by editing passages without restructuring. Often published before atoms-first research made the case widely.
โ– 
Shallow by design. The shallowness is an intentional architectural choice. Design documents explicitly defer atoms. Teacher guides acknowledge that substance-model misconceptions will form and treat them as acceptable. Published frameworks specify delayed mechanism. The entire curriculum would need to be restructured to fix the problem - you cannot get there by editing passages.

Color legend

Colors used throughout the Science Ed Ledger:

โ—
Green - Deep / Recommended / above 85%
โ—
Purple - Mid-Deep / Conditional / 50โ€“85%
โ—
Orange - Mid-Shallow / Request Rewrite / Shallow by Default
โ—
Red - Shallow / Not Recommended / Shallow by Design / high gap or error rates
โ—
Green - Deep by Design (depth classification)
โ—
Dark badge - FSE flagged (factual errors)

What about NGSS alignment?

Many of the programs we evaluate are aligned to the Next Generation Science Standards. NGSS alignment and foundation depth are independent dimensions. A program can be fully NGSS-aligned and still score Shallow if it delays atomic-molecular theory beyond the point where students are asked to explain phenomena. NGSS alignment tells you the curriculum covers the right topics. Our scoring tells you whether the foundational science underneath those topics is right. They measure different things.

Science kits

Kits are evaluated on the same instrument as curricula, but a kit is not a curriculum and scoring it as though it were would produce a number that means nothing. A box of materials with a procedure card is not attempting to teach mechanism, and marking it down for failing to do something it never claimed to do is not measurement. So kits run through a three-stage process, and the first stage decides which of the later stages apply.

Stage 1: Classification by contents

We open the kit and read everything in it. The classification comes out of what is actually in the box. It is an output of the measurement, never a premise we start from, and never taken from how the product is shelved or categorised by a retailer.

Experiment Kit

Materials and procedure only

The kit supplies materials, a procedure to follow, and an expected result. There may be a short explanatory paragraph, but instruction is not the product. The student does the activity; the understanding is expected to come from somewhere else.

Lesson Kit

Activity embedded in instruction

The activity sits inside instructional content that grounds the mechanism: the kit explains what the student is looking at in terms of the entities doing the work, before or around the hands-on part. The instruction is part of the product.

Lesson Kit is a status a kit earns, not a label it can claim. A kit is classified as a Lesson Kit only after its instructional content has been through the full pipeline. Until that happens it is an Experiment Kit for our purposes, whatever the box says.

Stage 2: Practices, scored for every kit

Every kit, both types, is scored on the same seven-practice rubric used for curricula. For an Experiment Kit this is the primary reported score, and it is a real one. A well-built experiment kit can honestly earn green here: students handle materials, follow a procedure, collect observations, and compare a result to an expectation. That is doing science, and the score says so.

Practices 6 and 7 are where the What-Why Gap becomes visible without any curriculum scoring at all. Constructing explanations and arguing from evidence both require the student to have causal pieces to reason with. A kit that supplies materials and a result, and no account of what the atoms and molecules are doing, will score well on practices 1 through 5 and thin on 6 and 7. That pattern is the finding.

A known limitation of this rubric on kits. Practice 3, planning and carrying out investigations, structurally disadvantages kits. A kit is a fixed set of materials with a fixed procedure; that is what makes it a kit and what makes it usable by a parent at a kitchen table. A kit cannot score full marks on a practice that rewards students designing their own investigation without ceasing to be a kit. We report practice 3 for kits because the rubric is applied identically across every product we score, but a low practice-3 result on a kit describes the format, not a defect in that particular kit. Read the practice 6 and 7 results as the informative ones.

Stage 3: The full pipeline, where there is instruction to score

A kit whose contents include instructional content, a guide that explains the science rather than only directing the activity, runs the full pipeline on that content: passage extraction, What-Why Gap, REI / LAE / AEC mapping errors, factual-error verification, and a Foundation Score. The verdict thresholds are the same ones curricula are held to. A kit that passes this stage is reported as a Lesson Kit and carries a Foundation Score alongside its practices score.

The product's own claim sets the bar

What a product claims to be determines which pipeline it faces. A kit sold as a box of experiments is scored on practices and is not marked down for the absence of instruction it never offered. A kit marketed as curriculum, or as a complete science program, or as sufficient to teach a subject, is scored as curriculum, on the full pipeline, because that is the claim being made to the person spending the money. Where this rule changes what a product faces, the entry page says so directly and quotes the marketing language that triggered it. Publishers can therefore predict how their product will be scored before we score it, by reading their own catalog copy.

What the scoring looks like on a passage

The scores on this site come from individual passages, so the clearest way to show what the instrument does is to put a scored passage next to a treatment of the same content that would not be scored. Both columns below are aimed at the same grade level and the same phenomenon. The difference is not length, vocabulary difficulty, or reading level.

Scored: REI-substance + AEC
"Thermal energy is transferred from the hot water to the cold spoon. Energy always moves from warmer objects to cooler objects. The tiny particles in the spoon gain this energy and the spoon gets warmer."

What the pipeline flags. Two mapping errors in three sentences. The first is REI-substance: energy is described as a thing that is transferred and that moves from one object to another, and that the particles gain and hold. Energy is a measure of change, not cargo, and a student reading this acquires a mental model in which it is stuff. The second is AEC: "tiny particles" names no entity. The student supplies one, and the model they invent inherits properties of the objects they can see, so the particles become small hot things.

This passage is also flagged for the gap if the lesson has not established atoms and molecules as the entities in play before asking students to explain the warming.

Not scored: mechanism complete
"The water is made of water molecules, and the spoon is made of metal atoms. Both are moving, and the molecules in the hot water are moving faster than the atoms in the cool spoon. Where the water touches the spoon, fast-moving water molecules collide with the spoon's atoms and push them, so the spoon's atoms start moving faster and the water's molecules slow down. When we say the spoon got warmer, we mean its atoms are moving faster than they were. Temperature is how we measure that motion."

Why it passes. The entities are named: water molecules, metal atoms. What they do is named: they move, they collide, they push each other. The outcome is defined in terms of the entities and their motion rather than as a substance that arrived. Nothing crosses from the water into the spoon: the water's molecules push the spoon's atoms and slow down in doing it, and "warmer" is given a meaning the student can hold onto.

Same phenomenon, same grade band, and not harder to read. What changed is that the student is given the actors and the action, so the explanation is theirs to build rather than a phrase to repeat.

The left column is the pattern the instrument finds most often across the programs on this Ledger. The right column is what the same content looks like when the causal pieces are present. Neither column is a quotation from a specific product: the scored passage is a composite of the substance-and-vague-entity pattern that recurs in the evidence files, written out in full so the error types are visible in one place. Real quoted instances, identified by program and page, are on the individual entry pages.

Transparency

Evidence is published at three levels, and each level states what it contains and what it leaves out. Every entry page carries the scores, the pattern counts, and passages quoted verbatim from the curriculum. Every entry page also links a curated evidence file, 8 to 12 passages selected to illustrate each pattern found, with the selection rule printed alongside the download. The complete classified evidence table is not posted publicly, because publishing every extracted passage would reproduce a substantial portion of a copyrighted work. It is available on request to researchers, journalists, and to the publisher of the program being scored.

Challenges to a score are published with our response, including the ones we lose. The full policy and the running log are on the evidence and challenges page.

Request a Review

Have a curriculum, kit, or program you'd like us to review? Drop us a note, tell us what you want to see, we'll test it and post the results.

Every curriculum, kit, or program runs through the same evaluation algorithm. We post what we find.

If you don't like your score and want to challenge it, we'll show you a sample data set so you understand how we got your score. If you want to change your score, fix your program and send us a new file for a new review.

shallowbydesign@proton.me