How we evaluate, score, and classify every curriculum on the Science Ed Ledger.
We read the curriculum - the student text, the labs, the teacher guides, the assessments - and extract every passage where a core science concept is taught. That includes passages about matter, atoms, chemical reactions, properties of materials, phase changes, energy, heat, sound, light, electricity, magnetism, cells, photosynthesis, rocks and minerals, stars, and other topics where the correct explanation involves atoms, molecules, electrons, or photons. We also extract observational passages (taxonomy, life cycles, anatomy, earth layers) to check for mapping errors and to assess subject-level anchoring. Every extracted passage is classified, the science is verified, and scores are computed. Sample data from our tables are available upon request. Our methodology is public. If you disagree with a score, the data is there to check.
Science curricula contain many different kinds of passages. Some make causal claims that require atomic-molecular grounding ("heat makes ice melt," "energy is stored in the battery"). Others describe and classify at the observation level ("frogs are amphibians," "the Earth has layers"). Our pipeline distinguishes between these because it would be unfair to penalize a K-1 lesson on animal classification for not mentioning atoms - that lesson doesn't need atoms to be correct.
Every passage is tagged as one of three types:
Observational passages can still contain mapping errors - a geology passage that says "heat makes the rock expand" is observational geology with a substance-model error. Every passage gets checked for mapping errors regardless of type. The passage type only affects the gap score.
For every mechanism passage, we ask: does this passage connect the concept to atoms, molecules, electrons, or photons? If the answer is no - if the curriculum asks students to explain why ice melts, why batteries die, why rocks are hard, why stars shine, without providing the atomic-level agents that would make a real explanation possible - that passage has a gap.
The gap score is the percentage of mechanism passages that lack atomic-molecular grounding. A 0% gap means every mechanism passage is grounded. A 100% gap means atoms are never introduced before explanations are expected. Observational passages are excluded from this calculation - they don't inflate the denominator.
We also check for scaffolding: if a passage doesn't name atoms in that specific sentence, but the same lesson establishes atomic grounding for the same concept elsewhere, we tag it as scaffolded and exclude it from the adjusted gap. The adjusted gap is the primary score - it measures what students actually experience, not whether every sentence repeats the word "atoms."
We check every science explanation for correct ontological categorization - is each concept placed in the right category? In science, some things are entities (atoms, molecules, electrons - things that exist and act), some things are processes (collisions, vibrations, rearrangements - things that entities do), and some things are emergent properties (heat, energy, force, sound, color - measurable outcomes that arise from what entities do). When a curriculum puts a concept in the wrong category, students build a mental model that is structurally wrong - not just imprecise, but categorically misaligned with how the science actually works.
For example: Is energy treated as a substance that can be stored and used up (entity), or as a measurable property that changes when atoms and molecules rearrange (emergent property)? Is force treated as something an object "has" (entity), or as the result of an interaction between objects (process)? Is heat treated as a fluid that flows from hot to cold (entity), or as the transfer of kinetic energy when faster-moving molecules collide with slower-moving molecules (process)?
These misplacements - what we call "mapping errors" - are not minor wording issues. Decades of research in science education has shown that when students categorize a process or emergent property as if it were a substance, the resulting misconceptions are uniquely resistant to correction. Chi, Slotta, and de Leeuw (1994) proposed that many of the most robust student misconceptions arise from placing concepts in the wrong ontological category - treating processes as things - and that correcting these misconceptions requires not just new information but an ontological category shift. Reiner, Slotta, Chi, and Resnick (2000) documented that students consistently reason about heat, light, and electric current as if they were material substances, and that this "substance-based commitment" persists even after direct instruction. Chi (2005) showed that misconceptions rooted in ontological miscategorization are more resistant to change than other types because the student's entire mental framework treats the concept as the wrong kind of thing - adding correct information to the wrong framework doesn't fix the framework.
This is why mapping errors matter more than they appear to. A curriculum that says "heat flows into the ice" is not just using informal language - it is placing heat in the substance category, and every student who reads that sentence adds one more piece of evidence to their mental model that heat is a stuff that moves around. That model will resist correction in high school physics because the student isn't holding a wrong fact - they're holding a wrong category.
We flag three types of mapping error:
The map error score is the percentage of student-facing passages containing at least one mapping error.
References: Chi, M. T. H., Slotta, J. D., & de Leeuw, N. (1994). From things to processes: A theory of conceptual change for learning science concepts. Learning and Instruction, 4(1), 27โ43. ยท Reiner, M., Slotta, J. D., Chi, M. T. H., & Resnick, L. B. (2000). Naive physics reasoning: A commitment to substance-based conceptions. Cognition and Instruction, 18(1), 1โ34. ยท Chi, M. T. H. (2005). Commonsense conceptions of emergent processes: Why some misconceptions are robust. Journal of the Learning Sciences, 14(2), 161โ199. ยท Slotta, J. D., & Chi, M. T. H. (2006). Helping students understand challenging topics in science through ontology training. Cognition and Instruction, 24(2), 261โ289.
Separately from framing, we verify whether the science itself is correct. A factual error is a statement that contradicts established science - not imprecise framing, but wrong. "T. rex evolved 170 million years ago" (it lived 68-66 mya). "Human cells have 3 billion base pairs on 46 chromosomes" (actually 6 billion on 46). An unbalanced chemical equation presented as balanced.
Factual errors are a disqualifier. Any curriculum with one or more confirmed factual errors receives an automatic โ Not Recommended verdict regardless of all other scores. A mapping error can be compensated for by a good teacher. A factual error teaches students something that is wrong.
We assess whether students engage in genuine scientific practices. Each practice is scored as Present, Partial, or Absent. The practices score is reported but does not affect the verdict. A program with excellent practices and wrong science still gets Not Recommended.
We evaluate seven practices adapted from the NGSS Science and Engineering Practices:
| 1. Asking Questions | Do students ask or respond to investigable questions that drive the lesson - not just answer teacher prompts? |
| 2. Developing and Using Models | Do students build, use, or evaluate models (physical, visual, or conceptual) to explain phenomena - not just look at illustrations? |
| 3. Planning and Carrying Out Investigations | Do students help design investigations with controlled variables, or only follow pre-written procedures? |
| 4. Analyzing and Interpreting Data | Do students organize observations, identify patterns, or draw conclusions from data they collected? |
| 5. Using Mathematics | Do students measure, count, calculate, or use quantitative reasoning - not just qualitative description? |
| 6. Constructing Explanations and Designing Solutions | Do students construct their own explanations of how or why something happens, grounded in evidence and mechanism? Engineering design is scored here, as the designing-solutions half of this practice, rather than as a separate eighth practice. |
| 7. Engaging in Argument from Evidence | Do students debate competing explanations, evaluate claims against evidence, or defend their reasoning? |
Each practice scored Present earns full credit (14.3%), Partial earns half credit (7.1%), and Absent earns zero. The total practices percentage and a color tier (green โฅ 70%, amber 30โ69%, red < 30%) are reported on each detail card.
Beyond the passage-level scoring, we check whether each subject area in the curriculum establishes at least one explicit connection to atoms and molecules. A biology unit that says "living things are made of atoms and molecules arranged into cells" is anchored - even if most of its passages are observational. A biology unit that teaches accurate descriptive biology with zero mention of atoms anywhere is unanchored - the "clean but empty" pattern.
Anchored subjects get credit for connecting to the atomic foundation. Unanchored subjects flag a structural disconnect - students learn accurate facts that float free of the explanatory framework they'll need later.
The Foundation Score starts at 100 and deducts for problems found:
Each percentage point of gap costs 1 point. A curriculum with a 12% adjusted gap loses 12 points. The gap captures the proportion of mechanism passages where atomic grounding is missing - the wider the gap, the more the student is navigating without a starting point.
Each mapping error costs 3 points. A curriculum with 9 mapping errors loses 27 points. Map errors are weighted more heavily because they don't just fail to teach the right thing - they actively install wrong mental models. Research shows these ontological miscategorizations are uniquely resistant to later correction (Chi, 2005).
Any confirmed factual error = Not Recommended. FSE overrides the foundation score entirely. A curriculum that teaches wrong science cannot be recommended regardless of what else it gets right.
The formula: Foundation = 100 โ gap% โ (map error count ร 3), capped at 0.
Every scored curriculum receives a depth classification on its detail page. The classification explains why the curriculum scores the way it does - not just what we found, but what produced it.
Colors used throughout the Science Ed Ledger:
Many of the programs we evaluate are aligned to the Next Generation Science Standards. NGSS alignment and foundation depth are independent dimensions. A program can be fully NGSS-aligned and still score Shallow if it delays atomic-molecular theory beyond the point where students are asked to explain phenomena. NGSS alignment tells you the curriculum covers the right topics. Our scoring tells you whether the foundational science underneath those topics is right. They measure different things.
Kits are evaluated on the same instrument as curricula, but a kit is not a curriculum and scoring it as though it were would produce a number that means nothing. A box of materials with a procedure card is not attempting to teach mechanism, and marking it down for failing to do something it never claimed to do is not measurement. So kits run through a three-stage process, and the first stage decides which of the later stages apply.
We open the kit and read everything in it. The classification comes out of what is actually in the box. It is an output of the measurement, never a premise we start from, and never taken from how the product is shelved or categorised by a retailer.
The kit supplies materials, a procedure to follow, and an expected result. There may be a short explanatory paragraph, but instruction is not the product. The student does the activity; the understanding is expected to come from somewhere else.
The activity sits inside instructional content that grounds the mechanism: the kit explains what the student is looking at in terms of the entities doing the work, before or around the hands-on part. The instruction is part of the product.
Lesson Kit is a status a kit earns, not a label it can claim. A kit is classified as a Lesson Kit only after its instructional content has been through the full pipeline. Until that happens it is an Experiment Kit for our purposes, whatever the box says.
Every kit, both types, is scored on the same seven-practice rubric used for curricula. For an Experiment Kit this is the primary reported score, and it is a real one. A well-built experiment kit can honestly earn green here: students handle materials, follow a procedure, collect observations, and compare a result to an expectation. That is doing science, and the score says so.
Practices 6 and 7 are where the What-Why Gap becomes visible without any curriculum scoring at all. Constructing explanations and arguing from evidence both require the student to have causal pieces to reason with. A kit that supplies materials and a result, and no account of what the atoms and molecules are doing, will score well on practices 1 through 5 and thin on 6 and 7. That pattern is the finding.
A kit whose contents include instructional content, a guide that explains the science rather than only directing the activity, runs the full pipeline on that content: passage extraction, What-Why Gap, REI / LAE / AEC mapping errors, factual-error verification, and a Foundation Score. The verdict thresholds are the same ones curricula are held to. A kit that passes this stage is reported as a Lesson Kit and carries a Foundation Score alongside its practices score.
What a product claims to be determines which pipeline it faces. A kit sold as a box of experiments is scored on practices and is not marked down for the absence of instruction it never offered. A kit marketed as curriculum, or as a complete science program, or as sufficient to teach a subject, is scored as curriculum, on the full pipeline, because that is the claim being made to the person spending the money. Where this rule changes what a product faces, the entry page says so directly and quotes the marketing language that triggered it. Publishers can therefore predict how their product will be scored before we score it, by reading their own catalog copy.
The scores on this site come from individual passages, so the clearest way to show what the instrument does is to put a scored passage next to a treatment of the same content that would not be scored. Both columns below are aimed at the same grade level and the same phenomenon. The difference is not length, vocabulary difficulty, or reading level.
"Thermal energy is transferred from the hot water to the cold spoon. Energy always moves from warmer objects to cooler objects. The tiny particles in the spoon gain this energy and the spoon gets warmer."
What the pipeline flags. Two mapping errors in three sentences. The first is REI-substance: energy is described as a thing that is transferred and that moves from one object to another, and that the particles gain and hold. Energy is a measure of change, not cargo, and a student reading this acquires a mental model in which it is stuff. The second is AEC: "tiny particles" names no entity. The student supplies one, and the model they invent inherits properties of the objects they can see, so the particles become small hot things.
This passage is also flagged for the gap if the lesson has not established atoms and molecules as the entities in play before asking students to explain the warming.
"The water is made of water molecules, and the spoon is made of metal atoms. Both are moving, and the molecules in the hot water are moving faster than the atoms in the cool spoon. Where the water touches the spoon, fast-moving water molecules collide with the spoon's atoms and push them, so the spoon's atoms start moving faster and the water's molecules slow down. When we say the spoon got warmer, we mean its atoms are moving faster than they were. Temperature is how we measure that motion."
Why it passes. The entities are named: water molecules, metal atoms. What they do is named: they move, they collide, they push each other. The outcome is defined in terms of the entities and their motion rather than as a substance that arrived. Nothing crosses from the water into the spoon: the water's molecules push the spoon's atoms and slow down in doing it, and "warmer" is given a meaning the student can hold onto.
Same phenomenon, same grade band, and not harder to read. What changed is that the student is given the actors and the action, so the explanation is theirs to build rather than a phrase to repeat.
The left column is the pattern the instrument finds most often across the programs on this Ledger. The right column is what the same content looks like when the causal pieces are present. Neither column is a quotation from a specific product: the scored passage is a composite of the substance-and-vague-entity pattern that recurs in the evidence files, written out in full so the error types are visible in one place. Real quoted instances, identified by program and page, are on the individual entry pages.
Evidence is published at three levels, and each level states what it contains and what it leaves out. Every entry page carries the scores, the pattern counts, and passages quoted verbatim from the curriculum. Every entry page also links a curated evidence file, 8 to 12 passages selected to illustrate each pattern found, with the selection rule printed alongside the download. The complete classified evidence table is not posted publicly, because publishing every extracted passage would reproduce a substantial portion of a copyrighted work. It is available on request to researchers, journalists, and to the publisher of the program being scored.
Challenges to a score are published with our response, including the ones we lose. The full policy and the running log are on the evidence and challenges page.
Have a curriculum, kit, or program you'd like us to review? Drop us a note, tell us what you want to see, we'll test it and post the results.