"However, in the end, I think that, just as we cannot understand many properties of matter without atomic and sub-atomic understanding, we cannot clearly understand many of the commonly used terms for forms of energy until we break them down again into the underlying particles and their interactions." Helen R. Quinn, Teaching and Learning of Energy in K–12 Education, p. 17

Who We Are

The Science Ed Ledger scores K–8 science programs on one specific thing: whether they connect what students are asked to explain to the atoms, molecules, electrons, and photons that actually do the explaining. That is the lens. It is not a general rating of program quality, and it does not claim to be neutral about depth. It is openly committed to the idea that causal grounding is what makes early science hold up later.

Where this came from

The instrument was not originally built to grade other people's programs. I built it to audit my own. I am Rebecca Woodbury, the author of Real Science-4-Kids (RS4K), an atoms-first K–8 curriculum, and I needed to find where my own books drifted. Where a concept got named but never grounded. Where the language slipped into treating energy or heat as a substance that flows. Where a fact had gone stale, which happened most in astronomy, because the numbers keep changing. The instrument was built to catch those things in my own work first.

The word "shallow" came from a customer describing what was missing in the programs they had already tried. Once the instrument worked on my books, the obvious question was whether it would say anything useful about anyone else's. So we are using the exact same pipeline, with nothing changed, to separate programs that are shallow (lack mechanistic explanations) from those that are not. We are also looking at the curricular design to see whether the shallow design is intentional (stated in the materials as a design choice) or by default (this is what everyone does so we do it too).

The disclosure, plainly

Because I am the person who built the scoring instrument and also wrote one of the programs it scores, there is a real conflict of interest. RS4K passes, and it is not the only program that does. Programs from other publishers pass too, and units within programs that score poorly overall sometimes have strong individual results. The instrument exists to give parents and educators a way to evaluate the programs on the market and make an informed choice based on what is actually in them. RS4K sits on this Ledger and runs through the same pipeline as every other program, with no separate track.

If NGSS alignment is a priority, we have noted each publisher's stated alignment and EdReports can provide full detail. If practices are the priority, we have scored those. If you want something deeper, something that takes a student from basic exposure to fundamental understanding of science concepts, we have provided the metric to evaluate where programs differ and how they rank.

Why we are testing it

One more thing is worth saying plainly. Every program on this Ledger, including the ones that score well, markets itself as effective. What none of them has done is subject its instructional sequence to an empirical test that could show it does not work. That includes RS4K, so far. The difference is that we are building that test. We are developing a research study designed to measure whether atoms-first instruction actually reduces the misconceptions our instrument identifies, or whether it does not. The study could confirm the thesis behind this curriculum, or it could disprove it. We do not know the answer yet. That is the point.

Building a scoring instrument is one kind of accountability. Subjecting your own program to a study that could prove you wrong is another. We chose both.

What we promise

We are not asking you to trust that we are independent. We are asking you to check.

The method is public. Every score is built from the same steps, and those steps are written out in full, including the exact formula, on the Methodology page. Nothing about how a score is reached is hidden.

Everyone runs through the same pipeline. There is no gentler path for any publisher, including the one I wrote. The same passages get extracted, the same errors get counted, the same thresholds decide the verdict.

You can challenge any score. If you think a result is wrong, we will show you a sample of the underlying passage data, and anyone is welcome to run the same analysis and dispute what we found.

What this is not

This is one instrument, measuring one property, run by people who are open about who they are. It is not a consensus body and it is not the last word. The longer-term hope is that evaluation like this expands, with more than one set of hands on it. That does not exist yet, and we are not going to pretend it does. Until it does, the honest description is the one above: a transparent tool, applied evenly, with its own author's program in the dataset and open to challenge.

The full scoring method, including every error type and how the Foundation Score is calculated, is on the Methodology page.

← Back to Science Ed Ledger

Request a Review

Have a curriculum, kit, or program you'd like us to review? Drop us a note, tell us what you want to see, we'll test it and post the results.