How to Evaluate Whether a Literacy Program Is Truly Evidence-Based

July 20, 2026

Six questions every curriculum committee should ask before signing a contract, because "evidence-based" is a starting point for evaluation, not the end of it.

Title graphic reading 'How to Evaluate Whether a Literacy Program Is Truly Evidence-Based,' beside an illustrated clipboard with a blue, green, yellow, and red checklist and a magnifying glass.

To judge whether a literacy program is truly evidence-based, look past the label and ask six questions: what evidence tier it meets, whether that evidence is independent, whether its design reflects the Science of Reading, whether it was studied with students like yours, what it takes to implement well, and how you will measure results locally.

Why "evidence-based" on the brochure isn't enough

Many literacy programs are now described as evidence-based, research-based, or aligned to the Science of Reading. Many use them loosely. For a curriculum committee weighing a contract that will shape instruction for years and cost a great deal, the phrase on the cover is only where the question begins.

Good curriculum matters, but no program succeeds on its own. Strong outcomes come from the interaction between high-quality instructional materials, knowledgeable teachers, supportive leadership, and consistent implementation. These six questions are designed to help committees evaluate not just whether a program has evidence behind it, but whether it is positioned to succeed in their own schools.

No single piece of evidence tells the whole story. A program can have a solid efficacy study but instructional design that the research has moved past. It can have beautiful, research-aligned design but no study showing it works. It can have both and still be impossible for your schools to implement well. Six questions, asked before the contract is signed, catch most of these mismatches.

The six questions at a glance

QuestionWhat a strong answer looks like
1. What evidence tier does it meet?A named ESSA tier (1–3) backed by a real study on this product
2. Is the evidence independent?Reviewed by an independent third party rather than the vendor alone
3. Does its design reflect the Science of Reading?Systematic phonics, decodable texts, knowledge-building; not cueing
4. Was it studied with students like ours?Similar grade bands, demographics, and settings
5. What does implementing it well require?Training, coaching, time, and materials you can actually provide
6. How will we measure results here?A local plan tracking teacher practice and student outcomes

1. What evidence tier does it meet, and for this product?

Start with the formal question. Ask the vendor which ESSA evidence tier the program meets, in writing, and for the exact product and edition you would buy. "Evidence-based" with no tier is not an answer, and a study on an earlier version or a different grade band does not transfer.

A strong answer names a tier (Strong, Moderate, or Promising) and points to a specific study showing a statistically significant, positive effect. A weak answer offers testimonials, a logic model, or the phrase "research-based," none of which is an evidence tier. If you are spending school-improvement funds, remember the program must reach Tier 1, 2, or 3, not Tier 4.

2. Is the evidence independent?

Who ran the study matters as much as the result. A study designed, funded, and published by the vendor still has value, but it carries an obvious incentive, and no one with distance from the sale has vetted it.

A strong answer points to evidence reviewed by an independent body, such as the What Works Clearinghouse or the Evidence for ESSA database, or to peer-reviewed research the vendor did not control. A weak answer rests entirely on the vendor's own internal numbers. Ask, plainly, who conducted the study and who paid for it.

3. Does its design reflect the Science of Reading?

Evidence of outcomes and quality of design are different things, and a committee needs both. A program might have an old efficacy study yet still teach children to guess words from context, or use leveled and predictable texts that the research has moved away from. An outcome study from years ago does not make outdated instruction current.

So look at the instruction itself. A strong program provides explicit, systematic instruction in foundational skills, uses decodable text appropriately for beginning readers, intentionally builds language and knowledge, and avoids encouraging students to rely on guessing strategies in place of decoding. The evidence tier tells you a study existed; the design tells you whether the program reflects how reading actually develops.

4. Was it studied with students and settings like ours?

A result found in one context does not automatically repeat in another. A program studied in affluent suburban schools may behave differently in a high-poverty urban district, and evidence from one grade band says little about another.

A strong answer shows the program was studied with students and in settings reasonably similar to yours, in grade level, demographics, and the kinds of challenges your schools face. A weak answer waves at a single study from a very different population. The closer the studied context is to your own, the more the evidence is worth to you.

5. What does implementing it well require, and can we provide it?

This is the question that decides whether any of the evidence will matter, and it is the one committees skip most often. Every program assumes conditions for its results: a certain amount of instructional time, specific materials, and, almost always, teacher training and ongoing support. The best-evidenced program in the world produces little if a school cannot implement it as designed.

A strong answer is honest about what the program requires and what support the vendor provides, including training and coaching over time rather than a single launch day. A weak answer implies the program works on its own, out of the box. Ask what implementation actually demands, then ask yourself honestly whether your schools can provide it. Evidence on paper and results in classrooms are connected only by implementation.

6. How will we know if it's working here?

An evidence base tells you a program has worked elsewhere. It does not promise it will work for you, and no responsible vendor should claim otherwise. The committee's last question is therefore about your own accountability: how will you know?

A strong plan defines this before purchase. It identifies what you will watch, including early signals like changes in teacher practice and, over time, student outcomes, and sets points to review whether the program is delivering in your context. A weak plan is to buy the program and assume the published results will follow. They might not, and you want to find out early enough to act.

Evidence is necessary, not sufficient

Running these six questions will tell you whether a program is genuinely evidence-based and worth a contract. It is worth keeping the larger truth in view: evidence is necessary but not sufficient. A program that clears all six still has to be taught well, by supported teachers, in a school whose schedule and leadership back the work. The strongest curriculum decisions pair high-quality instructional materials with the systems that allow teachers to implement them well. Choosing the program is only the beginning; sustained support is what turns potential into results.

Frequently asked questions

What does "evidence-based" mean for a literacy program?

Under ESSA, "evidence-based" has a specific meaning: the program is supported by a study that meets one of four evidence tiers, with the top three requiring a statistically significant, positive effect. In everyday vendor use, the phrase is often looser, which is why a committee should ask for the specific tier and study rather than accepting the label.

How do I know if a reading program is really evidence-based?

Ask which ESSA tier it meets and for which product, whether the evidence is independent, whether its design reflects the Science of Reading, whether it was studied with students like yours, what it takes to implement, and how you will measure results locally. A program that answers all six well is genuinely evidence-based; a label alone is not enough.

What questions should a curriculum committee ask a vendor?

At minimum: What evidence tier do you meet, for this exact product? Who conducted and funded the study? Does the program teach reading in line with the Science of Reading? Was it studied with students like ours? What does implementing it well require? And how will we measure whether it works in our schools?

Is "research-based" the same as "evidence-based"?

No. "Evidence-based" has a defined meaning under ESSA, tied to studies and tiers. "Research-based" has no such definition and often means only that a program is informed by research. For funding and for due diligence, a specific evidence tier matters more than the word "research-based."

Does an evidence-based program guarantee results?

No. Evidence shows a program has produced results elsewhere. It does not guarantee the same in your schools. Outcomes depend on implementation: teacher training, support, time, and leadership. That is why measuring results locally is one of the six questions, and why evidence is necessary but not sufficient.

Keep reading

These evidence-tier references follow the Every Student Succeeds Act's definition of "evidence-based" (ESEA as amended, § 8101(21)) and U.S. Department of Education guidance, including the role of the What Works Clearinghouse in judging study quality.

Don't miss new insights

Sign up for occasional updates and featured literacy resources.