Educational Resources

Stanford AI Tutor RCT Result: Access Isn't Intervention

Written by Danielle Brodetsky | Jul 22, 2026 5:44:48 PM

When school districts carved out dedicated instructional time for elementary students to work with a well-known AI reading tutor, the students barely used it. That is the finding from a new randomized controlled trial run by Stanford's SCALE Initiative, and it is the first Tier-1 research result to publicly puncture the pitch that AI tutoring is the answer to district budget pressure. "Having access to this AI tutor isn't the same as using it," Stanford SCALE research director Carly Robinson told Chalkbeat. For charter directors and special programs coordinators being courted by a wave of AI-native tutoring vendors, that sentence should stop the meeting.

What The Stanford RCT Actually Found

The Stanford SCALE Initiative ran a randomized controlled trial with elementary students in U.S. school districts. Some students were assigned to work independently with an AI literacy tutor. Others worked with an in-person tutor whose primary job was to keep students engaged with the same platform. Districts set aside protected time. Students had licensed access. The technology worked.

Students still did not use it much.

The through-line of the finding, reported by Chalkbeat and echoed in Stanford SCALE's own summary, is that neither providing access nor pairing the AI tool with a light-touch adult was enough to generate the usage patterns the platform's own model required to produce learning gains. According to Chalkbeat's coverage, the human-tutor condition moved usage by only one to four additional minutes per week, and many students never logged on at all. The RCT is one of the first independent studies of a widely-marketed AI tutoring product deployed at classroom scale, and it lands in a market where districts have been told repeatedly that AI can substitute for the labor of tutoring.

Why This Matters To Charter Directors And Program Coordinators

If you are a charter director, a special programs coordinator, or a federal programs lead, you are almost certainly being pitched an AI tutoring product right now, probably several. The pitch is consistent across the category: unlimited student access, adaptive content, on-demand availability, and a price point that looks attractive against per-hour human tutoring.

The Stanford finding matters because it directly attacks the load-bearing assumption underneath every one of those pitches. The assumption is that if you give students access to a good tool during protected time, students will use it enough to benefit. The RCT says no.

That is not a small footnote. That is the entire ROI calculation.

When was the last time you asked a vendor for their per-student weekly minutes-of-use data, broken out by subgroup, over a full semester?

If you have a Title I or Title III allocation being deployed against an intervention line item, ESSA's evidence tiers require you to fund what actually works with your specific students. A product that shows a compelling efficacy claim in a controlled pilot but generates near-zero engagement in your building does not clear that bar.

What The Research Actually Says About Tutoring That Works

The Stanford SCALE finding does not stand alone. It sits inside a much larger body of research that has, for a decade, kept saying the same thing.

A landmark meta-analysis of 96 randomized tutoring studies by Nickow, Oreopoulos, and Quan (NBER Working Paper 27476, 2020) found consistent, large positive impacts of tutoring on math and reading across grade levels, with effect sizes largest for programs that use teachers or paraprofessionals as tutors, occur at least three days a week, and are held during the school day. That evidence base is foundational to the ongoing synthesis work of the National Student Support Accelerator at Stanford, led by Susanna Loeb, which has translated the research into a set of design conditions that must be met for high-impact tutoring to produce those gains: three or more sessions per week, small groups (typically 1:1 to 1:4), tutor consistency across sessions, alignment to classroom instruction, and sustained duration across a full intervention window.

The NSSA Design Principles name engagement not as a byproduct but as a design requirement. Consistency of the tutor is called out explicitly. Frequency is called out explicitly. Relational trust is called out explicitly.

In our experience across two decades of running Tier 2 and Tier 3 interventions, the tutor-student relationship is not incidental to the intervention. It is the mechanism through which the intervention works. An AI tool with a login screen and an assignment window is not a mechanism. It is a resource. Those are different things, and the Stanford RCT is the empirical demonstration of the difference.

What A+ Sees In The Field

A+ Tutoring, a California K-12 virtual intervention provider working with charter LEAs and homeschool charter families, has run engagement-designed tutoring for two decades. The model is humans-first by design: consistent tutor assignment across the intervention window, small-group and 1:1 session structures, cadenced session frequency, and completion tracking that we report weekly to our partner schools.

The engagement design shows up in the outcomes. In A+'s 2024-25 iLEAD Math Tier 3 cohort, 75% of students (9 of 12) reached growth benchmarks on NWEA MAP. In A+'s iLEAD ELA Tier 3 cohort, 87.5% of students (7 of 8) reached growth benchmarks. Combined across both subjects, 80% of the Tier 3 cohort (16 of 20) posted MAP Growth at 3-6x national benchmarks.

Those numbers are not primarily a story about instructional method. They are a story about attendance and completion. Students who show up consistently to a small group led by the same tutor, week after week, learn. Students who log in twice and disappear do not.

If your current intervention vendor reported attendance and completion the way A+ does, would the numbers on the page match what the pitch deck promised?

We are not making this argument because it is convenient. We are making it because it is the only argument the Stanford RCT allows. Access is not intervention. Attendance is.

What School Leaders Can Do Next

Whether or not A+ ever ends up on your vendor list, the Stanford finding gives you a concrete set of questions to bring to every intervention conversation this year.

  1. Ask any AI-primary vendor for per-student weekly usage data from a comparable district deployment, disaggregated by subgroup, across a full semester, not a pilot cohort.
  2. Audit your current Tier 2 and Tier 3 caseload for tutor consistency. How many of your students had the same tutor for the full intervention window last year?
  3. Review your MAP growth data by intervention type. Separate students who completed 80%+ of scheduled sessions from those who did not, and look at the growth curves side by side.
  4. Pull your ESSA evidence tier documentation for every intervention product on your budget. Confirm the evidence base was generated under deployment conditions comparable to yours.
  5. If you sit on the federal programs side of the budget, pull your Title I or Title III per-vendor cost-per-completed-session-hour for the last academic year, not cost-per-license. That single number reframes every AI vendor conversation on your desk.

About A+ Tutoring

A+ Tutoring is a California K-12 virtual intervention provider partnering with charter LEAs, homeschool charter families, and district Tier 2 and Tier 3 programs. Our model is engagement-designed from the ground up: consistent tutors, small groups, cadenced frequency, weekly attendance and completion reporting. In 2024-25, A+ partner schools showed 75% of Math Tier 3 students reaching growth benchmarks, 87.5% in ELA Tier 3, and 80% in the combined Tier 3 cohort, at 3-6x national MAP Growth benchmarks.

Walk Your Intervention Vendor List With Danielle