A new AI tool lands in your inbox. The demo is polished. The claims sound familiar: less teacher workload, more personalization, better engagement, faster feedback.

The easiest question is, “What can it do?”

That is also the wrong first question.

On August 20, the U.S. Department of Education released new guidance urging states and districts to evaluate education technology by its instructional value, evidence, implementation conditions, transparency, and effect on student outcomes. The guidance offers five plain questions for every product. They are especially useful for AI because AI products can make impressive outputs long before anyone has shown that students are learning more.

This is not a procurement checklist to hand off to the technology office. It is a leadership discipline.

Because the real decision is not whether AI belongs in school. The real decision is what work we want it to support, what evidence we will accept, and what judgment must remain human.

The Five Questions

  1. What learning problem does it solve?

    Start with the learning problem, not the product category. “We need an AI tutor” is not a learning problem. “Grade 7 students need more timely feedback while explaining proportional reasoning” is.

    The sharper the problem, the easier it becomes to reject features that look exciting but do not serve the goal. A tool can save time without improving learning. Productivity is not automatically progress.

  2. When should it be used?

    “Available all the time” is not an instructional strategy. Leaders and teachers need to name the moments when AI adds value and the moments when it weakens the work.

    AI may help a student generate counterarguments before a seminar. It should not replace the seminar. It may help a teacher sort exit-ticket patterns. It should not make the consequential decision about which student is capable of advanced work.

    Use AI where it scaffolds thinking, expands access, or removes low-value busywork. Protect the moments where struggle, explanation, judgment, and relationship are the learning.

  3. For whom should it be used?

    A districtwide license can make us pretend that one tool serves everyone equally. It rarely does.

    Ask who benefits, who may be excluded, and who carries additional risk. Does the tool work for multilingual learners? Is it accessible to students with disabilities? Does it require connectivity or devices that some families do not have? Are the reading level, language options, and privacy terms appropriate for the students who will actually use it?

    “For whom?” also means identifying the adults who need support. A tool that works in a vendor demonstration may fail in a classroom if teachers do not have time, examples, coaching, and permission to adapt it.

  4. For how long should it be used?

    Time matters in two ways. First, how much of a lesson or school day should this tool occupy? Second, how long will the district pilot it before deciding whether to expand, revise, or stop?

    Set both limits before launch. A six-week pilot with a clear review date is a decision. An indefinite pilot is drift.

    The new federal guidance also asks schools to distinguish purposeful instructional technology from recreational screen use. That distinction matters, but it should not become an excuse to ignore screen exposure. Even a useful tool can be overused. The goal is not maximum utilization. The goal is the right use for the right amount of time.

  5. What evidence demonstrates that it improves student learning?

    This is the question most likely to slow the room down. Good. It should.

    A March 2026 review from Stanford’s SCALE Initiative examined more than 800 papers relevant to AI in K–12 education and found only 20 that produced strong causal evidence. The review found promising short-term gains while students had access to AI, but mixed evidence about whether those gains transferred when the tool was removed. The authors also found that pedagogical guardrails mattered: tools designed to support step-by-step reasoning were more promising than general-purpose systems that simply produced answers.

    That does not mean schools should wait for perfect research. It means leaders should stop treating vendor claims, usage counts, and teacher enthusiasm as proof of learning.

    Define what you expect to change. Collect baseline evidence. Listen to teachers and students. Review work samples. Look for different effects across student groups. Decide in advance what result would cause you to continue, change course, or stop.

What the New Research Says About Implementation

A national landscape study released this month by the National Girls Collaborative Project adds an important implementation warning. The mixed-methods study included 552 education professionals and eight focus groups with 47 educators. Sixty-seven percent reported using AI at least several times a week in their professional lives, yet 37 percent had received no AI-related professional development and only 11 percent had completed more than eight hours.

Educators were not asking for another tour of AI vocabulary. They wanted classroom examples, guidance on student use and policy, and advanced implementation strategies. The researchers found a strong association between more professional learning and greater confidence, more frequent use, and stronger readiness to teach AI concepts. Because this was not a causal experiment, it does not prove that training alone produced those outcomes. It does show leaders where the implementation gap is.

The pattern is hard to ignore: tools are moving faster than the systems around them.

Recent reporting from Charleston County School District shows what a more grounded response can look like. The Associated Press reported that the district began with policy development and input from teachers, students, and parents, then moved into training. Teachers used flawed AI outputs as material for analysis, not as a reason to either ban the technology or trust it.

That is AI literacy in practice: not merely knowing how to prompt, but knowing when to question, verify, refuse, and take responsibility.

A 30-Day Decision Routine for School Leaders

If your team is considering an AI product this fall, use the five questions as a short decision cycle.

  1. Name one learning problem and one student outcome. Keep the statement specific enough that a teacher could recognize it in student work.
  2. Identify the protected human work. Write down what the tool may support and what it may not decide, replace, or obscure.
  3. Choose a bounded pilot. Define the students, educators, setting, duration, privacy requirements, and professional learning support.
  4. Collect evidence beyond usage: work samples, teacher observations, student feedback, subgroup patterns, and one outcome measure.
  5. Hold a stop-or-scale review. Continue only if the evidence supports the learning goal and the implementation conditions are realistic.

For classroom teachers, the same routine can fit on a lesson plan: What is the learning goal? Where does AI enter? Where must students think without it? What evidence will show that the student—not the system—understands?

The Question Beneath the Questions

The Department’s five questions are useful because they move the conversation away from novelty and back toward purpose. But schools should add one more question:

What must remain human here?

The answer will be different across tasks. AI may draft, organize, suggest, translate, or surface patterns. It cannot carry responsibility for a child, repair trust with a family, understand every piece of local context, or own the consequences of a decision.

Schools do not need to prove they are using more AI this year. They need to prove they are making better decisions about learning.

Five questions can help.

Human judgment still has to answer them.