Teachers and school leaders were right to be careful, and the careful questions have changed. The early worries were about a tool that made things up and let students skip the work. Those have not gone away, but in 2026 the harder questions are about tools that arrive already installed, claims of learning gains that nobody has checked, and decisions made above the classroom.
This page takes the questions as they are actually being asked now. We think skepticism is the right starting point, and the useful kind is specific.
The failure has not disappeared; it has become harder to spot. Early models were wrong in obvious ways. Current ones are wrong less often and more plausibly, which is a worse combination for a student who does not yet know the material. Fabricated citations are the clearest case: a reference with a real-looking author, a real-sounding journal, and a volume number that does not exist.
The practical response is to stop treating verification as an occasional check and make it part of the assignment. Ask students to trace a claim back to a source they can open. When the source does not exist, that is the lesson, and it is a better one than any warning you could give in advance.
Detection is not the answer, and it is worth being direct about why. Detectors are unreliable in both directions, and they are disproportionately likely to flag writing by students who learned English as an additional language. A policy that rests on detection will produce false accusations, and it will produce them unevenly.
What works better is assignment design that makes the process visible: drafts, oral defense, in-class work, revision history, and reflection on where a tool helped and where it did not. This is more work to set up once and less work to police forever. It also gives you something to say to a student that is about learning rather than about catching them.
This is the question we spend most of our time on, and it is the one most likely to be answered badly.
An efficacy claim is at least three claims wearing one number: what was measured, who judged it, and whether it holds for your students rather than the ones in the study. A tool that improved scores on a vendor-designed post-test, graded by the vendor's own rubric, in schools that volunteered, has told you something much narrower than "it improves learning."
Ask what the comparison group did. Ask who chose which students got the tool. Ask what the effect was for the students who were furthest behind, because an average can hide a widening gap. And ask what would have to be true for the claim to be wrong, which is a question good evidence survives and marketing usually does not.
Procurement made without the people who will use it is the most common way a good tool becomes a resented one. It is also how a tool ends up deployed in a setting its evidence never covered.
Educators have standing to ask, before rollout, what the tool was tested on, what data it collects, who reviews its outputs, and what the exit path is if it does not work. If the answer to the last one is that there is no exit path, that is worth surfacing while it is still a purchasing decision rather than a sunk cost.
Models learn from text that carries the assumptions of the people who wrote it, and the outputs carry them forward: whose history gets centered, whose names get mangled, which dialects get treated as errors.
The part worth adding in 2026 is that bias is measurable, and that "we take bias seriously" is not a measurement. If a tool is being used to sort, flag, rank, or recommend students in any way, the question to ask is whether performance was reported separately for the groups you serve, or only in aggregate. An aggregate number can look fine while the tool works poorly for exactly the students who most need it to work.
This has become more consequential, not less, because tools now reach further into school systems and increasingly take actions rather than just producing text.
The questions worth asking of any tool touching student information: is our data used to train or improve the model, and can we turn that off; who at the vendor can see it; how long is it kept; what happens to it if we leave; and does this arrangement satisfy FERPA and, for younger students, COPPA. Get the answers in the contract rather than from the marketing page. A tool that connects to your student information system or takes actions on a student's behalf deserves a harder look than one that drafts text in a window.
This concern has aged well, and we think it is the right one to hold on to.
Teaching is not information delivery, and the parts of it that matter most are the parts that cannot be automated without being destroyed: knowing which student is having a bad week, and why the same mistake means something different coming from two different children. Tools that reduce administrative load can protect that work. Tools that stand in for it, whether an AI companion offered as a friend or an automated message where a conversation belonged, spend down something you cannot easily rebuild.
Our stance is augment, do not replace. It decides real design questions for us, and it is why we build human review into workflows rather than adding it afterward. The most useful question a school can ask about any AI tool is not whether it saves time, but what it does with the time it saves.
We are a research nonprofit rather than a vendor. If you are weighing a claim about an education tool, or want help designing a check you can defend to a school board, get in touch.
AI for Altruism (A4A) is a 501(c)(3) nonprofit organization.
Centennial, Colorado, USA
Copyright © 2026 AI for Altruism, Inc. - All Rights Reserved.
Please donate today.
We use cookies to analyze website traffic and optimize your website experience. By accepting our use of cookies, your data will be aggregated with all other user data.