AI is becoming ordinary. The evidence about its effects is still catching up
The most important question about artificial intelligence may soon stop being what the systems can do. It may become what repeated use does to people.
That shift is already visible in education. Generative AI has moved from novelty to ordinary infrastructure remarkably quickly. Students can use it to explain a difficult concept, critique an essay, generate practice questions or simply produce an answer. Those uses are not equivalent, yet public debate often compresses them into a single question: is AI good or bad for learning?
Recent developments suggest a more useful phase of research is beginning.
On 8 September 2026, OpenAI announced a $5 million programme for independent research on how generative AI affects people aged 13 to 17. The proposed research areas include emotional development, relationships, patterns of use, safeguards, AI literacy and differences across cultural and socioeconomic settings. The programme explicitly invites experimental, observational, qualitative and mixed-method work.
The announcement is notable not because company-funded research can settle the question. It cannot. OpenAI has an obvious institutional interest in how evidence about its technology develops, which makes disclosure, independence, publication and replication particularly important. The programme itself acknowledges this by requiring an independence and conflicts disclosure, including independence and credibility among its review criteria, and strongly encouraging researchers to make their findings publicly available.
What matters is the research question it reflects. “AI use” is too broad an exposure to be scientifically satisfying.
Imagine two students who each spend thirty minutes with the same model. One asks it to solve a problem and copies the answer. The other attempts the problem first, asks the model to identify weaknesses, challenges its explanation and revises the work. Recording both as thirty minutes of AI use would hide the mechanism that actually interests us.
The same difficulty appears in emerging evidence on educational outcomes. A 2026 meta-analysis in Humanities and Social Sciences Communications reported generally positive pooled effects of generative-AI-supported approaches on outcomes including academic achievement, higher-order thinking and writing. That is useful evidence, but meta-analysis does not make heterogeneity disappear. Effects can depend on the intervention, subject, learner, comparison group, study quality and the way the AI is incorporated into teaching.
There is also a measurement problem created by the speed of the technology itself. A study designed around one generation of models can be published into a world of more capable systems. Interfaces, safeguards and user behaviour change too. Evidence can therefore age unusually quickly even when the underlying study was rigorous.
For universities and schools, this argues against two easy responses. One is prohibition based on the assumption that every use substitutes for thinking. The other is adoption based on the assumption that access automatically improves learning. Both jump ahead of the evidence.
A better approach is to specify the outcome and the mechanism. Are we interested in exam performance, retention six months later, writing quality, confidence, critical reasoning, time saved, social interaction or dependence on external assistance? Does AI provide feedback after an attempt, or replace the attempt? Are gains concentrated among students who already know how to evaluate an answer? Do effects persist when the tool is removed?
Those are less dramatic questions than asking whether AI will transform education. They are also more answerable.
The broader lesson extends beyond classrooms. As AI becomes embedded in research, public services and professional work, capability benchmarks will tell us only part of what we need to know. A system can perform impressively in isolation while producing ambiguous effects once humans reorganise their behaviour around it.
The next frontier of AI research is therefore partly social science: not merely measuring the machine, but measuring the human-machine system that forms around it.