A living map of the research on AI and learning

How to Use AI Without Losing Your Ability to Think

What research tells us about learning with AI while preserving intellectual autonomy.

This question has been on my mind for several months. I stepped back from the announcements and individual studies to try to see the bigger picture: what connects this research, and what distinguishes an AI that provides answers from one that genuinely helps us learn?

This work began with a question: how can we use AI without losing our ability to think for ourselves? The more I read, the more that question seemed too narrow. Learning to think for ourselves also means learning with others, while our use of AI takes place in a world of finite resources.

So the question gradually broadened, until it became the one that now closes this overview:

What do we want to remain capable of?

Two caveats. I am not an education researcher: what I offer here is a reading of the available evidence, intended to contribute to the discussion, not close it. This work also began as a snapshot in June 2026. Since then, models, tools, and research have continued to evolve. Some of these conclusions will therefore need to be revisited.

This overview will be updated regularly. Depending on their significance, new publications will either be added as notes to the relevant sections or lead to a new section. Every change will be recorded in “Changes over time”.

Sources are collected at the end of the article. You can also view the original carousel as a PDF.

Scroll to begin

The problem

AI can improve performance and reduce learning.

In a randomised study of about 1,000 high school maths students in Türkiye, students using a standard ChatGPT scored 48% higher than classmates without it while they had it. When access was taken away, they scored 17% lower than students who never had it. A 2025 paper in Nature Reviews Psychology makes the underlying point: performance gains are not the same as learning.

Source: Bastani et al. (2025), PNAS; Yan, Greiff, Lodge & Gašević (2025), Nature Reviews Psychology.

Last sharpened Jul 2026

Better output is not the same as learning.

Why it happens

When AI does the thinking, the learning does not happen.

Most AI tools were built for work, not to optimise learning. At work, the goal is to finish the task with the least effort. But in learning, that effort is the point: it is what builds the capability. The task gets finished, but the understanding never develops. Researchers call this “metacognitive laziness”: the learner stops planning, monitoring and self-evaluating, because the AI always has an answer.

Source: Khosravi et al. (2026); Fan et al. (2025), BJET.

If the AI carries the effort, it also carries off the learning.

Why effort is the point

The struggle is not an obstacle to learning. It is the learning.

Learning scientists Robert and Elizabeth Bjork call these “desirable difficulties”: things like recalling an answer from memory, spacing your practice out, or trying a problem before you see the solution. They make learning feel harder now, but they make it last. When AI removes that effort, it can quietly remove the learning the effort produced. A 2026 review of 67 studies names the same mechanism, epistemic friction: without it, “AI-generated fluency can bypass the reflective struggle central to deep learning.”

Source: Bjork & Bjork; Li, Cui & Hagedorn (2026), Computers and Education: AI.

Protect the difficulty that does the teaching.

It starts with the person

When intelligence is plentiful, volition is valuable.

If the struggle is the learning, the first thing it depends on is the person doing it. Writing in The Atlantic, David Brooks argues that what will set people apart in an AI-saturated world is not how smart they are but their relationship to mental effort. His frame is need for cognition, the trait named and measured by John Cacioppo and colleagues, who reviewed more than 100 studies of it: at one pole, people who read dense books and play hard games for pleasure; at the other, the cognitive miser, who avoids effortful thought where possible. It correlates with intelligence, Brooks notes, but is not the same as it.

The evidence he marshals is largely about what happens when effort is offloaded. An MIT Media Lab team led by Nataliya Kosmyna measured brain connectivity dropping by up to 55% during ChatGPT use; Michael Gerlich (SBS Swiss Business School) found a significant negative correlation between frequent AI use and critical thinking; a Carnegie Mellon team led by Grace Liu found that after about ten minutes of AI-assisted problem solving, people who then lost the tool did worse than those who never had it. A study of endoscopists found precancerous-lesion detection fell from 28.4% to 22.4% once AI was withdrawn. From this he sketches three responses: the Productive Passenger, who lets AI think and stops noticing; the Reluctant Optimizer, who means to resist and gets pulled in; and the Mental Marathoner, who keeps the hard parts hard. His worry is a cognitive polarization between the two ends. His practical rule: ask AI for thinkers, not thinking, and treat it as a brilliant librarian, not an oracle.

Source: David Brooks (2026), The People Who Will Thrive in the AI Age, The Atlantic; on need for cognition, Cacioppo & Petty (1982), Journal of Personality and Social Psychology.

Last sharpened Jul 2026

The first variable is the learner, not the tool.

With what's already in their head

When answers are plentiful, background knowledge is the bottleneck.

Volition is not the only person-side variable. Stephen Fitzpatrick, a history teacher of thirty years, argues that the skill gating all of this is reading: AI output is fluent, confident and endless, and reading it with the skepticism it demands takes the vocabulary and background knowledge that only years of reading build. “The bottleneck to effective student AI use is mostly a problem of reading”, not writing, which gets the attention because writing is what we grade. Deep research makes the point sharper: finding sources is no longer the hard part; reading the report, and its sources, carefully is.

Marie Dollé, writing in French about the summer’s “brainmaxxing” apps, takes the same point one level deeper. Against the comforting idea that machines can hold the knowledge while we keep the judgment, she argues that judgment cannot form in a vacuum: “on ne problématise pas sans repères”, you cannot frame a problem without reference points already in your head. And she names a limit no better model fixes: what matters is often only knowable later, so no tool can pre-sort “the essential” for us. Consulting a fact on demand is not the same as having it in mind.

Daniel Susskind, an economist who has spent fifteen years studying AI and work, gives the same claim a mechanism. AI systems are what Geoffrey Hinton calls “idiot savants”: impressive on hard problems, wrong on easy ones, and nothing in the output tells you which one you are reading. So the basics are not a safe harbour, they are an instrument. Use AI critically rather than blindly, he argues, “keeping our basics sharp”, precisely so you can tell the savant from the idiot. Which is why he wants literacy and numeracy taught intensely even where AI already does them better, and calls it a no-regrets investment: one that pays off whatever the future turns out to be. The backdrop is not reassuring. Since 2009, literacy and numeracy have been falling among young people worldwide, according to the OECD’s PISA programme, and among adults too.

Source: Stephen Fitzpatrick (2026), It's the Reading, Stupid, Fitzy's History; Marie Dollé (2026), La star de l'été…, mariedolle.substack.com; Daniel Susskind (2026), The Guardian.

Last sharpened Sep 2026

You can only check a machine against what you already hold.

And by what we weren't looking for

Background knowledge helps us judge. It cannot tell us what will matter tomorrow.

Part of what shapes us comes from encounters whose importance we could not have measured in advance: a book, an idea, a person, an unexpected experience.

Marie Dollé sees here a fundamental limit of any technology that would sort the world on our behalf: it can learn what matters to us today, but how could it know what will matter tomorrow, when that also depends on events that have not yet happened and on who we will have become?

Research already documents this risk: writing with a language model makes texts more alike (Doshi & Hauser, 2024), and an assistant can nudge our opinions (Jakesch et al., 2023), or simply confirm what we already thought (Sharma et al., 2023).

Learning, then, is not only choosing better from what we already know. It is also staying open to what can shift our bearings.

Source: Marie Dollé (2026), La star de l'été…, mariedolle.substack.com.

Last sharpened Sep 2026

Preserve our capacity to be surprised, to wander off course, to come out changed.

Then the design

The same technology can help or harm. The design decides which.

In a Harvard physics study, a purpose-built AI tutor designed with proper scaffolding beat in-person active learning by 0.73 to 1.3 standard deviations, two to three times the usual bar for a substantial effect in education research. In the Türkiye study, the standard ChatGPT left students worse off once it was removed, while a guardrailed tutor version avoided that loss entirely. Same models, opposite outcomes, depending on how they were designed.

Source: Kestin et al. (2025), Scientific Reports; Bastani et al. (2025), PNAS.

The result is set by the design, not by the model.

And by the human behind it

Believing a human is paying attention changes how hard we try.

In a controlled study in a university creative-coding course, students received identical AI-generated feedback on their work. Those told it came from a human teaching assistant ran their code more, wrote more code, and spent more time on later work. They rated the feedback equally helpful either way. The effect on effort was large (d = 0.88 to 1.56). The content was the same.

Source: Morris & Maes (2026), Same Feedback, Different Source.

Same words land differently when we believe a human wrote them.

What only a human does

AI can help with the content. It rarely touches the rest.

Education does three things at once: it builds knowledge and skills (qualification), it helps you find your place among others (socialization), and it helps you become someone who thinks independently and takes responsibility (subjectification). AI tools mostly reach the first. They rarely, if ever, address the other two, and those are where a teacher does their deepest work.

Source: Gert Biesta; Wayne Holmes (2026).

AI can teach the content. A human helps you become someone.

Three ways AI can show up

An LLM, a tutor, and a learning companion are not the same thing.

An LLM answers your question. Faster work, less learning.

An AI tutor asks questions back, no matter what you actually need. Often frustration and drop-out.

An AI learning companion (Dr Philippa Hardman calls it a “study mate”) remembers where you got stuck and pushes you towards the thinking you avoid. Capability that lasts.

Source: Khosravi et al. (2026); Dr Philippa Hardman.

Aim for a learning companion, not an answer machine.

What education is for now

AI does not shorten what there is to learn. It lengthens it.

Holmes (UNESCO) asks it directly: if generative AI is this powerful, do we still need to learn? His answer is yes, and the list grows: on top of what we wish to learn, we now need to learn AI’s profound limitations, its impacts on human rights, social justice and the environment, and “perhaps most importantly, [to] learn how to think… critically.” The World Economic Forum (WEF) keeps the balance: rote memorisation may matter less, but “the process of mastering knowledge continues to develop broader capabilities” such as grit, curiosity, communication and critical thinking, and assessment must evolve to capture them.

Source: Wayne Holmes (2026), UNESCO Courier; WEF (2026).

The emphasis moves from having answers to judging them.

In a world of finite resources

Every answer has an energy cost. The task decides how large.

Learning with AI also happens somewhere physical: on servers, drawing electricity. A team from Hugging Face and Carnegie Mellon (Luccioni, Jernite & Strubell) measured it task by task: 88 models, a thousand requests per dataset, each run ten times. Sorting a text into categories used about 0.002 kWh per thousand requests. Generating text used over ten times more; generating images over sixty times more again. The least efficient image model used around half a smartphone charge for every picture.

The larger finding is about generality. A model built for one well-defined task used far less energy than a general-purpose model doing the same job: in the authors’ tests, from a few times less to about thirty times less. Their overall verdict is stronger still: “multi-purpose, generative architectures are orders of magnitude more expensive than task-specific systems for a variety of tasks, even when controlling for the number of model parameters.” For well-defined tasks such as web search, they see no convincing evidence that general models are needed, given how much energy they use.

So discernment has a physical side. Choosing the right tool for the task, a search, a classifier, a smaller model, or no AI at all, is also a way of taking the environment into account.

Source: Luccioni, Jernite & Strubell (2024), Power Hungry Processing, ACM FAccT.

Last sharpened Sep 2026

Use the smallest tool that does the job.

Ourselves, others, and the planet

Three levels of stakes. At each level, a relationship to attend to.

This map keeps meeting three stakes: what AI does to our thinking, to how we live and learn with others, and to the world it draws on. They are usually studied apart. Learning scientists measure thinking, ethicists debate fairness and power, engineers count energy. The declarations that try to hold them together, from the IEEE’s design guidelines to the Montreal and Toronto declarations, tend to do it around a single aim: human well-being.

In 2019 a group of Indigenous scholars, artists and technologists from Hawaiʻi, Aotearoa, Australia and North America answered those declarations with a position paper of their own, published with the Canadian Institute for Advanced Research and edited by Jason Edward Lewis (Cherokee, Hawaiian and Samoan), who holds a research chair in computational media at Concordia University in Montreal. Their objection is precise: “none of these efforts challenge the fundamental anthropocentrism of Western science and technology.” Their alternative is relational. Instead of asking only how a technology serves us, they ask what relationships it creates, and what each side owes. In their words, “while the developers might assume they are building a product or tool, they are actually building a relationship to which they should attend.”

Read this way, the three stakes become three relationships. With what we know: the historian Noelani Arista (Kanaka Maoli) asks whether “computer memory” will “replace experts and elders as repositories of knowledge”, and calls for institutions that train knowledge keepers “fluent in language and trained in computer science.” With others: tools should be built “collaboratively with, not on behalf of, and certainly not without the community”, because “The alternative is to have our worlds designed for us.” With the planet: the artist Suzanne Kite (Oglála Lakȟóta) asks “What is being offered to the Earth when we extract these mined materials?”, the question the previous section began to answer in kilowatt hours.

This is a lens, not evidence, and it asks for care. The contributors speak from particular nations and places, and say so: “This essay does not attempt to speak for all Lakota”. The paper does not forbid others from learning from it, but its participants “expressed concern about issues of appropriation and misuse of traditional knowledge”, so it is credited here as a way of seeing the question, not borrowed as a method. It was written in January 2020, before ChatGPT, about systems communities build and govern themselves, not chatbots owned by a few companies. And as a position paper, it opens a question rather than settling one.

Source: Lewis (ed.) (2020), Indigenous Protocol and Artificial Intelligence Position Paper, Initiative for Indigenous Futures and CIFAR.

Last sharpened Sep 2026

Ask what each relationship owes, and to whom.

So how do we make it help?

The challenge is also systemic. It runs through the tool, the classroom, and the system around them.

The tool: the guardrailed tutor and the learning companion are instructional design built into software: scaffolding, answers withheld, help that fades as you grow.

The classroom: co-design tools with teachers rather than deploying them on teachers, and set tasks that make learners compare, justify and revise what the AI produces. In the 67-study review, that scaffolding is what separated gains from cognitive offloading.

The system: AI only helps where the conditions are ready. As the WEF puts it, “learning outcomes will not be determined by technology itself, but by the conditions in which it is deployed”, and isolated fixes across policy, pedagogy and technology are unlikely to be sufficient.

In August 2026, the system layer got its first institution-scale test case. MIT’s committee on AI in teaching and learning refused patching (“this is not a moment for patches and duct tape”) and called for all three moves in concert: AI-aware redesign of every subject, an in-person social component in every subject, and permanent adaptation structures (a standing committee, per-school AI leads, a pilot fund), on a platform deliberately tied to no single AI vendor. One survey number shows what is at stake: MIT undergraduates report feeling more replaceable than capable (40% versus 34%).

Susskind adds the one move here with a precedent. In 1982 a British government report by Wilfred Cockcroft answered the electronic calculator not by banning it but by splitting mathematics in two: part of the time learning to work with a calculator, the rest learning to cope without one, and, the decisive part, examining both. That split is now the standard almost everywhere. He proposes the same structure for AI, in every subject, and calls it teach both, test both. His reason for putting the weight on assessment rather than detection is the part worth keeping: a teacher cannot know whether a student used AI alone in their bedroom, but nobody forgets sitting an exam they have only half prepared for.

Source: Bastani; Khosravi; OECD (2026); Li/Cui/Hagedorn; WEF; MIT AI committee (2026); Susskind (2026).

Last sharpened Sep 2026

Real progress needs all three moving.

What to take away
  • Preserve the difficulties that help us grow.
  • Make learning design a priority in the development of AI tools.
  • Recognize that human relationships shape us far beyond the transmission of knowledge.
  • Choose tools that strengthen our capabilities, rather than simply producing answers for us.
  • Cultivate a taste for intellectual effort.
  • Also seek, with and without AI, what challenges or shifts our ideas.

Wayne Holmes sums it up: learn AI’s limitations, “take into account its broader impacts on human rights, social justice and the environment, and, perhaps most importantly, learn how to think… critically”. In other words: a cognitive challenge, a societal one, an environmental one.

Three reasons to aim for what MIT describes as the highest-quality education: “of humans, by humans, in support of human flourishing, and for the betterment of humankind.”

And to ask:

What do we want to remain capable of?

Where this comes from

A map of the research, drawn over time and kept current as new sources arrive.

Download the original carousel (PDF)