The tools and methods for using AI in the academic workplace have improved dramatically over the past year or so. What has not improved at the same rate is the capacity to use them with any methodological seriousness. We now inhabit a landscape in which systems can quietly assemble collections of articles, interrogate those collections, automate routine tasks, and synthesise large volumes of information into serviceable arguments or teaching materials. Yet the distance between those who can work fluently with these tools and those who cannot is widening into something more structural: an AI capability gap.
This gap is not simply a matter of technical skill. It includes the ability to recognise the contexts in which AI is genuinely useful, and indeed ethical, and to understand which tool and which method belong to a particular task. This sounds straightforward, but it is precisely where much of the current commentary fails. We have a growing stack of articles measuring the effectiveness of AI in teaching or research, often authored by highly competent scholars attached to leading centres. The difficulty is that there is very little transparency in many articles and commentary about which systems are being used or which AI methods are being tested in these investigations, beyond an elusive gesture towards “prompting ChatGPT” (and this may also be a product of slow publishing decisions; what was true two years ago may not be true today).
Predictably, many of these articles begin with definitive statements: “AI is not capable of…”, “LLMs aren’t good for marking”, “LLMs hallucinate references”, “LLMs use my data for training”. Some of these claims are not incorrect in their narrow context, but they are rarely qualified with respect to the model, configuration, or method being deployed. If your study consists of a few generic prompts in the free version of ChatGPT, addressed to no corpus, under no constraints, then it is hardly surprising that the results do not meet the standards of scholarly assessment. The problem lies less with “AI” and more with an impoverished sense of method.
Part of this is historical. When OpenAI released ChatGPT 3.5 on 30 November 2022, it arrived in universities as a kind of technological thunderclap. The underlying research traditions were decades old, but the public reception was often shaped by Hollywood narratives and newspaper headlines rather than by any careful understanding of language models. The ethics of releasing a system of this scale to an unprepared public in this way are, at best, questionable. The early use cases- cheating, plagiarism, shortcuts- were framed in ways that suited outrage more than analysis, and they have proved extraordinarily sticky.
Hence the familiar refrains: “AI is only good for cheating.” “AI stole my data.” “AI can’t do maths.” “AI hallucinates references.” “AI is biased.” My short answer to many of these statements is that the system is being used badly. If one were to evaluate a statistical technique by throwing random numbers at it and then complaining about the outcome, we would rightly question the study’s design. Yet this is very close to what passes for evaluation of “AI” in much of the current literature: an undifferentiated category, a consumer interface, and a conclusion extrapolated well beyond the specific workflow.

The irony is that quite simple AI methods will neatly alleviate most of the problems that are rehearsed in these critiques. Creating a modest GPT or equivalent, uploading or linking a small set of key articles, turning off web access, and writing a few clear rules in a system prompt will address hallucinated references, reduce irrelevant content, and confine the model to a defined domain. This takes about twenty minutes, considerably less time than it takes to design and execute a research project aimed at proving that “AI is no good for marking student essays”. The lesser investment is in method; the greater investment is currently in complaint.
Good researchers already know not to use Google as their primary research method or to rely on Wikipedia as an authoritative source. We teach students to treat these tools as starting points, not destinations. The same should be said of language models. There are AI workflows that will yield results which can be integrated into serious research and teaching practices. There are also workflows that will lead the user into the digital wilderness, confirming whatever predispositions they had when they opened the browser. Both are technically “AI”. Only one deserves to be taken seriously.
Two implications follow. First, institutions need to provide better training for staff, not in AI as an abstract topic, but in the concrete use of language models as methods supported by appropriate tools. This includes understanding corpus design, data governance, model versioning, local policy constraints, and the relationship between automated outputs and professional judgement. It also requires discipline-specific methods and judgement rather than generic demonstrations of “what AI can do”.
Second, some unlearning is required. The early narratives about AI in universities, its presumed role as a cheating instrument, a data thief, a second-rate mathematician, have deeply shaped expectations. They underpin the reflexive response, “We tried AI and it didn’t work”, usually meaning “We typed a question into a chatbot and didn’t like the answer”. To make space for AI methods that can generate new knowledge, we need to retire the idea that the technology is either a shortcut or a threat, and instead treat it as either a small or a large part of the methodological toolkit, depending on the context.
Those who know how to situate AI within a meaningful methodology will quietly reorganise their work with AI as a partner. Those who do not will keep pouring bad wine into new bottles and then complain about the hangover. The probability that AI and LLMs will advance is high; the question is whether our methods will do likewise, or whether we will keep complaining about why the bottle itself is the problem.

Leave a Reply