100+methodology articlesWritersScribe
Systematic Reviews

Conducting a Systematic Review on ChatGPT in Higher Education

July 3, 2026·Dr. Priya Nair·5 min read
On this page

Few education technology topics have produced as much primary research as quickly as ChatGPT's role in higher education, and this creates a genuinely distinct methodological situation: a systematic review here is reviewing a body of literature that may be substantially larger, and substantially newer, than what a reviewer researching a more established educational technology topic would encounter.

Narrowing "ChatGPT in higher education" to an answerable question

This broad topic spans student use for coursework, instructor use for content creation, institutional policy responses, academic integrity concerns, and learning outcome measurement, each representing a genuinely distinct research question. A defensible review specifies which of these dimensions it addresses, which student or institutional population, and which specific outcome -- learning outcomes, academic integrity violations, instructor workload, or another clearly defined measure -- rather than attempting a single review spanning all of these genuinely different questions at once.

Study design diversity in this literature

Primary research on this topic includes randomized classroom interventions, but also a substantial volume of survey-based studies, qualitative interview studies, and descriptive case studies of institutional implementation, reflecting how rapidly and informally many institutions began engaging with this technology. Your eligibility criteria need to specify explicitly which of these study designs are eligible, and your appraisal approach needs tools matched to whichever mix you ultimately include -- likely requiring both quantitative and qualitative appraisal tools within the same review given this field's genuinely mixed evidence base.

Search strategy and this literature's specific publication pattern

A significant portion of early ChatGPT-in-education research appeared first as preprints or conference proceedings before, or instead of, formal peer-reviewed journal publication, given how quickly institutions and researchers needed to respond to a rapidly changing situation. Searching preprint servers and relevant education technology conference proceedings, alongside standard databases like ERIC and Web of Science, is particularly important for this specific topic given how much relevant early research exists outside conventional peer-reviewed journals.

Handling a genuinely fast-evolving evidence base

ChatGPT itself, and the broader landscape of similar tools, has changed substantially even within the relatively short time since this research area emerged, meaning studies from different points within even the past two years may be examining meaningfully different underlying technology. Explicitly noting the studied tool's version or approximate date within your extraction, where reported, and considering whether this affects how comparable your included studies genuinely are, is a specific methodological consideration this fast-moving topic requires that a more established research area would not.

Considering a living systematic review approach

Given how quickly new primary research continues to emerge on this specific topic, a living systematic review, discussed in more general terms elsewhere on this site, is a particularly strong candidate methodology here, letting your synthesis stay genuinely current as the literature continues to grow rather than becoming outdated within months of publication.

Academic integrity as a specific outcome requiring careful definition

If your review addresses academic integrity concerns specifically, defining exactly what counts as a relevant violation or concern across your included studies matters considerably, since this literature uses inconsistent terminology and inconsistent severity thresholds when discussing this outcome, and a loosely defined extraction approach here risks combining genuinely different concerns under a single heading.

Risk of bias considerations specific to this literature

Many included studies in this space are self-reported survey research or small-scale classroom studies without a comparison group, meaning standard RCT-focused risk-of-bias tools often do not apply cleanly. Appropriate appraisal here more often means tools suited to survey research and cross-sectional designs, or a qualitative appraisal tool where relevant, rather than defaulting to RoB 2 regardless of fit.

Why methodology support helps particularly here

The sheer volume and pace of this specific literature, combined with its genuinely mixed study-design landscape, makes this a topic area where a clear, well-planned protocol matters more than usual -- a vague question risks an unmanageable, inconsistent evidence base, while a properly narrowed one produces a genuinely useful, current synthesis of a topic institutions and researchers are actively trying to understand in real time.

Handling conflicting findings across a fast-moving evidence base

Given how quickly this specific literature has grown, it is common to encounter genuinely conflicting findings across studies conducted only months apart, sometimes reflecting real differences in student populations or institutional contexts, and sometimes reflecting how rapidly the underlying technology itself changed between studies. Discussing this kind of conflict explicitly, rather than glossing over it in favor of a tidier-sounding overall conclusion, is a more honest and more useful contribution given this field's genuine volatility.

A note on institutional policy as context, not evidence

Many published sources on this topic describe institutional policy responses rather than reporting genuine empirical findings, and distinguishing policy description from empirical evidence clearly in your synthesis, rather than treating a described policy change as though it were itself a research finding, keeps your review's evidentiary claims appropriately grounded.

A closing consideration on this topic's audience

Instructors, administrators, and policymakers are all actively seeking trustworthy synthesis on this exact topic right now, making a rigorously conducted review here genuinely more likely to reach and inform real decision-making than reviews on more slowly evolving, less publicly visible topics. Treating this responsibility seriously, with the same rigor expected of any other systematic review, is what makes the resulting synthesis genuinely trustworthy rather than merely timely.

A final word on staying current with this topic

Because this specific area continues to change so quickly, treating your first published review as a starting point rather than a final word, and building a realistic plan for when and how you might revisit it, sets appropriate expectations for both you and your readers from the outset. A review that openly commits to future updating, rather than implying permanence it cannot genuinely offer, tends to earn more lasting trust from a field that already understands how fast this particular area moves.

Bringing in support where your team's expertise has gaps

Given how genuinely demanding it is to stay fluent in both evolving AI capabilities and rigorous systematic review methodology at the same time, seeking targeted consulting support at the specific points where your team's confidence is thinnest, rather than attempting the entire review in isolation, is a reasonable and increasingly common choice for teams tackling this kind of fast-moving, cross-disciplinary topic.

#ChatGPT#higher education#systematic reviews