Conducting a Systematic Review on AI Ethics
On this page
- Why this topic often calls for a different review type entirely
- Narrowing "AI ethics" to a genuinely reviewable question
- When empirical evidence does exist
- Appraisal approaches for conceptual and normative literature
- Search strategy across genuinely different disciplinary venues
- Framework and taxonomy proposals as a distinct source type
- Reporting standards for this kind of review
- A practical starting point for this specific topic
- Stakeholder engagement as a distinct methodological consideration
- Balancing normative claims with descriptive findings
- A closing consideration on this field's genuine value
- A final word on humility appropriate to this topic
- Considering interdisciplinary readership carefully
AI ethics as a research area is disproportionately conceptual, normative, and qualitative compared to most systematic review topics, and this genuinely shapes which methodology fits -- a systematic review here often looks meaningfully different from a standard quantitative intervention review, and recognizing this early prevents forcing an ill-fitting standard template onto a fundamentally different kind of evidence base.
Why this topic often calls for a different review type entirely
Much of the AI ethics literature consists of conceptual argument, framework proposals, and normative analysis rather than empirical primary studies in the sense a standard systematic review synthesizes. Depending on your specific question, a qualitative evidence synthesis, or even a structured scoping review mapping the conceptual landscape rather than pooling empirical findings, may genuinely fit your research question better than a standard systematic review structure.
Narrowing "AI ethics" to a genuinely reviewable question
This topic spans algorithmic bias and fairness, privacy and surveillance concerns, accountability and transparency frameworks, and the ethics of specific AI applications like autonomous weapons or hiring algorithms, each representing a substantively different literature. A review committed to one specific ethical dimension, in one specific application context, produces a coherent synthesis in a way an overly broad "AI ethics" review cannot.
When empirical evidence does exist
Some AI ethics sub-topics, particularly algorithmic bias and fairness, do have a genuine empirical literature -- studies measuring actual disparate outcomes across demographic groups from specific AI systems. Where your review addresses this kind of empirical question specifically, standard systematic review methodology, including appropriate risk-of-bias assessment for the underlying empirical studies, applies more directly than it would to the field's more conceptual literature.
Appraisal approaches for conceptual and normative literature
Standard risk-of-bias tools do not meaningfully apply to conceptual or normative philosophical literature, and forcing RoB 2 or a similar tool onto this kind of source produces a meaningless appraisal. Where your review includes this kind of literature, a different quality consideration -- argument coherence, engagement with existing relevant literature, clarity of normative claims -- is more appropriate, though this moves further from standard systematic review appraisal into territory closer to a structured literature or scoping review.
Search strategy across genuinely different disciplinary venues
AI ethics literature spans philosophy, computer science, law, and social science venues, each with different indexing conventions and different terminology for closely related concepts. A comprehensive search needs to span databases relevant to each of these disciplines, and philosophy-specific databases in particular are easy to overlook if your search strategy defaults to conventional health or social science sources alone.
Framework and taxonomy proposals as a distinct source type
A recurring feature of this literature is proposed ethical frameworks or taxonomies for evaluating AI systems, and if your review's purpose includes mapping or comparing these proposed frameworks specifically, this is a genuinely scoping-review-shaped task -- mapping the landscape of existing frameworks -- rather than a systematic review synthesizing empirical findings toward a single pooled conclusion.
Reporting standards for this kind of review
If your review functions more as a qualitative evidence synthesis or scoping review given this topic's conceptual nature, following ENTREQ or PRISMA-ScR respectively, rather than standard PRISMA 2020 built around empirical intervention synthesis, reflects your review's actual methodology more accurately to readers and reviewers familiar with these distinct approaches.
A practical starting point for this specific topic
Before committing to a specific systematic review methodology, honestly assess whether your AI ethics question is genuinely empirical, conceptual, or a mix of both, since this determines whether standard systematic review methodology, qualitative evidence synthesis, or a scoping approach is the right fit -- a decision worth making deliberately at the protocol stage rather than defaulting to standard systematic review structure regardless of whether your actual evidence base supports it.
Stakeholder engagement as a distinct methodological consideration
Given how directly AI ethics questions affect real communities and real people, some reviews in this space benefit from structured engagement with affected stakeholders, not just synthesis of existing published literature, though this moves beyond standard systematic review methodology into territory closer to participatory research approaches. Considering explicitly whether this kind of engagement genuinely serves your review's specific purpose is worth deliberate thought rather than assuming standard literature synthesis alone is sufficient for every ethics-focused question.
Balancing normative claims with descriptive findings
Where your included literature makes explicit normative claims about how AI systems ought to be governed or designed, clearly distinguishing these normative claims from descriptive empirical findings in your own synthesis prevents your review from inadvertently endorsing a specific ethical position as though it were an established, evidence-based conclusion rather than one perspective among genuinely contested views in this literature.
A closing consideration on this field's genuine value
Even where full methodological consensus remains elusive in a field this conceptually contested, a carefully conducted review mapping the genuine state of argument and evidence still offers real value to readers trying to navigate a complex, rapidly developing area of concern. That contribution, modest as it may sometimes feel given how unsettled the underlying debates remain, is still genuinely worth making carefully and honestly.
A final word on humility appropriate to this topic
Given how genuinely unsettled many foundational questions in AI ethics remain, approaching your review with appropriate intellectual humility, presenting your synthesis as a careful mapping of current thinking rather than a definitive resolution of contested questions, reflects both good scholarship and honest engagement with this field's genuine current state.
Considering interdisciplinary readership carefully
Your review's readers will likely span philosophy, computer science, law, and policy backgrounds, each bringing different assumptions and different vocabulary to this topic, and writing with this genuinely mixed audience in mind, defining specialized terminology clearly rather than assuming shared background knowledge, makes your synthesis considerably more useful across this field's naturally interdisciplinary readership. That accessibility, achieved without sacrificing rigor, is what makes a review on this genuinely contested topic useful to the full range of readers actively engaging with it, from philosophers and technologists to the policymakers ultimately responsible for translating debate into action.