Everything, by date
94 papers, posts and guides, newest first.
2026
Welcome to the new cjblunt.com
After twelve years on WordPress, this site has a new look, a faster home and a fresh start for the WritePhilosophy guides.
The Pyramid Schema: The Origins and Impact of Evidence Pyramids
[This paper has been updated in 2026 from the 2022 preprint] Evidence pyramids are amongst the most recognisable artefacts of the Evidence-Based Medicine movement. Yet no study has established the origins of evidence pyramids, or analysed whether they offer any information beyond simple lists or tables. In this paper, I establish the origins of the first evidence pyramid and argue that the pyramidal turn is a retrograde step in evidence appraisal.
2025
Cultural Bias in Language Models: The Top 10 Test
How can AI-generated "Top 10" lists of cultural influencers expose the cultural biases and defaults that are built into our most popular language models? I propose and report an initial test using lists of the Top 10 most influential musicians.
Golem.AI: The Experimenter's Regress and AI Decision Systems
How can the concept of the experimenter's regress help us understand the problems in AI decision systems? I explore two forms of circularity that can underpin AI decision tools through the lens of Collins and Pinch's 'The Golem'.
2024
NotebookLM: Cast with Care
Google's NotebookLM "deep dive" feature is taking off in popularity. I subject three of my academic papers to the deep dive treatment, and reveal its tendency to subvert content for a happier ending.
New publication: Grading evidence from qualitative research
Announcing a new paper co-written with a EULAR working group, reporting the findings of a systematic review of qualitative evidence appraisal tools.
Four Myths about Generative AI in Education
I outline four common misconceptions about the use of Generative AI which are widespread in Higher Education debates about the use of these tools: that it is possible and practical to detect the use of AI in writing, that text produced by GenAI is bland, repetitive or predictable, that GenAI tools struggle to cite sources accurately, and that more creative or reflective assessments are harder to complete using AI.
Cryptic AI
As language models are fine-tuned to acquire more capabilities, we continue to seek new tasks to push the limits of Generative AI. In setting cryptic crossword clues, so far GPT-4 fails the test quite spectacularly.
2023
The AI Skills of Social Scientists
What skills do social scientists need to adapt to generative AI, and how should educators approach teaching them? A narrow focus on training everyone in technical skills is misguided - what's needed is authorial voice, leadership and management skills, and the critical force of the social sciences.
Leary Philosophers
Can a language model outperform old Edward Lear in describing philosophers in Limerick form? "There once was a Scotsman named Hume..."
'I did write that text': Ownership and Authorship Claims by Language Models
Will large language models acknowledge authorship of their own generated texts? Will a language model claim authorship or ownership of texts which it did not create? A mistaken comprehension of ChatGPT's abilities throws up a distinctive problem of intellectual property rights.
Perplexity, Creativity and the Zonkamoozle
What is the relationship between perplexity, creativity and novelty? Following on from 'Perplexing Perplexity', I set out to demonstrate that high perplexity texts are not always creative, and to showcase ChatGPT's ability to work with and even generate novel words - culminating in the Tale of Zonkamoozle.
Perplexing Perplexity
Detectors such as GPTZero use the property of 'perplexity' to try to detect AI authorship of texts. But I show that by writing in a specifically dull style, or engineering the prompt given to a language model, we can easily and systematically fool such detectors to label AI text as human and vice versa.
Machine Evidence II: The Abstract Setting
A recent study by Gao et al. (2022) validates the warning of 'Machine Evidence' (Blunt, 2019) that language models would soon become capable of beating detection attempts by human peer reviewers. This piece looks at the near-term steps that journal editors and conference organisers can take to prevent AI-generated abstracts bypassing their screening processes, along with a warning for the long-term viability of those strategies.
2022
[Superseded] The Pyramid Schema: The Origins and Impact of Evidence Pyramids
Evidence pyramids are amongst the most recognisable artefacts of the Evidence-Based Medicine movement. Yet no study has established the origins of evidence pyramids, or analysed whether they offer any information beyond simple lists or tables. In this paper, I establish the origins of the first evidence pyramid and argue that the pyramidal turn is a retrograde step in evidence appraisal.
To the Tune of Pure Reason
Can a new iteration of GPT-3 write pop songs, raps and limericks...about Immanuel Kant's Categorical Imperative? It is a moral duty to find out.
League Table Ranking and Name Recognition
Does the ranking of a UK university in league tables affect its reputation and name recognition? Using a novel data source from an online quiz, this research explores the complex relationship between league table rank and identifiability of universities.
The Machine Scientists: Iatrophysics and Selective Scientific Realism
The Machine Scientists - Giovanni Borelli, Jan Swammerdam and Niels Steensen - followed Rene Descartes in modelling the human body mechanically. The scientific success of this 'iatrophysical' programme, replete with rejected entities such as 'Animal Spirits', poses a problem for Scientific Realism. Using Psillos' moderate realism, this paper attempts to reconcile a selective realist position with the historical record.
'Modern' Philosophers by DALL·E 2
Can DALL·E 2 create images of ancient philosophers like Plato, Aristotle and Immanuel Kant as they'd look in modern day universities? Sort of. Should it? Definitely not.
Bias and the Myth of the Objective Average
Would you choose a black box AI surgeon with a 90% success rate over a human surgeon with 80% success? The answer exposes a fundamental and harmful assumption within dominant models of medical evidence.
Higher Orders of Evidence
DALLE 2 offers a far more powerful image generation AI than the popular open access 'Craiyon'/'DALLE Mini' model. How does DALLE 2 compare to DALLE Mini's visions of hierarchies and pyramids of evidence?
Visions of Evidence
How does a machine learning algorithm picture hierarchies of evidence and evidence-based medicine - and what do these visions of evidence remind us of the way we understand, order and assemble the information we use to guide clinical practice?
The origins of evidence pyramids
What was the first evidence pyramid? A deep dive into the murky history of this novel way to present an evidence hierarchy reveals a significantly earlier origin that previously presumed.
2021
Lumping and splitting: brain tumours in white Britons
A much-publicised report suggests that white Britons' brain tumour survival rates are lower than other ethnicities. But analysing the ethnicities categories used, and considering the diversity of the "brain tumour" label, complicates the picture, as the 'Dismal Disease' of Glioblastoma continues to confound.
The Jurassic Critique of Micozzi on Evidence Hierarchies
AI21 Labs have just released a public demo of their giant language model, Jurassic-1. At 178bn parameters, it rivals GPT-3. Feeding it my own work, it generated some interesting and potentially novel views on evidence hierarchies... and then attributed them to CAM researcher Marc Micozzi! Is Jurassic Micozzi's critique of evidential pluralism in medicine sound?
On the Global Summit: do critiques of evidence hierarchies favour chiropractic?
The Global Summit systematic review claims that spinal manipulation therapy is not effective in preventing any non-musculoskeletal disorders. But a breakaway group has challenged their findings, in part based on my arguments regarding evidence hierarchies. Are they correct? Does my critique undermine the Global Summit review? If so, does the evidence base favour chiropractic?
Imitating Imitation: a response to Floridi & Chiriatti
In their 2020 paper, Floridi and Chiriatti subject giant language model GPT-3 to three tests: mathematical, semantic and ethical. I show that these tests are misconfigured to prove the points Floridi and Chiriatti are trying to make. We should attend to how such giant language models function to understand both their responses to questions and the ethical and societal impacts.
Minding the Gaps: Statistical Misrepresentation in Attainment Gap Research
Political interests configure the stories we tell with data. Closing the gap in attainment between disadvantaged students and their advantaged contemporaries is pivotal to an agenda to use education as a positive social force. But both the measurement and representation of this gap is politicised, skewed and open to manipulation. This paper shows how two organisations with inverse aims represent—and misrepresent—their measure of the attainment gap to portray diametric trajectories in the pursuit of equal attainment.
Conversion Therapy: Evidence is Irrelevant
Pressure mounts upon equalities minister Kemi Badenoch to resign over the UK government's failure to ban conversion therapies. Attention has focused on the government's failure to publish research commissioned in 2018. But evidence about whether conversion therapy works is irrelevant: conversion therapy is not a medical intervention.
The Stochastic Masquerade and the Streisand Effect
What does Google have in common with Barbra Streisand? Since Google fired AI ethicists Margaret Mitchell and Timnit Gebru, our attention should turn to what they don't want us to read: "On Stochastic Parrots". Will the attempts to suppress this paper lead to it being overlooked, or will Google face Barbra Streisand's fate?
A new WritePhilosophy.com
A new version of WritePhilosophy.com is launching today. WritePhilosophy is a resource for students and teachers of philosophy, built on guides, articles and quizzes which help to immerse students in the concepts, language and principles of philosophical writing.
The True Causal Effect
Philosophy of medicine helps medical scientists to clarify their thinking. A maladept phrase like 'the true causal effect' serves to show why we need it.
2020
195 Hierarchies: A Systematic Database Update
The database of evidence hierarchies has been updated based on a new systematic review of the medical literature, and now contains over 195 hierarchies.
Causal Relevance (Philosophy of Diagnosis, Part 3)
This series of philosophical papers unpacks six philosophical issues in diagnostics and develops a pluralistic model of diagnosis. This paper analyses the role of causal relevance in diagnostics. Are diagnoses defined by their causal relevance to symptoms?
5 years: Hierarchies of Evidence in Evidence-Based Medicine
In the five years since the publication of Hierarchies of Evidence in Evidence-Based Medicine, what has changed and what lessons can philosophers learn?
Pathognomy, Sine Qua Non and Constitutive Matching (Philosophy of Diagnosis, Part 2)
This series of philosophical papers unpacks six philosophical issues in diagnostics and develops a pluralistic model of diagnosis. This paper presents a set of minimal constraints which any theory of diagnostics must satisfy based on pathognomy and sine qua non relationships.
Socrates in the Dungeon
What happens when you ask a machine learning language model tuned to create a D&D style adventure to instead produce a Socratic dialogue?
Automatic Gadfly: Socrates by Machine
I was inspired to see if AI language model GPT-2 could create some thought-provoking - or just weird - new Socratic dialogues.
Problems of Diagnosis (Philosophy of Diagnosis, Part 1)
This series of philosophical papers unpacks six philosophical issues in diagnostics and develops a pluralistic model of diagnosis. This introductory paper outlines the six roles of diagnostics and distinguishes and relates the six problems of diagnosis.
Are hierarchies history?
Even if the hierarchies of evidence are on their way out in EBM in theory, they could still be enjoying the spotlight in practice.
Dual Use Technology and GPT-3
Yesterday, AI researchers published a new paper entitled Language Models are Few-Shot Learners. This paper introduces GPT-3 (Generative Pretrained Transformer 3), the follow-up to last year's GPT-2, which at the time it was released was the largest language model out there....
Random Reflections: Cochrane and the Origins of Hierarchies
In 1972, Archie Cochrane publishedEffectiveness and Efficiency: Random Reflections on Health Services. In a little under 86 pages, Cochrane offers a wide-ranged but succinct delivery of his experience and his philosophy of evidence in clinical practice. It's a fascinating...
Aphorisms from the Automatic Philosopher
GPT-2 is a large language model capable of generating some of the most convincingly human-like text we are yet to see from artificial intelligence. Previously, I've used GPT-2 to generate reports of clinical trials and several paragraphs of an essay on irrationality in which...
Automatic Philosophising
The following philosophical musings were generated by a machine learning language model called GPT-2. They were created by a very weak version of GPT-2 which was released to the public in May 2019. The full version has nearly 5 times as many nodes and produces much more...
Crisis Thresholds: network demarcation and the Kuhnian turning point
In at least some of his work (e.g. 1962), Kuhn refers to a ‘crisis’ within the scientific community, which occurs when the build-up of anomalies becomes so substantial that most of the community begin to search for an alternative to the dominant paradigm. The crisis...
2019
A Ghost of Progress - How Hierarchies Become Fixtures
I have written extensively on hierarchies of evidence in evidence-based medicine. The origin story of hierarchies of evidence is a little contentious. Several sources in EBM cite Campbell and Stanley's 1963 classic "Experimental and Quasi-experimental Designs for Research" as...
The Dismal Disease: Temozolomide and the Interaction of Evidence
Evidence interacts. To understand the evidence base for an intervention we must look not only at individual studies, but the relationships between them.
All that Glitters is not... Evidence
Recently, I wrote about a new machine learning model called GPT-2 which was conspicuously not released by OpenAI. GPT-2 is a massive language model which can be used to generate often highly convincing text when given a prompt. Using the 'attention' framework, the language...
Machine Evidence: Trial by AI
Take a look at the following snippets from descriptions of clinical trials, thinking about how you'd rate the quality and strength of the evidence that comes from each:
Stop Fighting, You're Making Me SAD
Does medical science have to be adversarial to make progress? Does progress in patient care consist of weeding out treatments which are ineffective and replacing existing therapies with new and better alternatives?
Echoes of Evidence
Do you have an idea of what makes good medical evidence in your mind? Do you have ideas about the kind of evidence you’d pay attention to, and the kind of evidence you can safely disregard? Or ideas about the evidence that would strongly sway your beliefs and actions, and the...
The Authority of Evidence-Based Medicine
In the early part of the 20th century, the philosopher Ludwig Wittgenstein sought to demonstrate that metaphysical claims are meaningless. Statements which couldn’t be proven true in some way—through logic or evidence—weren’t even false, they had no meaning at all. But he ran...
The Positivity Machine: "Evidence-Based Alternative Medicine" and Grades of Recommendation
In his beat-poem Storm, the musician and comedian Tim Minchin lays out an apparent paradox. By his criteria, there can never be evidence-based alternative medicine. As soon as there’s evidence that a treatment works, it joins scientific medicine. Minchin goes on to list...
2018
The Parachute Problem: Extracorporeal Life Support and the Demand for Trials
The United States Parachute Association recorded 120 deaths while skydiving in America in 2008-2015. Most were due to human error, while others resulted from collisions with other parachutes or aircraft. Only 5 were due to equipment failure. With around 3 million parachute...
Medicine Needs Diverse Evidence
Do not centralize evidence. If you believe that all the evidence you need to make your decisions is of a single kind, you will be so much easier to mislead. Those who say 'I’ll believe it when I see this kind of study’, once their singular preference is declared, are so much...
Two Reactions to Philosophy in Medicine
I spend a lot of my time talking to healthcare practitioners of all kinds about evidence, formally and informally. I'm mostly focused on getting the word out that RCT results are not the only thing that matters in treatment recommendations, that hierarchies of evidence are...
The Avoidable Scandal: Benoxaprofen and Theories of Medical Evidence
Benoxaprofen, marketed as Opren in the UK, created a scandal in 1982 when it was withdrawn from sale amidst reports of over 60 deaths and thousands of adverse reactions. The question of who was at fault for this disaster has been addressed many times in the academic...
UBL and Variation
I've argued that information about variation in treatment effects is vital for doctors, patients and regulators alike. This information does not come from RCTs. Nonetheless, we can acquire high-quality, compelling evidence of variation. The case studies I've presented in the...
Failures of Diagnosis
Of late, a confusion has emerged in the diagnostic and philosophical literature concerning the ways in which diagnostic failure can occur. It's rare to see clear attempts at a taxonomy of diagnostic failure. Here, I want to disambiguate a number of ways in which clinicians...
2017
Pluralism and the Problems of Demarcation
In 1983, Larry Laudan proclaimed the “Demise of the Demarcation Problem” (Laudan 1983). I will argue that the ‘simple demarcation’ problem should indeed be abandoned. However, abandoning the quest for simple demarcation does not entail that demarcation problems are...
2016
Hierarchies of Evidence: Database of Hierarchies
Hierarchies of Evidence are a tool employed by many advocates of Evidence-Based Medicine. They are used to appraise evidence from a range of sources, as well as to teach medical students about evidence and evidence appraisal. My PhD thesis concerns the variation in hierarchies defended, and the range of philosophical interpretations of those hierarchies.
Against "Effective Treatments"
I am agitating for philosophers of medicine and philosophically-minded clinicians to lead the charge against the term “effective treatment”. This phrase has become ubiquitous in the medical and philosophical literature. But it is a misguided choice which misleads the public...
2015
An information ordering thing
Randall Monroe of xkcdstarted writing explanations of complex physical phenomena using only the most common 1,000 English words (or the most common 'ten hundred', as he'd have to put it). He recently released a straightforward application which allows writers to check which...
The Exhaustiveness Problem
Hierarchies of evidence can be classified as exhaustive or inexhaustive. A hierarchy is exhaustive if and only if it gives some ranking to any given piece of medical evidence - there's nothing it leaves out. But there's a dilemma which hierarchy authors face: either way,...
Two Dogmas of Evidence Hierarchies
Hierarchies of evidence in Evidence-Based Medicine (EBM) come in many varieties and have been very influential in medical practice and policy since the late 1990s. However, two fundamental problematic assumptions underpin the use of hierarchies of any kind in clinical...
2014
Philosophical Arguments (WritePhilosophy Guide)
A philosophy paper consists of an argument for a thesis. The quality of your paper will be judged primarily on how well your argument supports your thesis. But what is a thesis? How specific should it be? How do we construct arguments, breaking them down into a series of premises and a conclusion? What makes an argument persuasive?
Validity and Soundness (WritePhilosophy Guide)
We laid out two criteria for a persuasive argument: that the conclusion follows from the premises, and that the premises are all true. But what does it mean for a conclusion to "follow from" the premises? In deductive inference, we want to know whether an argument is valid and whether it is sound. What do these terms mean and how are they used?
Abstracts and Introductions (WritePhilosophy Guide)
Where to start? Writing an introduction or abstract for your philosophy paper can be daunting - and with good reason. The first paragraph of your paper is also the most important. But how do you write a good introduction? What should you include and what must you leave out? And how can writing an introduction help you to structure your paper?
Reading Effectively (WritePhilosophy Guide)
It sounds odd to say that you don’t know how to read academic papers. You start at the beginning, keep reading until the end—right? But the problem of being unable to cope with academic reading is probably the most common complaint amongst students at all levels. The truth is that reading for academic purposes is just not the same as other kinds of reading. You have to read actively, selectively and purposefully.
Analysing an Argument (WritePhilosophy Guide)
How can we read complex papers to extract the philosophical argument? What is the Principle of Charity? Reading philosophical texts can be challenging. Extracting the thesis and rendering the argument into premises-and-conclusion form is a great way to understanding the author's reasoning.
Structure (WritePhilosophy Guide)
"Poorly structured" and "unstructured" are very common criticisms in feedback on essays. But what does that mean, and how should you respond? Not all structures are created equal. This guide reviews six common essay structures to help you create a stronger structure for your paper.
Using Literature (WritePhilosophy Guide)
How many sources should a philosophy paper use? How do you cite these works and avoid accusations of plagiarism? How do you present other philosophers' ideas and arguments in your paper? Should you include quotations? These are very common questions which students face. Let's look at the ways you should and should not use literature in your essays.
Conceptual Analysis (WritePhilosophy Guide)
Philosophers perform conceptual analysis to understanding the meaning of terms from 'death' to 'love', 'good' to 'evil' and 'science' to 'art'. How do we evaluate conceptual analyses? What are necessary and sufficient conditions? What makes a conceptual analysis too strong or too weak?
Distinctions (WritePhilosophy Guide)
Philosophers often draw distinctions, dividing things into categories. How do we evaluate distinctions? What is a 'proper distinction'? What makes a distinction exhaustive or mutually exclusive? And what does this have to do with suckling pigs and mermaids?
Criticising an argument: Reductio & Dilemmas
Two powerful ways of criticising a philosophical position, reductio ad absurdum and dilemmas, and how to respond when they are used against your own view.
Counterexamples (WritePhilosophy Guide)
Some of the most powerful philosophical criticisms provide a counterexample to a claim. But how do we formulate counterexamples? And if there is a counterexample to a claim you want to defend or an argument you want to make, how should you proceed? We'll look at Gettier's famous counterexamples to find out.
Fallacies (WritePhilosophy Guide)
A fallacy is a mistake or error in reasoning. Fallacies can be accidental errors, or can be deliberately crafted to be misleading. Many different types of fallacy have been described across many fields. Understanding fallacies is useful for philosophers to identify fallacious reasoning in arguments which we are criticising, and to avoid committing them in our own work.
Formalising Arguments (WritePhilosophy Guide)
Formalising arguments allows us to look at the underlying logical form of the premises and conclusions. We can use this tool to determine the validity of the argument. How do we find the logical form of a proposition? This guide introduces the basics of propositional logic.
Truth Tables (WritePhilosophy Guide)
Truth Tables are a useful tool for analysing propositions and arguments alike. They allow us to give precise definitions for our logical connectives. They also allow us to show systematically that a proposition is a tautology or contradiction, and to prove that an argument is valid or invalid. How do we create truth tables?
Logical Equivalence (WritePhilosophy Guide)
What does it mean for two propositions to be logically equivalent? When can we swap out a proposition for another? Using the method of Truth Tables, we can formalise a way to discover whether two propositions are logically equivalent, and deploy this method to transform our claims into the most useful form to deploy them in arguments.
Editing (WritePhilosophy Guide)
How do we draft, edit and proofread philosophical papers to ensure that what we're turning in is high-quality? What's the difference between a Zero Draft, a First Draft and a Final Draft? This guide article gives you a range of tools, tests and techniques to improve your paper immeasurably through editing and proofreading.
Banned Words (WritePhilosophy Guide)
You can improve your philosophical writing by removing certain words and phrases. Some phrases weaken your papers by introducing vagueness and ambiguity, by wasting words, or by acting as a crutch to prevent you saying what you really need to say. Here we've compiled a list of "banned" words and phrases. By banning yourself from using these, you can improve your philosophical style.
Areas of Philosophy (WritePhilosophy Quiz)
Can you tell epistemology from metaphysics? Aesthetics from ethics? We have created this quiz to test your understanding of the terms for different fields or areas of philosophy.
Fallacies Quiz (WritePhilosophy Quiz)
Can you tell your ad hominem from your Eminem? Pick out a straw man from amongst the scarecrows? We created this quiz to test your understanding of the argumentative fallacies from the Fallacies guide.
Paradoxes Quiz (Write Philosophy Quiz)
Do you know your liar paradox from your Ship of Theseus? Can you tell a Catch 22 from an Unexpected Hanging? This WritePhilosophy.com quiz will test your paradox identification skills.
Logical Validity (WritePhilosophy Quiz)
Can you tell a valid argument from an invalid one? This quiz will allow you to practice identifying valid arguments. If you struggle with the quiz, go back and check out the 'Validity and Soundness' guide article.
Logic: Basic Propositions (WritePhilosophy Quiz)
Test your skills at formalising arguments by trying your hand at this quiz. Can you identify which sentences have which basic propositional forms?
Logic: Compound Propositions (WritePhilosophy Quiz)
This quiz will test your ability to formalise more complex sentences into propositional form. You should try this after you're confident in formalising basic propositions.
An Honest Mess
Evidence-Based Medicine is an attempt to simplify and streamline a complex reality into a more manageable structure. Reflective EBM proponents know that matters are very complicated. But they understand that practitioners have very limited time (and abilities) to deal with...
2013
Festive Fallacies: When the snowman brings the snow
When the nights are long and the days short, there is some solace in the familiarity of classic Christmas songs. But stuffed with pigs-in-blankets and addled on mulled wine, festive pop lyricists seem particularly prone to fallacy. What better way, then, to teach errors in reasoning than to dissect these Xmas classics? Humbug!
Strength of Recommendation - the Rudner Problem
Strength of Recommendation hierarchies such as SORT and GRADE go a step further than your standard hierarchy of evidence. Standard hierarchies tend to rank or rate the evidence provided by a study on a scale of quality, strength or validity. Strength of Recommendation...
Ps in a POD - definitions in circles
In my paper "The Disunity of Evidence-Based Medicine", I lay out four distinct definitional strategies employed by proponents of Evidence-Based Medicine, each conveniently beginning with P: Platitude, Paradigm, Principles and Process. I made a bad pun about "Ps in a Pod" in a...
The Disunity of Evidence-Based Medicine
Critics and advocates alike have expended much effort defining Evidence-Based Medicine. However, there has been little consensus about what “Evidence-Based Medicine” is. Some authors see ‘EBM’ as something one believes—a view about medicine. Others interpret EBM as something...
2012
RCTs - a note on terminology
For the most part, when I'm talking about an RCT, it's understood as a Randomised Clinical Trial, though I use them interchangeably because I can get away with it - but not everyone can.