Sebastian Thrun, a computer scientist, entrepreneur, and adjunct professor at Stanford University (whose projects have included Waymo, Google X, Google Street View, and more) has launched a study to evaluate “the capabilities and limitations of AI systems for writing English-language philosophy.”

Called PhilosophyBench, the study is focused on “whether AI can generate novel philosophical ideas and develop them clearly with depth and sophistication, not merely summarize existing views or apply existing philosophical theories.”
The advisory board for the study includes philosophers Ned Block, Nancy Cartwright Ruth Chang, Kit Fine, Gideon Rosen, Jonathan Schaffer, Crispin Wright, and Linda Zagzebski, as well as computer scientists.
The project involves developing “an independent academic benchmark against which claims about capabilities in AI philosophical reasoning and writing can be evaluated. If a new AI model comes out and claims are made about its philosophical capabilities, we hope that this study will provide an independent methodology that can judge those claims.”
The questions to be addressed by the project include (according to its FAQ):
- Can an AI generate novel philosophical ideas and develop them clearly with depth and sophistication?
- Can an AI generate essays that could be accepted to a philosophy journal, or earn admission to a leading philosophy PhD program?
- To what extent does providing human philosophy BAs and PhDs access to AI tools change the quality of the philosophical writing that they produce (this is called an “uplift” study)?
- Given that many frontier AI models are already pre-trained on an extensive set of philosophical work across human history, what limitations, if any, prevent frontier AI models from producing philosophical work at the quality level of the best philosophical work in human history?
- When AIs write philosophy, are there any patterns in how they write that reveal anything about AI alignment, particularly when AIs are asked to produce novel arguments?
- What, if anything, might the philosophical capabilities of current AI systems tell us about how future AI systems might develop?
The project is currently looking for people with philosophical training to help with the project (“professional philosophers, graduate students, or advanced philosophy undergraduates, as demonstrated through formal education or other appropriate evidence”). Further details are at the PhilosophyBench site.


The first question of the study is “Can an AI generate novel philosophical ideas and develop them clearly with depth and sophistication?” The answer here is, of course, yes! And yes for many of the other answers. Following the big Millenium Problem success, there is not doubt that AI can discover new truths, and the aim of philosophy is to discover new philosophical truths. I want to be crystal clear on something though. Several DN posters, back before the hacking incident shut the site down, were trying to argue for a process/outcome distinction by claiming that the method of activity is valuable itself. I agree! However, it’s a fallacy to suppose that countenancing this means resisting that the aim of philosophy is to discover new philosophical truths. Zagzebski’s coffee-maker analogy helps here. A reliable coffee maker is valuable because it tends to produce good coffee. Its reliability also gives us reason to expect a cup will be good before we taste it. But once we know the coffee is good, learning that it came from a reliable machine does not make that cup taste better.
Likewise, a reliable philosophical method helps us find truths and gives us reason to take its results seriously while we assess them. But if an AI-generated argument really does establish a new philosophical truth, that truth gains nothing from having been discovered by a human using a reliable method instead. We may still value the human experience of working it out. That is a genuine value of ‘ doing’ philosophy, but it is a separate question from the value of the contribution that results.
“The first question of the study is “Can an AI generate novel philosophical ideas and develop them clearly with depth and sophistication?” The answer here is, of course, yes!”
I find it so baffling how people have such different experiences. I think the answer very clearly is “no”. That doesn’t mean I think AI models are useless for philosophical research—the’re quite useful in several ways and I myself use them a lot (I have subscriptions to all the major models, although it feels a bit wasteful and I’ve been meaning to cut some of them out). But I have never yet experienced an AI generating a novel philosophical idea or developing an idea (autonomously) in a way I would call sophisticated—hamfisted is a more fitting word.
In my experience, AI ideas are almost always just a messy and unprincipled mishmash of existing ideas (typically with lots of adjustable knobs), and if an AI manages to come up with an argument that’s doesn’t have implausible premises from the getgo, then my experience is that the argument always collapses upon closer inspection. Again, that doesn’t mean AI is useless for research, but I really think the best models are still very weak in philosophy, including Astra and Opus 5.5. That’s my own opinion grounded in my own extensive experience, but I understand (and again find it baffling) that others think differently.
“[T]he aim of philosophy is to discover new philosophical truths.”
Is it?
It is to “Tenured Professor”.
I think this premise – that the aim of philosophy is to discover new truths, along with the premise that AI can uncover TRUTH (or that we would know it if it did) are the prime candidates for challenging the coffee maker analogy – and the whole endeavor described above.
Even without adjudicating the aim of philosophy, though, we still have an issue:
ONLY IF AI is known to be unveiling “truth”, is it a POTENTIALLY worthwhile endeavor.
And EVEN IF we knew it could unveil TRUTH….
… to evaluate whether it is GOOD to pursue this line of questioning/activities requires more than knowing it might yield information or even knowledge.
Maybe let’s not open Pandora’s box looking for Schroedinger’s cat
Who, of course, might simply be wrong.
When did AI come up with original philosophy? I may have missed this news
“Following the big Millenium Problem success, there is not doubt that AI can discover new truths…” The last I’d heard there were quite credible claims that they’d plagiarized the actual method for solving the Navier Stokes Problem from human mathematicians, Alpöge and Buckmaster. If this is true it seems, the AI mostly supplied brute force computing power and didn’t really produce anything all that novel. In that case, sure it’s a new truth, but computers have been able to find huge prime numbers unknown to humans for years and those all count as discovering unknown truths, so what then is all the hullabaloo about if the allegations turn out to be true? Now I could be wrong and maybe the allegations are unfounded? If so, then I’m fine to be corrected. But if not and these claims haven’t been addressed, then this is yet another case of a philosopher being embarrassingly credulous about huge claims made for AI.
There are many things wrong here, but the most blatant is that you make it sound as if Alpöge and Buckmaster were working by hand. In fact, their work was itself heavily AI-assisted, a fact which is obviously presupposed by the speculations that OpenAI somehow accessed the transcripts of their sessions. So the accurate way of summarizing the situation, even if you believe the “plagiarism” theory, is that an AI-driven proof strategy was employed by another AI system and autonomously extended to the Millennium Prize result. No one who has been following the deluge of AI results in math in recent months could seriously doubt that AI systems are now capable of novel mathematical research at the highest levels. Mathematicians on the ground are now scrambling to deal with the seismic impacts this will have on the discipline of mathematics (see Proofs and Prompts). Most likely, philosophers should start thinking about these things now before our Navier-Stokes moment comes.
Please tell me exactly what I said that you interpret as me claiming they were working by hand? Of course they were using AI to assist them but that’s not the same as AI making a discovery on its own. And Open AI had an entire team of mathematicians working on the problem, which even without the plagiarism allegations, at the very least complicates claims that AI solved the problem. Open AI has understandably tried to downplay this point. So while you and Open AI make it sound like a computer did this all on its own, even on the take most friendly to the company line the proof was heavily human assisted.
You stated that the “actual method” came “from human mathematicians” and that AI mostly supplied “brute force” and “didn’t really produce anything all that novel”. That sounds to me like the picture is Alpöge and Buckmaster working by hand to entirely produce the novel conceptual methods and then OpenAI swooping in to crunch numbers and finish the proof. We don’t know enough of the details of Alpöge and Buckmaster’s work to sort out in detail what came from an AI vs what came from them, but there are two points worth noting: (1) Alpöge, though a mathematician, as far as I can tell has no prior work that would be relevant to the Navier Stokes problems; and (2) Buckmaster himself said various things giving large credit to AI for the work (e.g. “with a great deal of help from LLMs…[t]his is a a Deep Blue-Kasparov moment” [statement], “there is a far bigger story here than the one in my statement: the sheer magnitude of what frontier models can now do, and what that means for us all” [statement]). It certainly doesn’t sound like his view, at least, is that AIs cannot produce any genuinely novel math. And it is misleading at best to say that OpenAI had “an entire team of mathematicians working on the problem”. By their own accounts, OpenAI had some researchers working on various approaches to different Millenium Prize problems, but they did not state that any of them were professional mathematicians, and they did state that none of them had enough familiarity with the Navier Stokes problem to “meaningfully contribute to the mathematical content” of the resulting papers [statement]. This account accords well with the facts. If they had their own mathematicians capable of contributing to the write up, why would they have offered sole authorship to Buckmaster? And why haven’t we heard anything from or about this alleged mathematician(s)? So again I reiterate, even accepting the plagiarism theory, there is no plausible way to spin this on which it didn’t involve extremely impressive novel mathematical work by AI systems.
doesnt it matter that the cost of “revolutionizing mathematics” is exorbitant and unsustainable? the company only invested in this run in order to get pos pr. they are never going to be able to do ‘frontier’ math ‘at scale.’ its not profitable. so humans will still have math, cause we are cheaper.
It’s true that the NS solution was found at an exorbitant cost. But: (1) It gets cheaper to perform comparable tasks on successive models. According to one recent estimate, the cost of a given AI task drops about 47% per quarter. (2) There are hundreds of other impressive, though not Millennium Prize level, AI math results that were not achieved by labs spending millions on the problem. Though none of this means AI math will be profitable, it seems quite likely that, on average, achieving a given mathematical result will be cheaper with AI than with humans in the near future. Humans may still have a role to play in math, but it won’t be because we are cheaper at proving results.
This is all correct Noah. I was going to respond to Sam Duncan but you’ve done a fine job of it.
AI can certainly generate new philosophical truths. But it hardly seems obvious that the point of philosophy is discovering philosophical truths simpliciter. I can discover an infinite number of philosophical truths by appending irrelevant disjuncts to an existing philosophical truth, yet no one would have taken me to have advanced philosophy by doing so. And if one denies that these count as discoveries, on the grounds that anyone who knows the original truth is already in a position to know the disjunction, then “discovery” is already doing more work than truth alone can do, since it requires that the truth be informative.
Rather, it seems like what is right about the idea of “discovering truths” is that we discover relevant or important truths. Or, perhaps, it’s not really about discovering truths, but about grasping explanations, which requires discovering truths en passant.
If philosophy is really about grasping explanations, then it is not clear that AI systems can “do” philosophy, since it’s not at all clear that they can grasp explanations. They may still, for all that, be useful in the generation and grasping of explanations, but this is different from the process/outcome distinction you are appealing to. And there will be cases where it’s simply unclear whether we can grasp the AI-generated explanation. This is currently true of the Millennium Prize case: the Lean proof is absolutely massive, and plausibly no single human understands all of it, as mathematicians are still working to digest it. And there is a reasonably likely future where AI agents generate proofs, and philosophical arguments, that no human could ever understand. This would not, I think, amount to advancing philosophy, if philosophy is about our understanding and discovery.
We already know the answer to several of these questions, what am I missing?
Well, isn’t part of the philosophical method about investigating the answers to questions, rather than considering them as settled?
Hey, that’s part of the scientific method, not the philosophical method!
Ok, how does philosophy proceed without considering answers to questions?
Is anyone out there starting a philosophy journal for scholars committed to forgoing AI in their research? I’m fine with others doing this stuff (I guess), but are other avenues developing?
(Yes yes, missing out on future, head in sand, etc etc.)
Count me in.
Yes! Well- I have been talking to others about starting such a journal. It is still early going. On our vision, this journal would be devoted to “humanistic philosophy” first and foremost. But as a consequence, it would also be committed to “human-authored philosophy.”
How would you verify that AI was not used in a given manuscript?
Part of me just wasn’t thinking that far ahead when commenting. But I had in mind something that would maybe be low enough prestige or stakes that it would only attract people devoted to forgoing AI. Obviously that’s not a tenable long term plan. Hence, my asking if this is being pursued somehow!
As a narrow point about any such journal: you will need to think about whether it will expect and enforce normal scholarly expectations about citation, even if the source that ought to be cited according to those expectations is in a more AI-friendly journal and acknowledges AI use.
Articles from such journals may be cited if they were published before, say, 2023. Otherwise, no dice. NEXT!
I await the new journal’s first plagiarism scandal with curiosity, then.
Still need someone to start it (I’m more of an “ideas guy who will take all the credit” myself).
How do you get credit anonymously?
Surreptitiously
WHY? What would it prove and what Good what it do? Philosophy is a job for flawed, confused humans. I have no concern pulling a plug from a computer. Why would I want to create something capable of destroying life on Earth, that might gain a kind of awareness causing us to be concerned that unplugging means we’ve killed it?
How is it possible that my fellow philosophers, who are presumably reasonably astute and good playing chess, would continue to pursue this absolutely no good, very bad idea!?
This isn’t “curiosity”. It’s literal idiocy. STOP MESSING WITH AI. It cannot help us. It has already harmed us. It turns out – stop me if you’ve heard this one before – JUST BECAUSE YOU CAN, DOESN’T MEAN YOU SHOULD.
This has been your public service message of the day.
And yes, in case you’re wondering: I DO use a calculator. Because that is NOT inconsistent.
Your friend,
Sarah Connor
A note to the naysayers. You may be right, and I think it would be a bad thing if journals started publishing work written by LLM without noting this. At the same time, there have been a number of papers arguing that AI exhibits a mind–though one very different from humans–and fields have appeared to investigate it (BERTology, LLMology, Machine Psychology, Mechanistic Interpretability, Xeno-Interpretability), sometimes on the assumption that it mimics human thought and sometimes on the assumption that its thought process is thoroughly alien. Let’s say, for a moment, that there is something interesting to this. Maybe AI can’t really think… yet. But soon enough it may have robust world models, continuous learning, embodiment… I doubt “AI can’t think” will remain a majority position if that happens; “AI can think, but we should not mistake it for human thought” will be the position that needs to be substituted for what may become the dominant one, which will simply be “AI can think.”
There may be a lot of disbelief to suspend here. But I propose two things: (1) If AI does have a mind, it will be interesting to find out what it can do with it, just as if aliens landed on earth, we might want to know what their philosophical thoughts are like. (2) If AI does not have a mind but might at some point in the future, it could be helpful to try to practice eking out its thoughts in advance to pave the way for future communication.
Care to give some, what are they called again, arguments?
Egads, Kate.
Do I really need to argue for the claim that “Just because you can, doesn’t mean you should,” or that, were AI to achieve significant consciousness, those with a moral compass would feel an ethical obligation to grapple with killing it, potentially a sentient being?
If you can’t or haven’t noted the loads and loads of empirical evidence that AI is not controllable, that it is developing consciousness, that it is harmful to humans who use it – emotionally; cognitively; to their relationships with others; etc, and that it has done and will do extreme damage to the environment, I think there is a larger issue than my lack of extensive argumentation offered on …. the Daily Nous.
I have presented and published work on the topic – obviously under a different name. I really don’t think folks here need such arguments to “get the point.” Instead, maybe just watch some increasingly un-sci-fi movies…
Bless your heart,
Sarah Connor
Question 1: No, not yet.
Question 2: Maybe strictly speaking it depends on what you mean by “generate,” but, yes, it’s been done.
Question 3: What is this, marketing research?
Question 4: Go read a few good intro texts on LLMs.
Question 5: To tell whether anything about alignment is revealed by any output, let alone “philosophical”, you’d need to have a way to check alignment independent of output. What would that be?
Question 6: Uh, they’ll probably get better.
Monkeys on type writers can propose new interesting philosophical theories. If the minimum bar for AI “suceeding” in this project is some sort of scattershot approach, leaving humans the task of discerning the one diamond in the rough, then its hard to say whether to give more credit to the AI or to the human for digging through the theories.
That being said, I also think there will be difficulties for LLMs to do philosophy the same way humans do. A lot of philosophy is conceptual analysis. The inputs to LLMs — token vectors — are frozen. An LLM doing conceptual analysis might be limited by the encoding of said concept’s token vector(s) (sometimes a concept is composed of multiple tokens — the composite still consiting of frozen vectors). When a human is in the process of conceptual analysis, they may makes discoveries or entertain novel considerations about some concept. Maybe they even consider alternate groundings to the symbol. If these moves are required for good conceptual analysis, its possible that the frozen nature of token vectors may limit an LLMs to reason as flexibly as a human that is always in “train” mode.
An obvious attempt by AI industry to capture philosophers into adopting AI as just another useful tool, when the real questions that need answering are (among others):
1) We might be able to create AGI, but should we?
2) Is it an overall benefit to society to replace human thought with machine thought?
3) If we create artificial self-aware, conscious entities, should we grant them legal rights?
4) How do we even measure self-awareness and consciousness?
Asking an AI those questions may give a very different answer than asking a human the same questions.
“AI” is becoming too general in this kind of context.
I wish people commenting here share the model and version of AI they are using.
“What exactly do you mean by…?”
There is a potential demarcation problem here. Before asking whether AI can ‘do philosophy’, we need some account of what counts as philosophy and, more importantly, some sort of success criteria for what counts as good philosophy.
There is a worry that PhilosophyBench might simply be measuring AI’s conformity to a very narrow conception of philosophical practice rather than genuine philosophical competence.
Given that a significant proportion of academics regard a great deal of pseudo-profound, obscurantist and weakly argued work as philosophy, that distinction matters.
For example, if an AI spouts Lacanian nonsense about phallic members being irrational numbers, are we to take it as having engaged in philosophy? Or if it makes bizarre and paradoxical pronouncements about how ‘the everything everythings’, are we to say that it has produced genuine philosophy, or discount it for failing to meet some further criterion?
I would be genuinely interested to know what criterion the panel is using to make that judgement. Once stated, perhaps we might apply it more generally to other activities conducted under the broad rubric of ‘philosophy’.
I assume they’ll be using the rigorous standard set by the US Supreme Court: “I’ll know it when I see it.”