If You Think You’re Refereeing an AI-Written Submission, What Should You Do?


You’re refereeing a paper and you come to suspect it was substantially and/or illicily written by an LLM. What should you do?

My suggestion:

  1. Keep in mind the possibility that you are mistaken as well as the possibility that the AI use falls within the journal’s rules and was disclosed during the submission process.
  2. Pause your work refereeing the piece, as your suspicions of illicit AI use may unduly negatively influence your assessment of the submission.
  3. Gather a few examples of passages that strike you as AI-written, briefly explaining why. (Do not feed the manuscript into an AI-detector, for two reasons: first, doing so may violate the norms of confidentiality referees are expected to abide by, and second, such detectors appear to currently be of questionable accuracy.)
  4. Write to the journal’s editor or managing editor expressing your concerns and sharing the evidence you’ve gathered.

With luck, the editor will look into the matter and either agree with you and relieve you of refereeing the piece, or the editor will convince you that you are mistaken and you can proceed to finish refereeing the paper.

But what if you aren’t so lucky?

One associate professor of philosophy wrote in with the following:

I recently accepted a referee invitation for a good generalist journal. I was completely convinced after about 20 minutes with the submission that it was mostly (if not entirely) written using AI tools of some sort. The paper was in my areas of research. The editor disagreed (or at least did not share my level of confidence), and so I was left with the awkward task of writing (brief) comments for the “author” to go along with my rejection. For obvious reasons, this seemed to me to be an obvious waste of everyone’s time. Nothing like this has happened to me before, and I referee fairly regularly. 

The professor is curious about the frequency with which reviewers are finding themselves in disagreement with journal editors and/or other reviewers on particular cases. Has this happened to you? What did you end up doing? How should referees and editors navigate disagreement on this?

guest

64 Comments
Oldest
Newest Most Voted
Inline Feedbacks
View all comments
Mark
Mark
20 days ago

and so I was left with the awkward task of writing (brief) comments for the “author” to go along with my rejection.

Or maybe the awkward task of telling the editor that your verdict is Reject but that you won’t be providing comments to be fed back into the model to generate the next version of the paper.

Sam
Sam
20 days ago

Have another LLM write the referee report.

Graham Harman
Reply to  Sam
17 days ago

An amusing solution in principle, but in practice reviewers are now often asked to promise in advance that they will not use LLM tools to generate reviews… Generally speaking, if I suspect that AI wrote an article, that manifests in tangible negative features that would lead to rejection even if a human had written it. The other day, for instance, I reviewed an article where at least 50% of the sentences were essentially just lists of terms separated by commas.

Kenny Easwaran
20 days ago

As an editor, I’ve received a bunch of these papers. As people may be able to tell from my comments here, I’m ideologically committed to the idea that there *could* be a worthwhile paper whose text is substantially AI-composed, but so far only one of the dozen or so papers I’ve had this suspicion about seemed promising enough to send out to a referee. (There was one other that I unfortunately identified *after* getting referee comments, and I apologized to the referee after the fact.)

As far as comments, I don’t think anything that is actually AI-authored is getting past the initial desk reject phase. If the paper looks like a real paper throughout, and makes a point and gives some sort of argument for it, there’s almost certainly a real human there who is thinking about the project. In some cases, it’s a person working across disciplinary boundaries, or across languages, relying on the AI to get the disciplinary norms and language right (but unfortunately, these AI systems aren’t actually good enough tools to be effective for this purpose yet).

I’ve been writing fairly substantial comments on a bunch of these in my desk rejections, pointing out formalisms they take pains setting up but never use, or ideas that they say they will argue for but no argument comes. I’ve been doing this because there’s often an interesting idea there that I really want to see someone write a good paper about!

But I’ll probably stop doing this when it’s just another paper written by AI, arguing that AI has constitute epistemic limitations and therefore can’t be a testifier or a knower (which seems to be the claim of about half of these AI-written papers).

Brian Earp
Reply to  Kenny Easwaran
19 days ago

Hi Kenny
I’m the EIC of JME Practical Bioethics (a BMJ journal, sister to Journal of Medical Ethics which I also edit) and we’re preparing a special issue on “Doing Ethics with AI” (as opposed to “the ethics ‘of’ AI”) where one of the main questions we’d like authors to debate is, under what conditions if any would it be okay, even desirable, for a philosophy/ethics journal to publish work substantially generated by AI, assuming it made a worthwhile contribution — and do we think it is even possible that a well-prompted AI could generate something that would count as a meaningful advance in the literature? Now, the focus is ethics research specifically rather than all of philosophy (because of the scope of the journal), but we’d welcome your philosophical reflections on the matter, perhaps in a short piece… I see Dan Greco’s commented below and has a similar attitude to yours; he’s submitted something already 🙂 More generally, anyone reading this, if you want to work through the philosophy and/or ethics of philosophers using AI to generate normative analyses and submitting them to journals – and what editors/journals should do, how readers should feel about this etc. – please let me know and I can send you the information on the special issue. I’m at [email protected] — cheers. Brian

Jade Schiff
Jade Schiff
Reply to  Kenny Easwaran
18 days ago

I hear what you’re saying and I get it. I think I’m more skeptical than you are, in part because of my reaction to this:

“If the paper looks like a real paper throughout, and makes a point and gives some sort of argument for it, there’s almost certainly a real human there who is thinking about the project.”

As AI continues to “get smarter,” or whatever we want to call it, that may become less and less certain and also harder to determine.

Daniel Greco
20 days ago

My attitude is similar to Kenny’s. So far I’ve seen some as an editor, and some as a referee. So far I’ve never thought: “this is a good paper, but it looks like it’s AI, so what do I do?” Rather, the kind of papers that are striking me as likely AI written are ones that go on for pages saying very little, or which have other serious substantive problems. So I’ll recommend rejection in a way that doesn’t have to cite their having been AI-written.

I suppose, like Kenny, I’m uncomfortable with the idea that an otherwise publishable paper–one that makes an interesting, worthwhile contribution to philosophy–should be rejected on grounds of AI authorship. Though that’s a bigger conversation. (Very roughly, my view is that just as a mathematical proof of a significant result is valuable because of the understanding it provides its readers, regardless of whether it’s mainly human or mainly AI-authored, likewise with a philosophy paper. If we get to the point that they’ve already gotten in math, where AI is playing a major role in generating important results, those results should absolutely be publicized, and the journal system is our discipline’s primary way of publicizing what we take to be worthwhile contributions that advance our collective understanding of philosophy.) But that’s also very much not what I’m seeing so far.

Eric
Eric
Reply to  Daniel Greco
19 days ago

I dont understand why this view is so controversial.

MBW
MBW
Reply to  Daniel Greco
18 days ago

I agree. If AI becomes a useful working tool, disclosing its use would make about as much sense as disclosing Word vs. LaTeX, or like any other methodology section. The rule should be that the author owns their work (and its flaws.)

Nick
Nick
Reply to  Daniel Greco
18 days ago

Just to be clear about what the argument is here (for those who claim to “not understand” it): we increasingly inhabit a social world where machines are doing everything for us. This has long been the case but it is accelerating rapidly. This (1) hugely empowers those lucky few who own the machines, it (2) removes or disincentivizes many activities which are constitutive of human flourishing, it (3) disconnects people from material and ecological reality, it (4) alienates younger generations who can only watch helplessly as viable creative life-paths are swallowed up by machines, and it is (5) a constitutive part of an economic system that is destroying our ecosystems and rendering parts of the planet uninhabitable.

So that’s freedom, health, meaningfulness and truth, all increasingly in danger from the mechanization of human activities. So, some of us believe that we should no longer collaborate in this process, out of a sense of integrity and loyalty to the human race and to the larger ecosystems of which we are a part. Others believe that this can all be outweighed by the benefits of “better products”.

What I would love to see is an argument showing how those quality products are actually going to outweigh the documented costs of mechanization. If 90% of excellent papers are written by AI in 2035, how exactly does this new philosophical value outweigh the costs to (say) the few surviving prospective graduate students, who will be robbed of a gift that we all enjoyed? Their scholarship is reduced to bot-prompting so that we can all enjoy a new and improved defense of the Quine-Duhem thesis? Those who want to enable this future should at least tell us why it is a good one.

Daniel Greco
Reply to  Nick
18 days ago

This might have been meant as a reply on a different thread–you directly quote “not understand”, I didn’t use that phrase in my comment, and the closest place something like that was used in the comments here was Eric’s comment below–but I’ll reply to it here.

I have the same anxieties you’re expressing about this whole process, and I sympathize with the preference to stand athwart history yelling “stop!” rather than commit to trying to find meaning in a world where so many of the activities which were previously part of a meaningful life now seem trivialized, as they involve producing stuff that can be better produced by AI. Like I said in my earlier comment, it’s a larger conversation, and I don’t particularly expect what I have to say to be all that persuasive, but here goes.

Imagine you’re a doctor working on curing (some form of) cancer. I’m sure that feels very meaningful, because it is. Perhaps almost to the point where if you learned tomorrow that all cancer had been cured, you’d be a bit disappointed, as you wouldn’t have available this meaningful life path left to you. You’d have some big decisions to make about what to do with your life, which would properly provoke a lot of anxiety. Still, it would be perverse, as a doctor, to hope that cancer won’t be cured, or still worse, to try to thwart someone else’s)= curing cancer, on the grounds that it would deprive you of a source of meaning. Right? Better to commit to finding alternative sources of meaning than to aim to keep everybody sick because you derive meaning from attempting to cure them.

That’s roughly how I feel about philosophy papers. If we’re writing papers not just as a game for ourselves where we’ve conned the public into paying for it, but because we think the existence of a philosophical literature somehow benefits people other than the philosophers who enjoy contributing to it–and I do think thisthen we face awkward questions when it turns out that those benefits can be enjoyed without professional philosophers creating them. I don’t think we’re there yet, but I can imagine that’s where we’re getting in the fullness of time.

What sources of meaning could be left in a world where it’s not just philosophy papers, and cancer cures, but all sorts of intellectual work that can be better done by computers than by us? I think the best thing written on this that I know of is the latter part of Bernard Suits’ book “The Grasshopper” concerning utopia. He’s not discussing AI in particular, but technological progress more generally, but I don’t think it would have looked all that different if he’d known what was coming. He thinks in a world where humanity faces far fewer genuine obstacles to achieving its aims, more of our meaning would have to come from game-playing, which he defines as, roughly, pursuing goals subject to constraints where you only accept the constraints for the sake of the constraint-driven activity that they make possible. You might think game-playing is trivial and meaningless, but I think if we intentionally thwart cancer-curing AI for the sake of preventing the obsolesce of human-directed cancer-research, we’re already treating cancer-research as a game, at least if you accept Suits’ definition (which strikes me as a pretty good one). Likewise, if we intentionally thwart understanding-enhancing philosobots for the sake of making possibile the activity of seeking-philosophical-understanding-without-AI, we’re treating philosophy as a game too. So I think once the possibility of AI is on the table, we face something like a forced choice between a very pessimistic picture where we have no hope of meaningful lives no matter what we do, or we instead reconcile ourselves to the possibility of deriving genuine meaning from game-playing.

Eric
Eric
Reply to  Daniel Greco
17 days ago

Fantastic post

Daniel Groll
Daniel Groll
Reply to  Daniel Greco
17 days ago

Thanks for this Daniel Greco. One difference, I think, between the “AI cures cancer” and the “AI does amazing philosophy” cases is that there’s no sense in which people need to appreciate the former in order to benefit from it, whereas in the latter they do (I think).

Certainly we don’t need to appreciate how the cure for cancer works and (less clearly) we could benefit from it without even appreciating *that* there is a cure (suppose, fort example, AI formulates a cure and puts it in our water supply such that all cancer disappears). But, it seems to me, there’s no point to having better philosophy papers if we cannot appreciate them. So in the philosophy case (and by extension a whole bunch of other meaningful endeavors), the meaning derives not just from the quality of the output but our ability to appreciate it. If we can no longer appreciate it, the quality of the output is irrelevant.

I’m not sure what this implies about handing over more (and more) of our philosophy to AI. But one might worry that by handing over more and more of our philosophy to AI, we will *lose* the ability to appreciate what it produces even (perhaps especially) if it is very very good, precisely because our ability to appreciate philosophy comes largely from our *doing* it.

Eric
Eric
Reply to  Daniel Groll
17 days ago

So dont stop doing it. What do you care what other people are doing? Thats their loss.

Daniel Groll
Daniel Groll
Reply to  Eric
17 days ago

Why the snark Eric? Serious question! My main point was that the philosophy case doesn’t map so neatly onto the cancer case. I acknowledged that I wasn’t sure what the implications are for whether/how much to use AI for doing philosophy.

But setting that aside: surely it’s not unreasonable to be concerned about other people (as individuals) and what the implications might be culturally.

Eric
Eric
Reply to  Daniel Groll
17 days ago

I wasn’t trying to be snarky. I was genuinely asking why you were assuming this would mean everyone would stop doing philosophy. Look at math. Pretty soon, very few people are going to be strong enough to prove things AI cant prove. But they still need a community do digest all the results that are coming out. And that community needs to be way better at math than you (I assume) or I am. So I don’t see why you think this leads to nobody doing philosophy anymore.

Daniel Greco
Reply to  Daniel Groll
17 days ago

Good points. Here’s an analogy that I’m not sure how I feel about. I like math. We home school our children and math is the one part I’m responsible for (my wife does the rest), and I like explaining it to them. As a summer side project I’m (slowly) working through a logic textbook with my eldest and we’re both enjoying it.

But I knew from pretty early in my life that I did not have anything approaching the talent to be a mathematician. And as I got into philosophy and enjoyed logic, I think I also knew I wasn’t going to be proving novel theorems, as much as I enjoyed my coursework. E.g., my college “Computability and Logic” course was wonderful; I felt like it gave me the chance to appreciate a sort of beauty that I would not myself be able to create. As much as I enjoyed the class, I didn’t come away from it aspiring to emulate Godel or Turing; it was clear to me I’d never be in a position to have insights like theirs.

If we get to a spot–I don’t think we’re there yet–where AI writes better philosophy than the best humans, then I think our relationship to philosophy might be more like my relationship to math. That is, we wouldn’t aspire to extend the frontiers of philosophy–that would seem as unrealistic as my aspiring to see what Godel or Turing could–we still might aspire to understand a lot. And to be more directly responsive to your point, the sort of understanding we want typically can’t just be achieved passively. If you want to understand math, you need to work through proofs yourself; you can’t just be a passive recipient of information. Similarly with philosophy and constructing arguments.

You might say that if the papers are out there, able to be produced on a whim, nobody will be motivated to read and understand them. Maybe! That would be sad. But I don’t think it’s at all inevitable, and I think the math analogy helps us see why. Most people who like math know they’re not gonna be Terrence Tao, and that’s fine. How much really changes about our relationship to math if even Terrence Tao knows he’s not gonna be Terrence Tao anymore? (Well, maybe a lot changes for him, but for the rest of us?) So maybe that’s where we end up with philosophy. In some moods, I think that’s not so bad at all. In others, it’s still kinda deflating.

John
John
Reply to  Daniel Groll
16 days ago

There is a logical problem here: we cannot infer that a person has lost the ability to appreciate something simply because he hasn’t done it. This is similar to claiming that the public doesn’t have the ability to appreciate a painting because they don’t know how to paint.

Felix
Felix
Reply to  Daniel Greco
17 days ago

…find meaning in a world where so many of the activities which were previously part of a meaningful life now seem trivialized, as they involve producing stuff that can be better produced by AI.

In many areas, that’s just empirically false though. It’s not “better produced by AI.” Unless one is aiming for a steady stream of mediocrity, I don’t at all feel we should be anxious about “AI” in this respect.

However, that’s just about the quality of the content. It is also the case that content of mediocre quality is often still serviceable enough to enter circulation and count as meaningful work. What Nick alludes to then is an understandable worry: that this content will displace genuine work, of whatever quality, and distort incentives such that there’s no point in even trying. Just use a chatbot.

Better to commit to finding alternative sources of meaning than to aim to keep everybody sick because you derive meaning from attempting to cure them.

“AI” is not going to cure cancer. So many of these hypotheticals rely on exaggerated notions of AI’s current or future capabilities rather than grappling with the present reality: Mass-produced mediocrity that’s passable enough to enter circulation and clog up attentional resources and time.

I’m not concerned about losing “sources of meaning” because an “AI” does something I care about “better” than I can. Because it doesn’t. I am concerned about the effects of people believing that it can and does though, among other things. Because it’s that belief that leads to institutions encouraging or, in some cases, mandating, that employees use “AI” and the result is an information environment where the content of what is produced doesn’t even matter anymore.

You might think game-playing is trivial and meaningless, but I think if we intentionally thwart cancer-curing AI for the sake of preventing the obsolesce of human-directed cancer-research, we’re already treating cancer-research as a game, at least if you accept Suits’ definition (which strikes me as a pretty good one). Likewise, if we intentionally thwart understanding-enhancing philosobots for the sake of making possibile the activity of seeking-philosophical-understanding-without-AI, we’re treating philosophy as a game too.

I would invert this point in its entirety. Handing over our thinking to an “AI” is “treating philosophy as a game,” where the content of the information produced does not matter so much as the fact that it is produced and can pass serviceably into “the literature” as “work.”

We aren’t “thwart(ing) cancer-curing AI for the sake of preventing the obsolesce of human-directed cancer-research.” We’re “thwarting” it because the present reality is that it is thwarting that research by creating an overabundance of spurious “work” that human beings ultimately have to attend to and evaluate if they’re going to cure cancer.

Much of your comment is focused on the “loss of meaning” aspect, which is what you seem to have gotten out of Nick’s comment, taking for granted the (lack of) realized (or realizable) benefits. However, it’s that which should hold concern here, as we are far more likely to see lack of benefit (or even outright detriment, as we are seeing), than a scenario where we’re genuinely worried that “AI” has gotten to all the problems before us and left us without meaningful work to do.

Daniel Greco
Reply to  Felix
17 days ago

You’re “fighting the hypo”. The way this thread started was that I was engaged in hypothetical discussion of a scenario where we’re encountering papers that, but for their provenance, would warrant publication in philosophy journals.

What I’m getting from your comment is that you don’t think that’s happening. As I said in my initial comment, I also don’t think it’s happening. Yet. Maybe you think the hypothetical is so outlandish as to not worth be considering? Or you worry considering it will make it seem more realistic than it is, which might lead people to think it’s been realized even when it hasn’t?

I suspect we disagree on how realistic it is. I’m reading a bunch of accounts from mathematicians to the effect that AI is already producing results that absolutely warrant publication in mathematics journals. They sometimes go on to lament–or actively seeking to prevent–the loss of meaning that seems to threaten. Stuff like this, this, and this. Reactions vary but nobody is saying that AI-written mathematics is just mass-produced mediocrity. It isn’t mediocre, and to insist that it is would be denialist.

On another thread you seemed to be denying that we’re already there with mathematics. Did I have you right? Are you still denying it, or have the results since then–the field moves fast, and there are a lot more AI-produced counterexamples to significant conjectures since the previous post–convinced you? Or maybe you think the extrapolation from math to philosophy is shaky? (There, I agree; it may be that the standards for good philosophy are too squishy to train AI to produce comparably valuable, largely unsupervised work the way it can be done in mathematics. But I wouldn’t bet my life on it.)

Eli Alshanetsky
Reply to  Daniel Greco
17 days ago

Really thought-provoking comment. Here’s a case that pulls the other way for me: suppose AI hands us an interpretation of Kant or Spinoza. “Here’s what’s load-bearing. It’s not X, it’s Y — I’ve checked more of this than you have.” And suppose it’s right. Would we say “awesome, thanks” and be done?

Your forced-choice assumes that once AI can create the good output, what you do is more or less redundant. But creating a good new thing isn’t the only serious relation you can have to something good. Another one is: getting why it’s good and not messing it up — keeping it going through dark periods, when materials are sparse, when subtly worse but impressive things turn up left and right and tempt you to change it. In many cases that’s way harder than creating the thing itself.

The old story had Islamic scholars keeping Greek philosophy on ice until Europe was ready for it. Even now the defenses mostly work by showing they produced new work after all. Sure they did. But the real erasure might be in the idea that creating something is the only way of truly owning it, and every other relation to it is second-rate. Translation, criticism, adaptation, teaching it to someone who’s lost the thread can be just as important, or more.

I suspect that idea is what makes lots of people frightened of AI. Not AI, but AI against the backdrop of that idea. Most human cultures, I think, would have found that idea deeply strange. I tend to be with them.

Daniel Greco
Reply to  Eli Alshanetsky
17 days ago

I think I agree with everything you say here except the idea that this in tension with what I was saying. I love this:

But the real erasure might be in the idea that creating something is the only way of truly owning it, and every other relation to it is second-rate. Translation, criticism, adaptation, teaching it to someone who’s lost the thread can be just as important, or more.

And that seems to me broadly consistent with the kind of thing I was saying about what our relationship to philosophy might be like in the future I’m imagining where it really is consistently better than us at producing original work. Or maybe what I wrote was indeterminate, but in that case I’m happy to just say that this strikes me as a kind of healthy imaginative preparation for how humans can continue to derive meaning from philosophy in a future that may or may not come–but which I don’t think can be ruled out–where AI is better than humans at producing novel philosophical writing.

Eli Alshanetsky
Reply to  Daniel Greco
17 days ago

Awesome, I didn’t think we were far apart. I was pushing against the grim either/or (no meaning, or making our peace with games), not your later point about understanding.

I’d add one other thing. We also want to be able to check the AI. Outside of math and a few other domains where you can verify things mechanically, you can’t really evaluate an answer without comparing it against genuine alternatives. If all the alternatives come from AI, that’s a problem. We still have to teach people to come up with good possibilities of their own.

Either way, I think there’s a meaningful life available that’s neither competing with AI at novelty nor just playing inside a sandbox that it creates for us.

Eric
Eric
Reply to  Eli Alshanetsky
17 days ago

Terrence Tao has a very recent post (I forget where, probably X) where he basically says this is true even in math. You can use lean to check whether an AI proof is formally correct, but there is still a ton of work left in translating these proofs into a form humans can understand more easily. And then there is the even more difficult work of figuring out what it all means about the world of mathematics.

It’s hard for me to imagine a world where there is no role for humans to play (unlike the cancer case) precisely because humans are the direct consumers here (where I would call the direct consumers in the cancer case the tumors, or whatever). But that doesn’t mean that there isn’t a role for machines to play in generating perfectly good arguments and papers.

Eli Alshanetsky
Reply to  Eric
17 days ago

Absolutely, there’s a paper on exactly this issue by Thurston that I always loved:

https://www.math.toronto.edu/mccann/199/thurston.pdf

I feel like this is now more relevant than ever. There’s also the broader question of who, ultimately, controls the overall direction where math is going. Even if AI is super helpful in generating proofs and explaining them to us, that may still not be enough to ensure we end up where we want to be, mathematics-wise.

Daniel Greco
Reply to  Eli Alshanetsky
17 days ago

Yes good point, maybe there’s nothing really gamelike about what’s involved in “owning” (in the sense you explain in your comment) a work of philosophy that you didn’t create. Likewise with math. Trying to get yourself from the state of being confronted with a valid proof of some significant concept to internalizing and understanding it needn’t be a game; you might not be avoiding any shortcuts, and it’s just that can’t get the understanding by further prompting; you’ve got to do the work. So if achieving that sort of understanding, whether in math or philosophy, is meaningful–which I think it is–then it’s a way between the horns of the dilemma I was posing.

Alice
Alice
Reply to  Daniel Greco
17 days ago

good exchange. But the game-like aspect is not undermined by how serious the activity feels to yourself or how much effort it takes. A professional player takes their sport very seriously, even though it is just a game. Doing philosophy, understanding definitely *feels* not a game. But this is precisely what makes Greco’s game analogy a substantial claim. If one is receptive enough, rather than insisting on how mediocre AI currently seems to them, this picture should rightfully blow one’s mind.

Eli Alshanetsky
Reply to  Alice
17 days ago

Totally agree that games can be genuinely meaningful. But philosophy and mathematics aren’t games in Suits’ sense. Change the rules of chess and you get a different game; change math’s and you get an error. Say AI generated all the papers and got them right. Recognizing what actually matters, inhabiting it, and making sure a run of impressive-looking ideas doesn’t slowly take us somewhere we can’t retrace our steps from is a huge and necessary part of doing philosophy. And we can’t do any of it if no human is in a position to check the AI, which takes some independent capacity to come up with the ideas ourselves.

Daniel Groll
Daniel Groll
Reply to  Daniel Greco
17 days ago

Thank you all for such an interesting discussion! To further bolster the non-pessimistic projection (and echoing some of Eli and Eric’s points): the fact that the kind of philosophy AI might do will interest us only insofar as we can understand it (or, perhaps more broadly, appreciate it), suggests it is unlikely that using AI to do philosophy will get to the point where we are not meaningfully (that is, only game-playing-ly) doing philosophy. For: once we get to that point, the AI’s outputs will no longer be of interest to us.

More broadly — and perhaps this a very obvious point — we need to distinguish between activities where the benefit to us need not go through our understanding/appreciation and those where it does. In the former cae, we might have good reasons to see what AI can produce, even if what it produces elludes our understanding (because, say, it gives rise to a new treatment. Contra Eric, I wouldn’t say the tumor is the “consumer”. We are. But our consumption doesn’t need to go through our understanding). But in the latter case, it makes no sense to have the AI stretch beyond our capacities to understand what it produces. And so, that’s an argument for not worrying so much about letting the AI do philosophy, since our interest in having it do philosophy will always be tethered to our ability to menaningfully engage (via understanding) with what it produces.

I don’t particularly like the conclusion, but maybe that’s just conservatism on my part.

Eli Alshanetsky
Reply to  Daniel Groll
17 days ago

I’d make an even stronger claim. Even when the benefit doesn’t depend on our understanding (e.g. elevator safety), we still need people who can tell the difference between genuine safety and thirty years of apparent safety before an unforeseen condition brings every elevator down.

Eric
Eric
Reply to  Eli Alshanetsky
16 days ago

This is such a great subthread. Daniel Groll: I really appreciate your open-mindedness, especially after I came off as snarky.

Vilhelm Agdur
Vilhelm Agdur
Reply to  Daniel Groll
15 days ago

Staying in the example of mathematics, it’s worth keeping in mind that there isn’t a perfectly watertight division between “pure” mathematics, where we need human understanding of the results and integration of those results into our understanding, and applied mathematics, where the results can speak for themselves.

If an AI makes some progress in functional analysis that no human is truly able to appreciate, but that result is then used in another AI-based result to improve some optimization algorithm, and suddenly producing carbon-neutral concrete is economically feasible, does it matter that the pure mathematics wasn’t appreciated by human mathematicians?

The mathematics story is clearly plausible – the hope that our pure results will some day be useful to the applied mathematicians is the most common justification for why the public should fund mathematics departments, after all. Can we imagine a similar case in philosophy?

An AI philosopher develops a new better account of the semantics of counterfactuals and free will, which no human is able to fully appreciate, but which another AI applies to ethics, and eventually we have a clearer understanding of how to format welfare programs?

It sounds less plausible to my ears than the concrete mathematics example, but is that a problem with AI in particular or is it a problem with the picture of how “abstract” philosophy can be made useful at all?

(NB: As perhaps my choice and quality of examples showed, I’m a mathematician with a philosophy interest, not the other way around, which probably affects my view on the issue.)

Eric Steinhart
Reply to  Nick
17 days ago

You wrote: “Their scholarship is reduced to bot-prompting”. That’s both absurd and tragic. That thought is one of the biggest problems with the ways philosophers reply to the increasing power of AI.

If prompting is how philosophers think of using AI, then that reflects very badly on philosophers. We’re supposed to be experts on thinking – but prompting is all you can think of when it comes to using AI? Seriously? If so, you’ve already stopped thinking.

How about: philosophers start building AIs and training AIs. Or philosophers start using their skills to design AIs that demonstrate even greater cognitive excellence. Or philosophers starting using AI to reverse engineer philosophy, or to illuminate and correct specifically human biases in philosophy. There are ten thousand ways we can use AI to advance philosophy. Prompting is not among them.

If prompting is all we can think that philosophers can do with AI, then we’ve have already stopped thinking. At that point, the discipline of philosophy is no longer worth preserving.

Michel
Reply to  Eric Steinhart
17 days ago

Is anyone in this fever dream paying the real cost of compute?

Kenny Easwaran
Reply to  Michel
15 days ago

Yes. The price of compute has been falling drastically, and it’s quite plausible that it will continue to do so. A year ago it would not have been economically feasible for me to pay for an AI to write a customized data structure for me to store all the facts that go on my CV, in a form that is easy to customize to a shortened CV and also to the format of the .docx form the university makes me submit every three years for merit review, even though Anthropic and OpenAI were running big operating losses by not passing the costs on to customers. But this year, they’re basically breaking even on costs, and it only cost a small fraction of the compute included with my $20/month subscription to do all this.

There may eventually be obstacles to getting the cost of LLM-type compute down further, but it seems not to be slowing down just yet.

kjk
kjk
Reply to  Daniel Greco
17 days ago

AI as it now exists is not, and is not foreseeably ever likely to be, a morally or politically neutral resource. To sit back and muse about what publication norms would or will apply on the day when LLMs start emitting graduate-level philosophy papers is to ignore AI’s litany of dangers and harms, and its central placement in class warfare and the military industrial complex; by effectively treating such things as matters of course, such discourse collaborates with the interest groups driving AI growth. These groups want nothing more right now than for democratic populations to think of “AI-transformation” and “data center buildout” as a fait accompli to be welcomed or at least not interfered with.

Also, and forgive me for getting personal, but consider whether, if you weren’t so fortunate as to be able to home-school your kids, if instead you had no choice but to send them to public school where every day they were herded toward screens, if you lived in a lower-class neighborhood down the street from a brutal black metal monument streaming smog and ruining your peace outdoors with its whistle, if you were being coerced to pay higher utility costs to subsidize the thing’s grid suction, if you weren’t tenured and far removed from the threat of ever losing your current income level or your professional self-respect to LLM replacement, perhaps you might feel less bothered by the idea of denying space in academic journals to AI slop.

Eli Alshanetsky
Reply to  kjk
17 days ago

The class and environmental risks should be front and center. But I worry that “denying space in academic journals to AI slop” strengthens the actors you’re worried about. There’s no reliable way to verify, so the rule gets enforced by suspicion. An established name is protected by social trust. Someone with no reputation, writing in their second language, gets flagged, since detectors are famously bad at telling AI from non-native English, and what’s left is an editor’s sense of what human writing sounds like, which is mostly a sense of what insiders sound like. Meanwhile the corporate labs never needed peer review, which is one of the few places where a good idea can beat a funded one on its merits, and weakening it costs those companies nothing.

Last edited 17 days ago by Eli Alshanetsky
Kenny Easwaran
Reply to  kjk
15 days ago

Humans are notably not morally or politically neutral either. Humans have a litany of dangers and harms, and play very central roles in class warfare and the military-industrial complex. To ignore these features while asserting that they matter when they apply to something other than humans seems like missing the point.

We shouldn’t pretend that these systems, and our friends and colleagues, are morally or politically neutral, and we shouldn’t pretend that they aren’t guilty of harms. We should acknowledge that, and try to make them better.

Jessie Ewesmont
Jessie Ewesmont
19 days ago

Reject, with the comment that it’s AI generated. Easiest review ever.

Matt L
19 days ago

as well as the possibility that the AI use falls within the journal’s rules and was disclosed during the submission process.

I am curious about what the idea is here. It seems to me that this sort of disclosure should also be made available to the referees. Maybe that’s wrong, but I don’t think it’s obviously wrong, and I’d be interested to hear what people have to say about it.

Eric
Eric
Reply to  Matt L
19 days ago

Maybe say first why it should? Isnt the referees job to determine the quality and originality of the paper, and that seens orthogonal.

Matt L
Reply to  Eric
19 days ago

I suppose it might depend on if I think authors are being candid on their declarations. (We have students make such declarations, and I’m sure many are not candid.) But, if they were, and they explained how AI was used, it would help me come to a better conclusion on the paper, I think. Maybe that’s wrong, but it doesn’t seem obviously so to me.

Eric
Eric
Reply to  Matt L
19 days ago

you haven’t said a single word about why it should help you come to a better conclusion on the paper. you added some irrelevant stuff about whether they are candid. assume they are.

I understand the view that AI writing is slop. Maybe most is. Maybe all is. I don’t understand the view that you should be refereeing a paper that you can’t use your own judgment to decide if its good or not–and that you need to know facts about what tools were used to create it.

Matt L
Reply to  Eric
18 days ago

Okay. I’ll say a bit more. Here are a few things. I am sometimes asked to review things that are not squarely in my wheel-house. How much time I spend looking up claims, or how deferential I am to an author’s citations, can be influenced by if they disclose they used AI in a literature search or in research. Similarly, if I am wondering why something is written in a particular way, and what I might say about changing it, whether AI was used to help with working in a different language can help as well. (I have lots of experience with these things from working with student papers, among other things.) Now, I do expect lots of people to lie about this stuff, sadly enough. That’s one reason why I expect AI declarations to not be much use.
Another way is that if someone says no AI was used, and then I get a couple of obviously made-up citations, I’ll know that there is simply no point in reading further – that the person has not only wasted my time, but has been deceitful about it.

I don’t expect everyone to think those are the right paths. And, as I noted, I expect a lot of people will simply lie. But, if there are going to be such declarations, I think they should be made avaiable to the referee.
(For what its worth, I think that at least some conflict of interest statements should also be made available.)

Eric
Eric
Reply to  Matt L
17 days ago

I don’t think you should be refereeing such a paper any more than I think you should be grading a paper if you need to know whether a male or female student wrote it to grade it.

Matt L
Reply to  Eric
17 days ago

if you need to know whether a male or female student wrote it to grade it.

I’ll admit that I have no idea where you got this from.

Eric
Eric
Reply to  Matt L
17 days ago

Then you’ve never read the most important paper in the history of AI and the philosophy of AI.

But my point is: if you are too out of your wheelhouse to evaluate a paper without knowing irrelevant facts about its provenance, decline to review it. I don’t review papers I don’t know the literature well enough to evaluate them on their own merits.

Matt L
Reply to  Eric
16 days ago

Then you’ve never read the most important paper in the history of AI and the philosophy of AI.

You’re probably right there. You’re going to need to say more to me here.

my point is: if you are too out of your wheelhouse to evaluate a paper without knowing irrelevant facts about its provenance, decline to review it. I don’t review papers I don’t know the literature well enough to evaluate them on their own merits.

In these cases, I usually tell the editor that this isn’t a core area of mine (it’s often something I have written a bit on, but it’s not a main area) and ask them if they still want me to do it. They almost always say yes, because it’s very hard to get referees. (I’m an editor myself, so I know.) If it’s really an area that’s too distant for me, I say no, but I rarely get such requests. So, I think you’re still off here.

Eric
Eric
Reply to  Matt L
16 days ago

In his classic paper on machine intelligence, Alan Turing defines the immitation game by first giving the example of whether a person can tell a man from a woman by what they type.

Matt L
Reply to  Eric
15 days ago

Well, I have read that paper. I assumed you must mean something else, because it’s pretty clearly not relevant to this discussion. Too bad! I thought this might be helpful.

J.P. Loo
19 days ago

If Pangram works, that solves one problem (how to decide whether it’s AI-generated),* though not others (principally, what to do about it). I suppose I’m quite undecided about that. In certain areas (e.g. philosophical logic), we esteem largely formal papers of a kind that would be no worse off for having been written by an AI. I’m not sure we should think the same of e.g. phenomenology. (And a final complication is that I suppose in principle we might one day create AI systems that are saliently autonomous, conscious, or what have you, and so could write phenomenology papers just as interesting as the best written by humans.)

* Although there are some good outstanding criticisms of Pangram, one interesting point is that you can just email them with a half-baked idea and (n=1) they’ll give you 100,000 tokens to experiment.

Jessie Ewesmont
Jessie Ewesmont
Reply to  J.P. Loo
19 days ago

I recently read a blog post where the author was experimenting with Pangram. When he fed it a full essay (which he manually wrote), it came back 100% human generated. When he fed it a fragment of that same essay, it came back 100% AI generated. When he fed it a fragment of that fragment, it came back 100% human generated again. This is a pretty odd result, and it makes me distrust Pangram out of worry that it’s inconsistent.

Rob Hughes
Rob Hughes
19 days ago

If a paper contains fabricated references, serious misrepresentation of sources, or blatantly false statements of fact, that is reason enough to reject the paper without further comment.

If a paper doesn’t contain fabrication but is bad, and it appears LLM-generated, I would write up a review. I would try to keep it short. I might comment about my suspicions in the confidential letter to the editor, but I would reject the paper on grounds of poor quality.

If a paper appears to have merit, but it also appears to violate the journal’s policy on LLM use, I would reach out to the editor.

I think there is little chance of an entirely LLM-generated paper in philosophical value theory being good enough to deserve publication. I think there is a high chance a journal could receive an interesting paper that is mostly human-written but has some LLM editing or additions that violate the journal’s (legitimate) policy on LLM use.

It is unethical to submit someone else’s unpublished manuscript to an “AI detector” without their explicit consent. It is unethical to use “AI” systems to review papers.

a nonnative english speaker
a nonnative english speaker
18 days ago

AI tools help nonnative speakers to make their sentences more natural. They also help them to correct grammatical errors. Now, the problem is that using them for these purposes sometimes makes the prose look AI-generated. But that does not mean that AI actually wrote the paper.

Gorm
Gorm
Reply to  a nonnative english speaker
18 days ago

But the prose look AI generated because they were AI generated. That is a problem

Matthew Braham
Matthew Braham
18 days ago

See my recent post on the Meta-Epistemological Reasons thread: its the case of mathematicians checking an AI generated result on the 87-year-old Jacobian conjecture: it was recently disproved by Claude’s Fable model and amusingly while the prompting mathematician was watching the World Cup final. The discussion is very instructive.

Anya
Anya
Reply to  Matthew Braham
17 days ago

I think the jacobian disproof is indeed telling but largely because it is a disproof. Philosophers are really good at coming up with counterexamples, but those are only one part of what we do and they are I think the easier part by far. Some really significant progress has come from them (e.g. Frankfurt cases vis a vis the principle of alternate possibilities springs to mind), of course. But I wouldn’t think AI will render us obsolete by being better than us at counterexamples. Indeed I think that might generate more fruitful theorizing (as Frankfurt cases maybe did) that it can’t do). But of course this is speculative and short-to-medium-term.

Matthew Braham
Matthew Braham
Reply to  Anya
17 days ago

There is some misunderstanding here: I am not indicating obsolesce but rather the need to reconsider our attitude and there maybe something to learn from how mathematicians go about their business. Counter-examples are both disproof but also interesting stimulators of new thought on a topic — the Frankfurt case as you state. So its quite possible AI could do this.

David Sobel
David Sobel
17 days ago

I am quite curious about a related, but not exactly on topic question. How good is AI at writing a philosophy paper at this point? Can anyone say that they read a paper that they are sure was (mostly) written by AI that was otherwise good enough to be accepted into a legit journal?

Philipp Stehr
Reply to  David Sobel
16 days ago

Yascha Mounk claimed a while ago that he was able to basically generate a good political theory paper from scratch: https://writing.yaschamounk.com/p/the-humanities-are-about-to-be-automated

On a first look, the paper isn’t great, but I can see it make it past peer review at a low-mid tier journal in political theory.

Lynette Reid
Lynette Reid
Reply to  Philipp Stehr
9 days ago

As a journal editor, I respond to papers that read like this (i.e. whose authors are probably using AI in ways they didn’t disclose) that negative parallels do not constitute arguments.

Eric Steinhart
Reply to  David Sobel
16 days ago

Your question is exactly the one we ought to try to answer. (Right now, nobody knows — people have opinions, but there’s no evidence.) I think the APA ought to try to answer this with an experiment. A sort of Turing test for philosophy papers.

Brian Weatherson
Brian Weatherson
Reply to  David Sobel
15 days ago

I think it depends what we mean by ‘paper’. With a bit of work I’m pretty sure I could get one to write a 1000-1500 word discussion note that was ok journal quality. A full length paper though seems further away. As others have noted, it’s really hard to get the models to make the kind of long range connections you expect to see in an 8000-10000 word paper.

Jukka
Jukka
16 days ago

As this problem is what I am increasingly facing, I think particularly the third point was a good advice.

But I’d add a point (5): write in the review that, if necessary, you can continue a review in a subsequent review round with new points. That is a reasonable position that also editors should understand because no one should be compelled to carry out a thorough review of slop.

In addition, I think submission systems should nowadays have a form for delivering feedback to editors about low-quality and slop peer reviews. As it stands, many editors do not seem to see it themselves.