Illicit Use of AI by Philosophers Refereeing for Journals
In 2024, a study found that “7–17% of the sentences in the reviews [of computer science manuscripts] were written by LLMs”. It was only a matter of time before this spread, and now it appears to have reached philosophy.

Last year, a philosophy PhD student in the US submitted a paper to a well-known philosophy journal.
They write:
The paper was rejected a few months ago; the first reviewer left very detailed feedback and suggested substantial R&R while seeming generally positive about the paper, the second reviewer suggested rejection. At first, I appreciated both of their feedback, and this was sort of my first go at submitting for publication anyway.
However, I recently showed this feedback to someone else who thought the first reviewer (the more positive reviewer!) sounded like AI. [This] hadn’t occurred to me when I first received it, but it now seems very clear that it was AI-written. Although I know they aren’t the most reliable tools, I also checked with a couple of AI detectors and they come back as highly confident the text is 100% AI generated.
I believe this is the first time someone has written to me about this happening in philosophy, which means it is almost certainly not the first time it has happened.
Has it happened to you? (This is philosophy, so start by checking the reports that seemed relatively nice.)
It’s worth discussing. We can start with why, as things stand now, if you are asked to referee a submission for a journal, it would probably be wrong for you to use AI’s like ChatGPT, Claude, Gemini, etc., in doing so (except, sometimes, in a very limited capacity). Here is why:
1. Let’s start with you feeding the manuscript into an AI. One problem here is that you have no reason to believe that the author of the manuscript agreed to the manuscript becoming training data for an AI, or for whichever AI you happen to use. Given that there is well-known controversy over this, you shouldn’t assume the people you are working with would agree it’s okay.
When you agree to referee for a journal, you agree to abide by the publisher’s review policies, and various academic publishers explicitly prohibit in their reviewer guidelines uploading a manuscript under review into an AI. For example, Oxford Academic’s policy for reviewers states:
It is prohibited to upload project proposals and manuscripts, in part or in whole, into a Gen AI tool for any purpose. Doing so may violate copyright, confidentiality, privacy, and data security obligations.
2. Now let’s turn to the assessment of the submission and the writing of the report. The journal editor asked a specific person to review a manuscript: you. If you accept their invitation to review the paper, you agree to review it. If you were then to give the manuscript to one of your colleagues or graduate students to review, you would not be doing what you agreed to do. If you did this and then didn’t tell the editor of the journal for which you’re reviewing, then you’re fraudulently submitting another’s work under your own name. The fraud is still there when it’s not a colleague or student but an AI to which you’ve handed over the task.
Again, when you agree to referee for a journal, you agree to abide by the publisher’s review policies, and some publishers outright prohibit the use of AI-written referee reports. Elsevier’s policy, for example, states:
Generative AI or AI-assisted technologies should not be used by reviewers to assist in the scientific review of a paper.
An additional concern is that authors typically submit their work for peer review under the reasonable expectation that, if the work is of sufficient quality, they may receive constructive feedback on it from fellow experts. The use of AIs to write referee reports not only fails to meet this expectation, but does so at a cost. As the graduate student who wrote to me put it, “I find it very frustrating because if I wanted feedback from an LLM I would just ask it, instead of waiting months for feedback from a journal just to send me AI-written feedback.”
3. Even more limited use of AI that doesn’t involve uploading the submission into an AI or having the AI generate comments on it can be problematic. It might be acceptable for a referee who has read a manuscript and written up their report to then feed the report into an AI for rewriting (say, for tone, clarity, grammar, translation), check the rewritten version for accuracy, and then submit that AI-written version. But even the permissibility of this varies across publishers. Elsevier, for example, prohibits it:
This confidentiality requirement [restricting the uploading of the submission to an AI] extends to the peer review report, as it may contain confidential information about the manuscript and/or the authors. For this reason, reviewers should not upload their peer review report into an AI tool, even if it is just for the purpose of improving language and readability.
Taylor & Francis, meanwhile, states in their policy that it may be permissible:
Generative AI may only be utilised to assist with improving review language, but peer reviewers will at all times remain responsible for ensuring the accuracy and integrity of their reviews.
Note that individual journals may have more restrictive policies.
At some point, as the technology improves, our norms, policies, practices, and expectations may change. And some changes may be welcome, as our current academic publishing system certainly has its problems. But for now, if you’ve agreed to referee a paper, don’t try handing off that work to an AI.
Discussion is welcome. One thing to keep in mind, philosofriends, is that this is largely a matter of policy. So while we might imagine various “in principle” examples of the permissible use of AI in refereeing, these may be of limited value in figuring out which institutional rules we should adopt for imperfectly influencing the behavior of many people who differ in many ways.
Chatbots are not my peers. End of story.
I am completely aghast and dumbstruck. I wondered for a moment whether this was an April fool’s joke, albeit, a bit late for that. How could anyone possibly result to such a debasement of the integrity of the peer review process by having a so called “AI” (though I question the ‘intelligence’ part of the “I”) review a paper. This is an absolute defilement of intellectual integrity and a shameful state of affairs that we are discussing this here today. I can understand an outsider to philosophy or to academia (the business type, say) looking for a ‘time saver’, as it were, but for a philosopher to look themselves in the mirror after defiling the peer review process in this way is beyond me.
Sure, but it seems like this moralized take on peer review misses that the majority of reviewers are overworked academics with almost no incentive to write reviews. The amount of uncompensated labor that we are asked to do as academics is insane and we should not be suprised that people are refusing to spend time doing it.
Also, this response assumes that human reviewers are putting thoughful time into producing reviews focussed entirely in truth-seeking through collective inquiry. This is obviously false. Many reveiwers barely read the paper and are often motivated by their own gatekeeping biases that trade more in advancing their conception of what philosophy papers should look like. It is unclear that the process is some sacred practice as you describe it.
We miss the point, and risk strengthening labor inequality in academia, when we pretend that peer review is some holy institution.
I agree reviews are uncompensated labor. If you think it is too much work, you should just say no rather than using an AI. I also agree some reviewers do a sloppy job even if they use AI, and that peer review is not a sacred and unimpeachable process. Still, it strikes me that the flaws of peer review mean we ought to try to make it better in whatever ways we can – by not using AI, and doing our best to set our bias to the side – instead of treating it unseriously. As flawed as it may be, it still determines who gets published, and therefore who gets hired and who gets tenure. These are weighty stakes, so we should do our best as peer reviewers.
Peers are wrong all the time. Peers do a bad job all the time. But they are, at least, peers. A chatbot is not a peer.
I teach 8-11 courses a year. I also referee 20+ items a year. I’m not superhuman, and neither is the effort involved. Not everyone need do so much; but if someone hasn’t got the time, then a simple ‘no’ is preferable to offloading the task onto an RA, a graduate student, an undergraduate, a random stranger, or a chatbot.
It should be obvious that peer review is part of the suite of professional duties that make up an academic’s professional life, a part of the system that generates our salaries. Treating it as uncompensated labor is among the delusions of those academics that think of themselves as part of the working, rather than the ruling class. Sure, if you’re an adjunct, no incentive to review. But also no expectation that you will.
I do not think it is at all obvious – at least for all academics. This very much depends on the academic’s job description, teaching load, etc. Perhaps for Chicago profs it is, but I certainly would not generalize across all philosophers in academia
I know that’s not your point but even though academics are underpaid, I don’t think people with PhDs are part of the working class. As for uncompensated labor, the point is that refereeing counts as service to the profession, which is part of the job for at least most tenure stream academics.
So just to get this straight if someone with a PhD ends up driving a school bus or working as a plumber they are not working class but on the other maybe Elon Musk and Zuckerberg are since they don’t have college degrees? If only those privileged adjunct bourgeoisie would stop oppressing working class folks like Zuck and Musk.
It’s like someone didn’t read their Marx
I’m sympathetic to some of your gripes but the basic argument here is something like “If important institution X is in bad shape or flawed then we shouldn’t complain about any changes that make it worse.” I mean I see this style of argument a lot but it’s just mad. And I can moralize peer review in that I think it’s important and *should* be done well while admitting it’s very often not. (I’ve got plenty of reviewer 2 stories myself).
Here is a sound argument. There are two premises. One states an obligation; the other is an empirical claim. The conclusion is that in a frequent subset of cases where one agrees to peer review an article, one ought to employ an AI.
Premise 1 (The Obligation Principle): In every case where one agrees to peer review an article, one ought to employ the most reliable means available for identifying errors in that article.
Premise 2 (The Empirical Claim): In a frequent subset of cases, employing an AI constitutes the most reliable means available for identifying errors in an article.
Conclusion: Therefore, in a frequent subset of cases where one agrees to peer review an article, one ought to employ an AI.
I direct your attention to the final paragraph of my post, where I write:
One thing to keep in mind, philosofriends, is that this is largely a matter of policy. So while we might imagine various “in principle” examples of the permissible use of AI in refereeing, these may be of limited value in figuring out which institutional rules we should adopt for imperfectly influencing the behavior of many people who differ in many ways.
Fair enough Justin. I have revised accordingly. Can we agree that this revised argument is now sound?
Urmason’s Argument (Revised):
Premise 1* (The Obligation Principle): In every case where one agrees to peer review an article, one ought to employ the most reliable journal-compliant means available for identifying errors in that article.
Premise 2* (The Empirical Claim): In a subset of cases, employing an AI constitutes the most reliable journal-compliant means available for identifying errors in an article.
Conclusion*: Therefore, in a subset of cases where one agrees to peer review an article, one ought to employ an AI.
This still feels like it doesn’t address the policy question. Compare this argument
That’s clearly a bad argument. The right *policy* might tell people to do something suboptimal in some cases, because if we set the policy for unusual cases like this, too many people who should be following the sensible policy would not.
It’s at least coherent to believe
A. Some reviews would be improved by AI. (As you say below, that’s not plausible if AI use is mindless cut-and-pasting, but it is a bit more plausible if it involves using the AI as a double check.)
B. If journals allowed AI, the average review quality would go down (perhaps because too many people would slip from ‘double check’ to ‘just write it for me’.)
That’s one reason the policy question is central.
(The confidentiality question is separate, but I kind of think that’s orthogonal to the AI question. Using cloud storage to move a review file between devices raises confidentiality risks, using an on-device AI does not.)
Thanks Brian. So that’s a fair enough point about the distinction between individual optimal actions and overarching policy, but I think my revised argument actually already accounts for this. By specifying that the means must be ‘journal-compliant’ in Premise 1, the argument defers to the journal’s policy. So if a journal implements a blanket ban on AI (like your speed limit), then AI is no longer a ‘journal-compliant means,’ and the obligation to use it dissolves. However, if we are pivoting to debate what that journal policy should be, I think the speed limit analogy at that point breaks down. A better analogy I think is banning calculators in advanced engineering exams because some students might rely on them instead of thinking. If a reviewer is lazy enough to mindlessly copy-paste an AI output, they were likely going to write a poor, perfunctory review anyway. Editors already serve as a quality-control backstop against bad reviews. Implementing a blanket, unenforceable ban on AI doesn’t stop lazy reviewers but really just serves to prohibit conscientious reviewers from using a powerful diagnostic ‘double check’ to catch errors they might otherwise miss, and this is something that will ultimately harm the peer review process.
I hope this was true. If science in general is considered, many editors nowadays merely “count votes” via the buttons provided by manuscript submission systems. Many of them also send AI slop for peer review, making it a vicious cycle.
Both premises are wrong. Premise 2 is wrong because of the well known propensity of AI to hallucinate, especially relative to the well known propensity of reviewers to pick apart even subtle errors in a manuscript. (That’s relevant, since “most reliable means” is comparative.)
The first premise is also wrong. Suppose I am in a department with many other intelligent and astute senior philosophers. It might be that they’re more skilled at identifying errors than me. But it wouldn’t be permissible, let alone obligatory, to ask them to give comments on the manuscript that I simply copy and paste in my review. The editor has asked me to do the review, not them.
I think both of these critiques rest on treating AI as an autonomous substitute rather than an assistive tool. Regarding the second premise, it is a false dichotomy to pit human expertise against AI hallucinations; the “most reliable means” is not human or AI, but human plus AI. While AI can hallucinate, human reviewers are subject to fatigue and cognitive bias, so using an LLM to flag potential issues for the human to independently verify effectively mitigates the weaknesses of both. Regarding the first premise, the analogy of passing a paper to a senior colleague fails because it conflates a moral agent with a non-sentient tool. Giving a manuscript to a colleague violates confidentiality (obviously) and entirely outsources the intellectual judgment the editor explicitly asked of you. Conversely, using an AI as a diagnostic aid is functionally no different than using a calculator to check a paper’s math or a search engine to check for plagiarism; you are still applying your own final, synthesized judgment in such a way as to be fulfilling your obligation to the editor while utilizing the best available tools to do so.
I think something much stronger than Premise 2 is actually true. In some subset of cases, a blanket policy of saying “there are no errors” may constitute the most reliable available method for identifying errors in a manuscript. These might also be cases in which a blanket policy of saying “there is a serious error on page 7 that makes this article unpublishable) may also constitute an equally most reliable available method for identifying errors in a manuscript.
These are, of course, cases in which a person has overcommitted themself and isn’t going to be able to do anything at all reliable. These are not good cases, but we know they exist. In these cases, despite the frequency of hallucinations, an LLM is likely to do better than either of these policies.
But I think the failure of Premise 1 still means that one shouldn’t use an LLM in these cases. (At least, certainly not one run on the cloud from a for-profit company.)
Looks valid to me. But sound? I, for one, am not sure either premise is true.
Meanwhile in the real world one of the main tells for when my students use AI on one assignment is that the robot goes on and on about how important fairness and integrity are for act utilitarians. And on another it loves to bullshit about how important the virtues of patience and authenticity were for Aristotle. Now feel free to tell me how this doesn’t happen with the $500 dollar version that I’m never going to buy and so I’m entirely unsuited to impugn the glory and wisdom of Claud, Chat GPT, or whatever.
I think there are several problems with premise 1, but I want to focus on premise 2. This premise doesn’t align with my experience at all, albeit I don’t use AI much so I have limited data to work from.
Just as one piece of anecdata, I was curious about six weeks ago how good LLM systems have gotten at philosophy. I fed a rough, ~5000-word first draft of a paper to mine to Claude Opus 4.6 (with “thinking” enabled). I knew I needed to add a whole section positioning my paper within a major literature I hadn’t addressed at all, and that was the main weakness of the draft. I asked Opus 4.6 to evaluate the draft. It missed that problem entirely, raised a couple of trivial worries (one of which was based on a misunderstanding), and then declared the paper already publishable pending minor revisions. It was, effectively, useless. Since then, I’ve realized it also failed to point out that an inference I drew in handling an objection didn’t follow without a questionable assumption about which there’s a sizable literature, which I learned only after sending the paper to a peer for feedback. I’ve had to majorly rethink the paper accordingly.
Suffice to say, I wasn’t impressed with Opus 4.6.
Of course, this isn’t proof that there does not exist a sizable set of cases in which Claude or some other LLM provides more reliable feedback than any other means. I’m just flagging doubt about it.
right, so your experience with Claude isn’t something I’m going to dispute — but I think it actually highlights a misunderstanding of how Premise 2 works in practice. You tested the AI as an autonomous senior colleague by giving it an open-ended prompt (‘evaluate this draft’). Current LLMs are undeniably not that great yet at macro-level philosophical positioning (though Opus 4.7 is much better as is GPT 5.5) and often default to polite, surface-level nitpicks when given a zero-shot prompt like the one you gave. However, notice that Premise 2 doesn’t claim AI is the best means for every type of error detection; it just claims it is the best means in a subset of cases. The ‘most reliable means’ of employing AI isn’t asking it to review the paper entirely, but using it as a targeted diagnostic tool. So if you instead prompt an LLM to execute narrow, specific tasks (say eg like ‘extract the premises in section 3 and check for formal validity,’ or ‘read my defense against Objection X and generate the strongest possible counter-argument’, it frequently catches subtle and tricky stuff that human fatigue might miss. The point here being that the failure of an LLM to independently identify a missing literature gap doesn’t disprove its unmatched reliability when properly directed at the subset of tasks it actually excels at
Do you have an example (or multiple examples) of an LLM catching subtle and tricky stuff that human fatigue might miss? Like others I have not yet found good use cases for LLMs but I am open to the possibility I am using them wrong. Actual examples would help me evaluate claims like yours.
When I was recently revising a Stanford Encyclopedia entry, I ran it through Claude for copyediting. (Since this was (a) my work, and (b) already on the open web, there aren’t any confidentiality issues here.)
As well as having a whole bunch of useful suggestions for how to clarify things, it noted that one of the cross-references was simply wrong. I said I’d discussed something earlier in section 3, but it was really section 4. (Or vice versa, I can’t remember which.)
It’s not the biggest deal, but (a) I’d missed it on previous drafts and so had SEP reviewers/editors, and (b) I’d expected it to be the kind of thing the LLM was *bad* at, since it normally does better at short-range things than long-range things. And it did require more than pattern matching; it required connecting the content of the referring sentence to the content of the earlier sections.
Does this matter for journal refereeing? Maybe not. But it is the kind of fussy detail that the LLMs are already getting to be better than many humans at catching.
At best I’d call that subtle but not tricky. Subtle is even pushing it – I think most people’s eyes glaze over when they see a section # unless they are looking for that information specifically. So it’s not surprising to me that lots of people would miss it, and I’m not sure I’d characterize their reason for missing it as one that depends on its subtlety. But in any case that doesn’t sound tricky to me.
I have tried similar things with a few of my papers. So far I also haven’t gotten good results. Maybe there are better sets of instructions that would make it work better. But someone who is turning to an LLM to save time with their refereeing probably isn’t spending the time to carefully craft a usable set of instructions for an LLM to effectively review a paper.
It’s much more plausible though that someone could jot down a bunch of notes while reading a paper, and an overall verdict, and then usefully have an LLM turn that into the paragraphs that an author can usefully respond to.
Premise 1 is false. When I’m asked to do peer review, I’m consulted as an expert. If it’s best to use some method other than my own expertise, then I ought to decline to do peer review and recommend that the editor do the other thing instead. For example, I might decide that I really don’t know anything about this topic but I know someone who does. In that case, I’d recommend a different reviewer. It would be wrong for me to take the assignment and that farm out the work to my more qualified colleague— or, at least, wrong to do so without OKing it with the editor beforehand.
Interesting line.. so you are absolutely right that an editor is consulting you for your specific expertise and it would be unethical to farm that out to a colleague even if they’re a smart one. But I think this objection conflates outsourcing your expertise with augmenting it. So when you say, ‘If it’s best to use some method other than my own expertise, I ought to decline,’ you are assuming that using an AI replaces your expert judgment. But the ‘most reliable means’ advocated in Premise 1 is actually ‘Expertise + AI,’ not AI alone. Consider an analogy here, so if a paper relies on a complex formal logic proof or statistical data, etc you wouldn’t decline the review simply because logic-checking software or a calculator is a ‘method’ more reliable at math than your unassisted brain — you’d use the tool to verify the mechanics, and then apply your philosophical expertise to evaluate the implications. My take here is that using an LLM as a diagnostic tool (to check for structural consistency etc) is functionally identical. You aren’t farming out the work to an agent; you are using the best available tools to ensure your expert evaluation is as rigorous as possible, which is exactly what the editor is hoping you will do
I’ve received one referee report that I’m 95% sure was AI generated. Honestly I’m just happy it has only been one so far…
I would rather get two flat rejections from human reviewers than two AI-generated unconditional acceptances.
We’d better let the Philosophical Review know about this preference promptly.
R R Urmason,
This rather nasty response to someone’s statement is uncalled for and, to my mind, downright rude. Have you any basis for this sarcastic implied-assessment of the merits of the prior commenter’s work?
I thought the joke was that most people’s experience submitting to Phil Review consists in getting two flat rejections from human reviewers.
Lighten up, Adjunct, it’s just a bit of banter
Imagine the possibility where journals don’t send through the comments when rejecting papers.
I agree completely with the conclusion that reviewers should not use AI and should not paste into chatbots the manuscripts they’ve been asked to review. A really nit-picky point regarding #1: it seems worth clarifying that you can easily opt out of having your conversations used as training data. E.g., in Claude, click on your name and then “Privacy” and then uncheck “Help Improve Claude.”
Also, the AI detectors are extremely unreliable. I treat them as providing practically no information. I fed in a manuscript of mine (which of course I did not use AI to write at all) and several detectors were extremely confident that it was AI-generated (think 80-90%). So although I believe that some reviewers in philosophy have used chatbots by now, the evidence cited (based on AI detectors) seems to me very weak.
One other related nitpicky point – there are various open-source LLMs that people can run locally (Mistral, Gemma, LLaMa, Qwen, Kimi, etc.) and if you run it completely locally, everything definitely stays private. Also, many universities have set up similar internal systems, sometimes even using versions of relatively high-end models from Anthropic, OpenAI, and Google, that again avoid the privacy/data security objection.
They still aren’t going to do a great job if you give them the job as a whole, but they might be usable for turning notes for a review into an actual review, or for providing a second opinion about whether there are any relevant points you missed while reviewing.
Yeah, it’s actually alarming how many colleagues and admin think these are reliable tools.
It is never a good sign, if you take it as an indicator of the expertise with which the topic is approached, that the assumption that anything that’s processed by a model becomes its training data is made unquestionably.
If you trust these companies to do what they say they are going to do, so much so that you are willing to upload someone’s unpublished work which you are very much not supposed to share with anyone, then I think you have undue trust in these companies. Specifically, you are trusting their security procedures (even though every day it seems like people are using LLMs of all things to poke holes in companies all over the world!), their technical acumen (are they really not hanging on to the data? Or if they are, are they really hiving it off from their training?), and their honesty and commitment to stick to promises (don’t even get me started…).
I have never received a review that I felt was written by an LLM, but there was one time when I served as a reviewer and was sent the other reviews the paper received. One of them was unambiguously written by an LLM. The review had literally nothing to do with the paper. It was about a completely different topic, but appeared to have been generated based on just the title of the paper. It was so strange and definitely very disheartening.
I hope that everyone will agree that this is wrong, but even if most everyone does, this problem is only just beginning. As many others have said before, we’re careening toward a vicious cycle: referees are already overworked, and given the competitiveness of the discipline and the imperative to publish or perish, it’s only a matter of time before AI-driven research output increases dramatically, thus overworking referees even more. I don’t know the solution, but I doubt that clear policies around AI will help unless the whole incentive structure changes.
Consider this a “modest proposal:” reputable journals should get together and form a loose consortium and agree to share both author and reviewer data. Any author or reviewer suspected of using AI to write either their papers or reviews *in a way that violates the journal’s policy* should be given an opportunity to explain their use. If, by majority vote of a journal’s editors, the author is found to have violated their AI policy, then they should be banned from submitting any material to any of the journals in the consortium.
This would have the secondary benefit of more clearly marking predatory/scam journals.
That would change part of the incentive structure, so it could work in that respect. But there’d be a risk of false positives with disastrous consequences for the author, and just as clever students quickly learned how to cover their tracks, so too would authors and referees.
If an author (and we’re talking about people either in PhD programs or people with PhDs and year – if not decades – of experience) can’t convince a majority of the editors of a journal that they actually wrote something then maybe we kill two birds with one stone here and we help to reduce the overall number of bad submissions (human and AI).
On the other hand, our solutions have to be responsive to changing conditions so if we get to the point where human and AI submissions are indistinguishable then, as I’ve suggested in other posts before, maybe it’s time for us to reconsider “the paper” as the best measure of philosophical activity. I’ve suggested that “the talk” might return to prominence in such an environment and, to be honest, I think it would be better for the discipline if talks were given more weight than they are now. Maybe it’s win-win.
So a weakness in this proposal is that it makes it rational to never review another paper, as doing so comes with a non-negligible chance of never being able to submit a paper again; this bad result could arise, following a given reviewing task you perform, via either a triggering a true positive (if you use AI to review and they catch it) or a false positive (if you don’t but they thought you did) and where banishment from submission is a known possible consequence of either.
It would be rational according to a narrow sense of ‘rational,’ according to which it’s already rational for most people never to review another paper (because you get no real benefit for reviewing, or at least no benefit that outweighs the cost of reviewing). Practically nobody in the profession is this kind of ‘rational,’ so your point is neither here nor there.
If you think that the proposal would in fact cause too many people to stop reviewing, you can say that, but it’s needlessly obfuscatory to make this out to be a point about rationality. Anything is rational in this sense, given certain preferences. The question is what preferences people in fact have.
It’s not clear to me that many people have preferences such that the proposal would change things. Given the increase in reviewers needed due to the onslaught of AI-written papers (I’ve reviewed one already, and lord knows how many editors have caught before sending them out for review) I think there’s a good case to be made that we’d get more genuine papers reviewed if AI users were cut out from the system than if we don’t do anything. It at least seems to me to be an open question.
But then what’s the incentive to have an AI review it? As the begrudged PhD mentioned, the author could do it themselves. Alternatively, it could be a feature of submission pages
Nothing says rigorous peer review like a referee pasting your paper into an AI and waiting for wisdom to rain down. Suddenly every report has the same tone: polite, vaguely impressed, and deeply confused about your equations. “Consider expanding the discussion” has never felt so algorithmic. I imagine a secret Slack where referees compare prompts instead of insights. Somewhere, two chatbots are debating your methodology while their humans sip coffee. Honestly, if the AI accepts the paper, I’m tempted to cite it as Reviewer 3. At least it responds faster and doesn’t hold grudges from conferences in 2014, either way.
Here’s my solution. Reviewers have to be open about AI usage. Authors are then told about the AI usage and are given a chance to respond.
Then the author can just say, “Actually, you got this completely wrong.”
And the AI will be like, “What an astute observation! You are absolutely correct. Your paper should be published immediately!”
While my first reaction was certain outrage, I quickly remembered that I have had a large number of presumably human “reviewers” hallucinate and attribute to my papers outrageous claims which I never made, rejecting my actual document on the basis of a straw man which they invented. And I have sometimes gotten better feedback from chatgbt about, e.g., which authors to read to supplement a budding argument sketch. I think that there may be *some* journals which might benefit from having AI feedback, which might be somewhat more objective and neutral than some human reviewers. They certainly shouldn’t *replace* human reviewers. But maybe should supplement them. Perhaps even assist editors who can get that feedback first to simply assess whether this is a new idea or a recycled/badly written one. A human should of course *always* double-check this and see if the ai review passes an initial smell/sense test. If the AI agrees with the human editor/reviewers, there you go; if it disagrees, time for a second look.
This doubtless says more about my disgust with some of the absurd things that human reviewers have said about papers which they clearly didn’t read carefully, or approached with great bias, than about my faith in AI.
To answer the first question, yes, I believe this did happen to me, only not as author but as the other reviewer. I reviewed a piece some months ago — maybe it was that of the PhD student you’re talking about — and at some stage I was able to see my fellow referee’s report. To me, it was pretty clearly written by AI. It had generalized high praise, coupled with some unusually detailed (if easily executed) suggestions for revision, and lots of gratuitous summaries of the paper, of a style I’ve never seen in reviews before. I won’t go into all the signs, but I was pretty sure. I didn’t share my concerns with the journal (an extremely highly regarded moral philosophy journal), but I discussed them in the Cocoon.
As it happens, I’m now reviewing a paper where I believe the exact same thing happened, with a fellow ref’s report that is strikingly similar to the other one, as though a common template was used! That only increases my suspicion about both reports, obviously.
My sense is that the other reviewer/s did not wholly hand over their assessment of the manuscript to some LLM. Rather, they read the manuscript quickly and superficially, judged it to be of publishable quality, and then let AI handle the nitty gritty of generating a review, one which they probably requested be both decisively favoring publication, engaging with the details, and recommending very precise, targeted changes it would be easy to implement.
I’m thinking of raising it with the journal editors, but I’m hesitant to smear a colleague with the insinuation that they “cheated” in the review — having only the indirect evidence of the referee report (damning as it is, in my view).
I’m a non-native speaker and I use AI to proof-read every report that I write. Perhaps, some of reports “sound like AI” but the content is from real philosophers.
I think it depends a lot on what prompts you use. If you just ask “please flag all typos and grammatical mistakes” then it’s fine. If you ask to modify your writing, then you might lose your voice and specific style.
I am a native English speaker with proficiency in French and German, proficiency which I have worked very hard to develop. If I were to rely on AI to proofread my writing in my non-native languages, I feel it would deprive me of valuable practice and eliminate my incentive to continue improving. It would, in short, make me worse at my languages.
This is surely an odd response. One might have other ways to practice. This is written for someone else’s sake
Even amongst philosophers, I’m finding the level of discussion this is generating quite surprising. It’s clearly wrong! Asking AI to do it for you is *not* what you’re being asked to do when you’re asked to provide peer review! Just say no to the review request!
I understand why in theory it is immoral for a referee to feed a submitted manuscript into an AI-model. I get the argument, of course. The author didn’t give their consent for the upload.
However, on a purely practical level, most likely, the author themselves has already run their manuscript through some AI. Moreover, one day the paper will be published somewhere, and thus it will be accessed by the various AI models.
I am not saying it is right to use AI to review—I don’t think it is—but just that the “author consent” argument seems very weak to me.
I have never run a paper of mine “through some AI.” Where do you get the idea that “most likely” authors have?
Same. If I find out, during the refereeing process, that an author has done this, I will stop refereeing and tell the editor to find someone else. I absolutely refuse to referee chatbot work. Period.
You should consider doing so, it’s incredibly helpful for finding typos at minimum
Isn’t that what proofreading and (later) copyeditors are for?
Try using Microsoft Word’s spellchecker.
Or printing it out. Or reading aloud. Or changing the font/colour. Or asking a friend. Or having someone else read it to you out loud. (Or making fewer mistakes!)
What an obtuse comment
Which, the report that I would stop refereeing, or the real and proven copy editing techniques?
Can’t tell if you’re sarcastic or 90 years old
But… but… Socrates said the same thing about writing right guys? And AI is going to continue proliferating so we should all just accept it, right? Guys?
Most of my time as a reviewer is already wasted on six cases: (1) papers that contain obvious flagrant technical errors in logic and mathematics; (2) papers that are just gossip; (3) papers that make trivial points, often by adding yet another epicycle; (4) papers that beat some dead horse even deader; (5) papers that have no awareness of the relevant literature; (6) papers that don’t sustain any line of reasoning at all but just wander. None of these cases involves deep or subtle issues.
I’d be very grateful for an AI that could weed out these cases before they get to me. I increasingly decline reviewing tasks because I get swamped with these six cases. I bet editors would be grateful for an AI that could weed out these cases. We all know that the journal review system is strained. I suspect that with some effort we could train an AI to weed out these cases. That would be helpful to our whole publishing ecosystem.
More deeply, as has already been mentioned here, we’ve all encountered reviewers who’ve done really bad jobs reading our papers (they were lazy, incompetent, biased, hateful, etc.). The day is probably far off when AIs can do better, on average, than humans on these deeper issues. But if that day comes, AI would be helpful there too.
If “none of these cases involve deep or subtle issues,” one wonders why so much of your precious time is wasted with them. If the problems in these papers are really that obvious, then just point out as much and recommend rejection. Yes, I’m sure you deal with an immense volume of submissions, and perhaps the majority of them really are as bad as you say, but generally this sentiment reads as a dismissive and condescending lecture from somebody who is already comfortably situated in the academy, to the effect that he shouldn’t have to be bothered with reading work that doesn’t please him. As incomprehensible as this might seem, many young scholars today, scrounging their way through an outrageously abysmal job market, would love to switch places with you.
No one is disputing the fact that the review system needs to be reformed, but I suspect that there are ways of doing so that do not involve acquiescing to AI infiltration.
I waste time because I actually do my job: I read the whole paper, and I flag errors and try as kindly as I can to explain why they are errors. The fact that an error is not technically deep does not mean it is easily visible in a paper. Moreover, I try to see what’s valuable in the paper and to point that out to the author for possible future revision.
This isn’t about work that “pleases” me, it’s about work that fails to meet basic standards for being sent out to reviewers. Editors are overworked, reviewers are overworked, and the strain on the system doesn’t help junior scholars.
An AI that could quickly find errors would be of great help to junior scholars: they could use it to produce papers with greater chances of publication.
I suspect that most people commenting on the review process in this thread are people with experience as reviewers, and so are established professors, so it’s hard to see the relevance of your hostility or your ad hominem.
Yes, any personal criticism is an ad hominem. The question you should be asking is whether or not it’s fallacious, i.e., irrelevant to the point being made.
Here is the rationale: You say you are frustrated with the volume of lackluster submissions. This is perfectly reasonable, but late-career tenured academics often lose sight of what great privileges they enjoy precisely by virtue of being late-career tenured academics, and so tend to regard the strictures of their position as a ridiculous imposition. I am pointing out that your personal circumstances are what allow you to voice this complaint in the first place, and these circumstances are in fact enviable. But you are not used to being called out in this way, and so you wave the point off by calling it an ad hominem. I am actually trying to make a point de homine, you see.
The basic issue I am raising is that your frustration with the quality of many submissions is not a good enough reason to embrace AI-based solutions. The care and attention with which you approach the task of evaluation are extremely admirable, and I hope to be fortunate enough to get the chance to emulate them one day. I just think it’s a shame that you regard this as a burden that you’d prefer to outsource to an LLM.
Everything I’ve said benefits junior people far more than senior people.
If you’re seriously concerned about career asymmetries, you should support development of these kinds of AIs.
I’m assuming here that AI filters would be available to authors before they formally submit their articles to a journal, or that they be used like automatic and anonymous desk-rejects.
I believe AI filters or “conceptual proof-readers” or “sounding boards” would be extremely helpful for junior people. It’s demoralizing to get rejections, especially if they’re based on the kinds of issues I mentioned.
Conceptual proof-readers would make reviewers more sympathetic, because they would have more confidence in the quality of what they were getting. And that they were getting something that is more likely to advance their field. It would signal to reviewers that the paper presents a good idea in a sound way. And it would decrease the number of reviews they were doing so they’d spend more time on those fewer reviews. They’d be more inclined to give helpful feedback.
Junior people would avoid wasting precious time on papers that just aren’t going to get published. Such AI assistants would be able to give early-career scholars much-needed feedback. In particular, good feedback on technical issues is often the hardest to find.
And junior people suffer far more from rejections than senior people, obviously because rejections can destroy junior careers, but also because junior scholars have much shorter time-frames.
I think junior people should get more help. Of course, I’m assuming that the AIs are sufficiently competent to do the job they’re being asked to do. Why would this conceptual proof-reading be any worse than spellchecking, or getting help from a computer on your grammar as a non-English speaker? It wouldn’t. It would be a useful service.
I think this is one of the best things our field could do for junior people. We should be providing junior people with more and better tools for success.
I have a somewhat less optimistic view of the long-term consequences of embracing AI. For one thing, I think that acquiescing to the “faster and cheaper = better” style of reasoning that seems to be favored by AI apologists is a move that will work out to the detriment of later generations of philosophers. It’s not hard to imagine what carrying that way of thinking to its logical extreme would look like, and it doesn’t bode well for philosophy as a discipline. Maybe this is just slippery-slope thinking, though.
At any rate, since your position has now changed from “the problem with the current system is that it’s annoying for me” to “the problem with the current system is that it’s not as friendly to juniors as it could be,” I no longer have much of a bone to pick.
Either the text produced passes the standards of the discipline or it doesn’t. The etiological story shouldn’t matter.
I had thought this was why we liked reviews to be blind. If it’s slop. Don’t teach it and don’t attend to it.
If AI is bad at this particular job, it ought to be the case that journals find a way to police this *in order to* remain relevant.
If that’s not the case, and we’re still convinced this is bad ‘for philosophy’ then I don’t see why such policy should be hiding the reality of disciplinary standards from practicioners
But I’m still confused.
P is in position to do A.
Such a position is envious. Not everyone can do A.
Therefore it is wrong/criticize me for P to do A?
I see no issue with providing an AI with my notes on a manuscript after I have read it and having the AI produce a well-written response based on my notes. I then review and edit the response before submitting it to the journal.
I view this as not much different than if I had an admin assistant I could provide my notes to who would then draft a response letter on my behalf.
I’m still reading the manuscript, I’m still evaluating it based on my expertise, and I’m editing and approving the response I submit. I’m just not typing the response with my own fingers or dictating it into my word processing software with my own voice. But…how much of that is what really matters?
i agree no issue in principle but it will likely sound like AI which may cause editors or authors to think it’s AI all the way down and to on that basis feel shortchanged or wronged. so just be aware of that possibility
IMO, this is unacceptable use.
As far as your admin assistant goes, in that case they’re not just typing up what you’ve dictated/jotted down (note: I am also leery of how this worked historically, when a secretary or student did the work and got no credit); they’re actively organizing it, choosing how to present it, phrase it, etc. This is co-authorship.
I believe that at this point you’re just being dogmatic
Have you never encountered, in the history of philosophy, texts cobbled together from a student’s notes on a Great’s lectures? Do you really think those are just the same as if the Great had written it themselves?
Michel: I think whether this is acceptable depends on the details – and in particular what happens post-AI. I have reviewed papers where I have read the paper myself, written my own notes regarding evaluations/concerns/needed revisions etc., and given an LLM my notes and asked it to tidy it up (organise, put into full sentences, etc.). Crucially, I then both read and edit the LLM output to make sure that the meaning and substance of the original points I wanted to make are preserved and clear, and often end up adding flourishes here and there that better reflect my voice/ideas. From my point of view, this reduces time, cognitive burden and labour that aren’t necessary for the quality or integrity of the review process, without in any way stopping the review being genuinely my own and based on my own expertise. I think this is really different (ethically) from asking an LLM to read a manuscript and generate its own review. I also think (though it’s philosophically more nuanced) that it’s different from the student/secretary case you object to, since the LLM neither puts in creditworthy effort nor is ultimately responsible for any of the conceptual decisions (assuming, of course, a thorough evaluation and editing process of the LLM output).
Brien: As above, I think whether or not there is or is not “an issue” with this kind of process does depend on what happens between the stage of obtaining the LLM output and the submission of the review.
Anya: Good possibility to raise, but I think that if you engage in the genuine post-AI checking and editing process I suggest above, this is less likely.
At least two philosophy colleagues of mine have received clearly AI-generated reviews in the past year or so. Another was asked by an editor with whom she had a connection to review a paper to replace another reviewer who just submitted AI-generated slop
– even though it wasn’t in her area of expertise. The editor was just desperate to try to get some on they knew would actually produce a genuine review. It makes me so angry – not only is it unethical with respect to the author and editor, but it can only serve to undermine our discipline and trust in academia more broadly at a time when that trust is under threat. If you’re not going to actually properly review, have the guts to openly decline.
The legal profession has an obligation to confidentiality that trumps anything you’re describing in point #1, as do medical providers. While consumer AI subscriptions sometimes force you to allow your prompts for training data, OpenAI, Anthropic, and Google have zero data retention agreements that make it possible to ensure that the data is not used for training.
One could also run an LLM locally, such as the open source models out of China or Facebook’s llama 4.
Thus, #1 (which I agree is a trumping objection) is defeasible so long as the reviewer has taken the proper care.
#2 is very important right now, but only insofar as the norms and policies are as they are now. Also relevant: Elsevier et al are partly guarding against AI in papers because certified non-AI text written by experts is extremely financially valuable… when sold to the AI companies. I’m not sure how much their profit motives should motivate us to comply.
Still, your fundamental point is #2 just seems right: always it seems like a waste of time to merely share AI reviews without adding value oneself. Supplementing your own reading and criticism, as you indicate might be possible in #3, however, seems like a possible middle road, and I don’t believe you’ve addressed it here.
I am still tempted by the following norm, though: if an author has not herself tested her paper against LLM chatbots and responded to the main concerns and objections, they have not exercised due care in writing. At some point in the not-too-distant future, it seems pretty plausible that “did you bother to check for counter arguments and respond to objections from Claude?” will become the equivalent to “did you even bother to proofread?” or “your lit review missed the main paper / recent work in this subfield” and the obligations will shift.
Certainly not yet! But someday. And there doesn’t seem to be a clear argument here against that eventual change in policy that’s not circular. In that era, sharing the ChatGPT criticisms would be like saying, “By the way, Google pretty clearly mentions Author X when you look up this topic, did you fail to do basic research?”