If You Think You’re Refereeing an AI-Written Submission, What Should You Do?


You’re refereeing a paper and you come to suspect it was substantially and/or illicily written by an LLM. What should you do?

My suggestion:

  1. Keep in mind the possibility that you are mistaken as well as the possibility that the AI use falls within the journal’s rules and was disclosed during the submission process.
  2. Pause your work refereeing the piece, as your suspicions of illicit AI use may unduly negatively influence your assessment of the submission.
  3. Gather a few examples of passages that strike you as AI-written, briefly explaining why. (Do not feed the manuscript into an AI-detector, for two reasons: first, doing so may violate the norms of confidentiality referees are expected to abide by, and second, such detectors appear to currently be of questionable accuracy.)
  4. Write to the journal’s editor or managing editor expressing your concerns and sharing the evidence you’ve gathered.

With luck, the editor will look into the matter and either agree with you and relieve you of refereeing the piece, or the editor will convince you that you are mistaken and you can proceed to finish refereeing the paper.

But what if you aren’t so lucky?

One associate professor of philosophy wrote in with the following:

I recently accepted a referee invitation for a good generalist journal. I was completely convinced after about 20 minutes with the submission that it was mostly (if not entirely) written using AI tools of some sort. The paper was in my areas of research. The editor disagreed (or at least did not share my level of confidence), and so I was left with the awkward task of writing (brief) comments for the “author” to go along with my rejection. For obvious reasons, this seemed to me to be an obvious waste of everyone’s time. Nothing like this has happened to me before, and I referee fairly regularly. 

The professor is curious about the frequency with which reviewers are finding themselves in disagreement with journal editors and/or other reviewers on particular cases. Has this happened to you? What did you end up doing? How should referees and editors navigate disagreement on this?

guest

16 Comments
Oldest
Newest Most Voted
Inline Feedbacks
View all comments
Mark
Mark
1 day ago

and so I was left with the awkward task of writing (brief) comments for the “author” to go along with my rejection.

Or maybe the awkward task of telling the editor that your verdict is Reject but that you won’t be providing comments to be fed back into the model to generate the next version of the paper.

Sam
Sam
1 day ago

Have another LLM write the referee report.

Kenny Easwaran
1 day ago

As an editor, I’ve received a bunch of these papers. As people may be able to tell from my comments here, I’m ideologically committed to the idea that there *could* be a worthwhile paper whose text is substantially AI-composed, but so far only one of the dozen or so papers I’ve had this suspicion about seemed promising enough to send out to a referee. (There was one other that I unfortunately identified *after* getting referee comments, and I apologized to the referee after the fact.)

As far as comments, I don’t think anything that is actually AI-authored is getting past the initial desk reject phase. If the paper looks like a real paper throughout, and makes a point and gives some sort of argument for it, there’s almost certainly a real human there who is thinking about the project. In some cases, it’s a person working across disciplinary boundaries, or across languages, relying on the AI to get the disciplinary norms and language right (but unfortunately, these AI systems aren’t actually good enough tools to be effective for this purpose yet).

I’ve been writing fairly substantial comments on a bunch of these in my desk rejections, pointing out formalisms they take pains setting up but never use, or ideas that they say they will argue for but no argument comes. I’ve been doing this because there’s often an interesting idea there that I really want to see someone write a good paper about!

But I’ll probably stop doing this when it’s just another paper written by AI, arguing that AI has constitute epistemic limitations and therefore can’t be a testifier or a knower (which seems to be the claim of about half of these AI-written papers).

Brian Earp
Reply to  Kenny Easwaran
14 hours ago

Hi Kenny
I’m the EIC of JME Practical Bioethics (a BMJ journal, sister to Journal of Medical Ethics which I also edit) and we’re preparing a special issue on “Doing Ethics with AI” (as opposed to “the ethics ‘of’ AI”) where one of the main questions we’d like authors to debate is, under what conditions if any would it be okay, even desirable, for a philosophy/ethics journal to publish work substantially generated by AI, assuming it made a worthwhile contribution — and do we think it is even possible that a well-prompted AI could generate something that would count as a meaningful advance in the literature? Now, the focus is ethics research specifically rather than all of philosophy (because of the scope of the journal), but we’d welcome your philosophical reflections on the matter, perhaps in a short piece… I see Dan Greco’s commented below and has a similar attitude to yours; he’s submitted something already 🙂 More generally, anyone reading this, if you want to work through the philosophy and/or ethics of philosophers using AI to generate normative analyses and submitting them to journals – and what editors/journals should do, how readers should feel about this etc. – please let me know and I can send you the information on the special issue. I’m at [email protected] — cheers. Brian

Daniel Greco
Daniel Greco
1 day ago

My attitude is similar to Kenny’s. So far I’ve seen some as an editor, and some as a referee. So far I’ve never thought: “this is a good paper, but it looks like it’s AI, so what do I do?” Rather, the kind of papers that are striking me as likely AI written are ones that go on for pages saying very little, or which have other serious substantive problems. So I’ll recommend rejection in a way that doesn’t have to cite their having been AI-written.

I suppose, like Kenny, I’m uncomfortable with the idea that an otherwise publishable paper–one that makes an interesting, worthwhile contribution to philosophy–should be rejected on grounds of AI authorship. Though that’s a bigger conversation. (Very roughly, my view is that just as a mathematical proof of a significant result is valuable because of the understanding it provides its readers, regardless of whether it’s mainly human or mainly AI-authored, likewise with a philosophy paper. If we get to the point that they’ve already gotten in math, where AI is playing a major role in generating important results, those results should absolutely be publicized, and the journal system is our discipline’s primary way of publicizing what we take to be worthwhile contributions that advance our collective understanding of philosophy.) But that’s also very much not what I’m seeing so far.

Eric
Eric
Reply to  Daniel Greco
1 day ago

I dont understand why this view is so controversial.

MBW
MBW
Reply to  Daniel Greco
2 hours ago

I agree. If AI becomes a useful working tool, disclosing its use would make about as much sense as disclosing Word vs. LaTeX, or like any other methodology section. The rule should be that the author owns their work (and its flaws.)

Jessie Ewesmont
Jessie Ewesmont
1 day ago

Reject, with the comment that it’s AI generated. Easiest review ever.

Matt L
1 day ago

as well as the possibility that the AI use falls within the journal’s rules and was disclosed during the submission process.

I am curious about what the idea is here. It seems to me that this sort of disclosure should also be made available to the referees. Maybe that’s wrong, but I don’t think it’s obviously wrong, and I’d be interested to hear what people have to say about it.

Eric
Eric
Reply to  Matt L
1 day ago

Maybe say first why it should? Isnt the referees job to determine the quality and originality of the paper, and that seens orthogonal.

Matt L
Reply to  Eric
20 hours ago

I suppose it might depend on if I think authors are being candid on their declarations. (We have students make such declarations, and I’m sure many are not candid.) But, if they were, and they explained how AI was used, it would help me come to a better conclusion on the paper, I think. Maybe that’s wrong, but it doesn’t seem obviously so to me.

Eric
Eric
Reply to  Matt L
7 hours ago

you haven’t said a single word about why it should help you come to a better conclusion on the paper. you added some irrelevant stuff about whether they are candid. assume they are.

I understand the view that AI writing is slop. Maybe most is. Maybe all is. I don’t understand the view that you should be refereeing a paper that you can’t use your own judgment to decide if its good or not–and that you need to know facts about what tools were used to create it.

Matt L
Reply to  Eric
41 minutes ago

Okay. I’ll say a bit more. Here are a few things. I am sometimes asked to review things that are not squarely in my wheel-house. How much time I spend looking up claims, or how deferential I am to an author’s citations, can be influenced by if they disclose they used AI in a literature search or in research. Similarly, if I am wondering why something is written in a particular way, and what I might say about changing it, whether AI was used to help with working in a different language can help as well. (I have lots of experience with these things from working with student papers, among other things.) Now, I do expect lots of people to lie about this stuff, sadly enough. That’s one reason why I expect AI declarations to not be much use.
Another way is that if someone says no AI was used, and then I get a couple of obviously made-up citations, I’ll know that there is simply no point in reading further – that the person has not only wasted my time, but has been deceitful about it.

I don’t expect everyone to think those are the right paths. And, as I noted, I expect a lot of people will simply lie. But, if there are going to be such declarations, I think they should be made avaiable to the referee.
(For what its worth, I think that at least some conflict of interest statements should also be made available.)

J.P. Loo
1 day ago

If Pangram works, that solves one problem (how to decide whether it’s AI-generated),* though not others (principally, what to do about it). I suppose I’m quite undecided about that. In certain areas (e.g. philosophical logic), we esteem largely formal papers of a kind that would be no worse off for having been written by an AI. I’m not sure we should think the same of e.g. phenomenology. (And a final complication is that I suppose in principle we might one day create AI systems that are saliently autonomous, conscious, or what have you, and so could write phenomenology papers just as interesting as the best written by humans.)

* Although there are some good outstanding criticisms of Pangram, one interesting point is that you can just email them with a half-baked idea and (n=1) they’ll give you 100,000 tokens to experiment.

Jessie Ewesmont
Jessie Ewesmont
Reply to  J.P. Loo
20 hours ago

I recently read a blog post where the author was experimenting with Pangram. When he fed it a full essay (which he manually wrote), it came back 100% human generated. When he fed it a fragment of that same essay, it came back 100% AI generated. When he fed it a fragment of that fragment, it came back 100% human generated again. This is a pretty odd result, and it makes me distrust Pangram out of worry that it’s inconsistent.

Rob Hughes
Rob Hughes
3 hours ago

If a paper contains fabricated references, serious misrepresentation of sources, or blatantly false statements of fact, that is reason enough to reject the paper without further comment.

If a paper doesn’t contain fabrication but is bad, and it appears LLM-generated, I would write up a review. I would try to keep it short. I might comment about my suspicions in the confidential letter to the editor, but I would reject the paper on grounds of poor quality.

If a paper appears to have merit, but it also appears to violate the journal’s policy on LLM use, I would reach out to the editor.

I think there is little chance of an entirely LLM-generated paper in philosophical value theory being good enough to deserve publication. I think there is a high chance a journal could receive an interesting paper that is mostly human-written but has some LLM editing or additions that violate the journal’s (legitimate) policy on LLM use.

It is unethical to submit someone else’s unpublished manuscript to an “AI detector” without their explicit consent. It is unethical to use “AI” systems to review papers.