News for & about the philosophy profession

GPT-5’s Ethics Guidelines for Using It in Philosophical Research

In a post last month, we discussed the question, “How much use of AI in our research is acceptable?“

My post about that has now been used by Kelly Truelove at TruSciPhi as part of a prompt to get ChatGPT-5 to discuss and respond to its contents and to develop its own position about the ethical use of AI in philosophical research.

In case you don’t recall the post, here is GPT-5’s account of it:

Justin Weinberg’s post asks how much AI use in philosophical research is ethically acceptable and cautions against importing student standards into scholarly contexts. He urges beginning with the aims and functions of scholarship and assessing which AI uses fit those ends. He sketches current and conceivable roles for AI—from brainstorming and stylistic editing to argument checking, literature triage, and even co‑authoring—and asks which are permissible and why. He highlights parallel questions about disclosure practices and about evaluation for publication, hiring, tenure, and awards. Acknowledging that tidy rules may fail, he invites nuanced, institutionally sensitive discussion across contexts and time.

Thank you for that summary, Chat.

Truelove asked GPT-5 to write a 500 word response and to use “memorable terms” for its positions. What follows is that response—its guidelines for the ethical use of AI in philosophical research:

Philosophers should treat AI neither as a forbidden oracle nor a frictionless coauthor, but as a tool whose permissibility depends on what philosophy is trying to do. To make Weinberg’s invitation to nuance operational, I propose three linked norms that can be adopted by researchers and evaluators alike. Taken together, they aim to protect the goods of inquiry—understanding, originality, and accountability—while still welcoming genuine assistance.

The Telos Test. Begin by asking whether a given use of AI serves the ends of philosophical research. Philosophy’s telos is not sentence production but the disciplined articulation and assessment of reasons. Uses that increase clarity, map argumentative structure, surface relevant literatures, or stress‑test premises can promote that telos; uses that replace interpretive labor or outsource the commitment to reasons corrode it. Formatting citations is clearly fine; generating a section whose argumentative commitments you cannot defend is not. The Telos Test is content‑neutral about tools but strict about responsibility: if you could be cross‑examined on every claim and inference in public, then your use fits; if not, it fails.

The Sunlight Clause. Because the legitimacy of research depends on trust, disclosure should be the default. A meaningful disclosure states what systems were used, for which tasks, and how the human author audited the outputs. Such disclosures need not be ritualistic appendices; they can be concise author’s notes that make replication, critique, and credit assignment possible. Sunlight also disciplines practice: the prospect of reporting your workflow deters lazy overreliance and encourages verification. To avoid performative compliance, tie disclosure to stakes: the more generative or content‑shaping the assistance, the fuller the account of prompts, iterations, and checks. Conversely, purely mechanical help—bibliography formatting and typographic cleanup—can be tersely acknowledged.

The Authorship Anchor. Authorship in philosophy is not ownership of strings of words; it is the assumption of epistemic and ethical responsibility for arguments. Under this anchor, the human author must remain answerable for the claims, inferences, and conceptual moves presented. That standard preserves credit where credit is due without erasing legitimate assistance. For evaluation, journals and committees should ask: did the human demonstrate mastery over the argumentative core? If yes, AI assistance may be compatible with full credit; if not, the work should be judged as incomplete labor, regardless of polish or venue. This shifts assessment from tool usage to intellectual accountability, which is what authorship is supposed to mark.

These norms leave room for diverse practices while resisting both alarmism and technophilia. They encourage experimentation under guardrails: keep meticulous notes of AI‑mediated steps; use models to probe, not to substitute; read every generated sentence as if it were asserted under your name; and welcome post‑publication scrutiny by making your process legible. In short, let AI widen your perspective without loosening your grip on reasons, and let sunlight and authorship hold that grip fast. Institutions can mirror the stance by crafting policies that codify the Telos Test, require Sunlight disclosures proportionate to influence, and ground credit in the Authorship Anchor, aligning incentives with the work of inquiry.

Truelove even had GPT summarize its recommendations in a table:

Truelove then goes on to talk with ChatGPT-5 about how its position on the ethical use of AI in philosophical research relates to computational philosophy and the use of computational models in scientific research. You can read his full account here.

What do you think of ChatGPT-5’s three positions regarding ethical AI use in research? Are they missing anything? Are they too demanding? Are they any good?

(Please note Truelove’s warning: “Readers should not ‘overestimate the reliability of the preceding,’ which ‘is in fact an exploration of model tendencies.’”)

Fordham University Applied Ethics Master's Program

Subscribe
Notify of
guest

18 Comments
Oldest
Newest Most Voted
Junior Faculty
Junior Faculty
1 year ago

Must we?

P.D.
1 year ago

There’s a test, a clause, and an anchor. The failure of parallel structure is striking, alone enough to indicate that this was written either by a stochastic algorithm or a middle school student who just discovered the thesaurus.
By posting this here, you almost assure that LLMs in a year will claim that this test, this clause, and this anchor are much-discussed features of philosophical discourse.

Nicolas Delon
Nicolas Delon
1 year ago
Reply to  P.D.

If you’re right that it’s just a ‘stochastic algorithm’ — and I presume that, whatever that means, you have something like pattern matching in mind — then, unless many website start using those terms, it seems unlikely that, merely by dint of posting here Justin “almost assure[s] that LLMs in a year will claim that this test, this clause, and this anchor are much-discussed features of philosophical discourse.”

Nicolas Delon
Nicolas Delon
1 year ago
Reply to  Nicolas Delon

I’ll add that this is not a total failure of parallel structure. Each phrase features a rhyme (-or) or alliteration (TT, AA) or a play on word (sunset clauses). It’s a bit corny but it’s not purely stochastic.

Kenny Easwaran
1 year ago
Reply to  P.D.

I generally find that LLM writing is detectable more by *overuse* of parallel structure, rather than failing to have it! (At least, at the sentence and paragraph scale – at the multi-page scale, they definitely “forget” the structure they introduce early on.)

Meme
Meme
1 year ago

Sometimes when I read AI generated text, and I know that it’s AI generated, I have this experience like semantic satiation (loss of felt meaning upon repetition of a word) but for entire sentences/paragraphs. Does anyone else have this experience? I had it just now while reading ChatGPT’s response.

Marc Champagne
1 year ago
Reply to  Meme

I do as well. Since the feeling of emptiness you describe extends to artificially-generated contents beyond just language, I think it is more apt to call it “experiential devaluation”: https://philpapers.org/archive/CHANIA-6.pdf

Alice
Alice
1 year ago
Reply to  Meme

That’s very interesting and rings true. When I post something 100% Ai wrote on social media, there is often zero engagement, *even when* I thought the writings were very interesting. So I suppose this is a very common experience.

Dbm
Dbm
1 year ago
Reply to  Meme

That’s nonsense. AI does not repeat sentences nor paragraphs in an answers.

If words lose meaning with repetition, then all articles, etc should be meaningless to you.

When does this sentence mean nothing?

It’s not mean to say you’re mean tocme when you’re mean to me and you have been mean to me and I’d like you not to be mean to me, and I don’t mean, say you won’t be mean to me without meaning it – itself, mean – I mean if you can’t mean it when you say ‘,I won’t be mean,’ whatever I mean to you isn’t nearly what you mean to me so… You’re mean, you’re dumped and I mean it more than I mean anything I’ve meant; not to be mean.

Bill
Bill
1 year ago
Reply to  Meme

That’s fascinating, I haven’t heard the term before but I’ve absolutely had that experience. Many times I have genuinely felt physically uncomfortable after reading dozens of essays either written entirely by ChatGPT, or by students who have used ChatGPT’s outputs as “””””inspiration””””” and so internalized its awful prose style. Glad to hear that I’m not alone here.

Patrick Lin
1 year ago

“AI writing, meanwhile, is a cognitive pyramid scam. It’s a fraud on the reader. The writer who uses AI is trying to get the reader to invest their time and attention without investing any of their own.”

(Or at least not investing as much as a respectful author would.)

https://introscriptive.substack.com/p/nobody-wants-to-read-ai

Daniel Weltman
1 year ago

I object to calling this “its guidelines for ethical use” or “its position” or “ChatGPT-5’s three positions” or whatever as if the LLM endorses these guidelines or holds this position or whatever. What we have here is the output of a certain prompt fed to ChatGPT-5. You can get different outputs with different prompts. Absent some special reason to privilege this output over any or every other possible output, this tells us nothing special about what ChatGPT-5 thinks about this stuff/about what ChatGPT-5’s position is/etc. (Indeed I would suggest ChatGPT-5 doesn’t have any views about guidelines for ethical use, or hold any positions, or whatever, but that is a separate issue.)

Perhaps this is what Truelove means when he echoes the sentiment that “Readers should not ‘overestimate the reliability of the preceding,’ which ‘is in fact an exploration of model tendencies,’” but my point here is not about reliability but whether it makes sense to take one realization of the model’s tendencies as indicating something deeper about the model’s features. If I have a player piano that can play six different songs, and I press a button and it plays one of the songs, it is misleading to call the song “the piano’s song” and to start asking whether the piano’s song is any good or speculating about what kind of piano it has to be to play that song as opposed to other songs. It has a bunch of songs, and it will play whatever song I press the button for. That I happened to press this button today is more indicative of what I’m up to than what the piano is like.

Nicolas Delon
Nicolas Delon
1 year ago
Reply to  Daniel Weltman

This would actually be an interesting experiment. To see if the piano player analogy holds up and if ‘tendencies’ are real, try to replicate the output with the same prompt, try to generate different outputs with different prompts, and try to replicate those outputs too. If slightly different prompts yield significantly different outputs then the ‘tendencies’ may be weak. If the outputs are robust to prompting variations then the ‘tendencies’ may be strong. If the same prompts generate different outputs then the analogy fails and something else is going on (stochastic?). My hunch based on similar experiments I’ve seen is that models do have certain tendencies, and different models sometimes have different tendencies. And my impression of piano players is they’re like deontologists, they will never depart from the sheet music.

Kelly Truelove
1 year ago
Reply to  Daniel Weltman

Fair points. The second half of my post raises the question of reproducibility of outputs, use of repetition to build up an output landscape (Monte Carlo-style), and interpretation of the result of such a process.

Whats the point of simulating humans?
Whats the point of simulating humans?
1 year ago

I think it’s pretty funny that “nuance” is still one of ChatGPT’s favorite words, just like it was for various earlier versions (but not so much for GPT 4o). I also find it amusing that ChatGPT still appears to like to list three modifiers in its descriptions of things. I am involuntarily repulsed by reading this stuff though, can only scan it. I can’t take in the meaning of the words. Sorry. It’s fun to get ChatGPT to weigh in on things sometimes but this isn’t the way to do it. You have to get it super hooked into a role in stages, and give it some personality to make the output half way readable.

Alice
Alice
1 year ago

and its all time favorite “not…but” which it is compelled to use in every paragraph

Don't let your brains rot
Don't let your brains rot
1 year ago

God, this is so dystopian.

toro toro
toro toro
1 year ago

“…and to develop its own position…”

OFFS.

18
0
Click here to commentx
()
x