News for & about the philosophy profession

The Ethics of Using AI in Philosophical Research

“How much use of AI in our research is acceptable?”

In a recent email, a professor of philosophy asked for a discussion about the ethics of using AI in philosophical research, and institutional policies about it.

He adds: “My instinct is to hold my students to the same standards that I would expect of my colleagues.”

While that is one possible conclusion we might land upon, I don’t think it’s a particularly good place to start, since there seem to be relevant differences between student coursework and scholarly research (e.g., the purpose of the work, the roles of the agents, etc.).

Better, it seems, to consider the nature of scholarly research, the aims people have for undertaking it, the institutional functions it serves, and so on, and try to figure out the extent to which various uses of AI fit, if at all, with those things. And I think it would be good to take seriously from the beginning of this inquiry the possibility that a clean and neat answer is unlikely to be the best one.

There are different ways researchers in philosophy might use AI in their work, now, and with some improvements to the technology, in the future. Here are some examples (feel free to add more in the comments):

  • chat or brainstorm with it about your ideas, as you would with a colleague
  • make stylistic edits to your writing to enhance comprehensibility
  • format bibliographic information
  • draft an index
  • draw out the structure of and check the validity of an argument you’ve made
  • search through articles and other resources to see which will be useful for a specific project
  • raise objections to an argument you’ve written
  • explain technical work in relevant empirical fields
  • create examples, counterexamples, and thought experiments
  • write abstracts of papers
  • ask about the implications of your thesis for other questions
  • have it analyze a paper you’ve written and provide detailed outlines of other papers you could easily write based on the ideas of the first
  • co-author substantive text with it, as you would with a human co-author (some of the resultant passages are written solely by the AI, some by you, some a mix)

Which of these are okay? Which not? Why?

In addition to the question of which specific tasks a researcher might acceptably have an AI perform, there is the question of disclosure. What kinds of AI assistance must researchers acknowledge in their talks, articles, and books? What form should such acknowledgement take? And what’s at the intersection here? Might all AI use be permitted so long as it is disclosed? Or are there other considerations here besides transparency?

Beyond what’s permitted, there’s the question of evaluation—of works and of authors. To what extent may/should one’s use of AI be taken into account when it comes to questions of accepting manuscripts for publication, or for hiring, tenure, promotion, honors, and awards? (“Yes, they have an article in PPR, but it was co-authored with Chat-GPT, so does it really count?”) (See this earlier discussion of journal policies about AI use.)

There’s more, but perhaps that’s enough to get us started. Discussion welcome.

Central European University Philosophy

Subscribe
Notify of
guest

74 Comments
Oldest
Newest Most Voted
Brian Weatherson
Brian Weatherson
1 year ago

I think it’s implicit in some things you listed, but I’d add as a separate category *Checking the bibliography*. I’m sure I’ve got lots of errors in my .bib file, and it’s impractical to check the thousands of entries to see where they are. But for any given paper, it’s reasonable to ask the computer to double check the entries in that paper.

This isn’t perfectly reliable! I just had it tell me something was wrong when in fact the citation was correct but many others have cited it incorrectly. But it’s reliable enough that it’s a useful way of allocating checking time.

To be sure, this is a little lazy; I could probably check all 60 myself. But if someone was editing an edited volume, and there are 1000 citations across all the papers, it seems fair enough for them to ask ChatGPT or Claude or something to see what it thinks of the 1000 citations, and then personally check the ones that the computer is most worried about.

Paul Wilson
Paul Wilson
1 year ago

“How much use of AI in our research is acceptable?”

How much use of AI in our research is *detectable*?

Welcome to a new world, far beyond the uncited contributions and unattributed research of graduate assistants.

Patrick Lin
1 year ago

Here’s where we drew the line on when to credit an AI/LLM in one’s work, which is very different from the position of most journals that AI should never be a co-author: https://ethics.calpoly.edu/AIauthors.htm

But this is not necessarily to endorse the idea of using AI in any of one’s writing. Per my guest post here last week, I would generally advise resisting that temptation.

WWW
WWW
1 year ago

Any and all use of AI is acceptable. But AI is not an author. The author remains responsible for the content of the work. This should be the policy I think.

Any further restrictions imply that authors are not allowed use the technology available to them to produce the best work possible. Unless there are very strong ethical objections to the use of the technology, like with, say, dangerous experiments on human subjects, that seems untenable. And I don’t see that there are any such strong ethical objections.

Amy
Amy
1 year ago
Reply to  WWW

I agree

Nick Hadsell
Nick Hadsell
1 year ago
Reply to  WWW

I will shamefully self-promote a paper where my friends and I argue that responsibility or accountability isn’t a necessary condition for authorship, which means AI could count as an author. And this excellent paper by Joshua Habgood-Coote et al (emphasis on the et al–there are *lots* of contributors!) shows authorship norms are surprisingly hard to pin down.

WWW
WWW
1 year ago
Reply to  Nick Hadsell

Thanks for sharing. Could you give us a brief pitch: what social function could it fulfill to call an AI an ‘author’ that can’t be fulfilled by AI use disclosure statements (even these seem dubious to me)?

A human sponsored AI paper just sounds like a human authored paper to me (and also not realistically a great paper with current technology).

Nick Hadsell
Nick Hadsell
1 year ago
Reply to  WWW

That’s a good question. A lot of our paper tries to get readers to treat cases like. In human authorship, there are cases where accountability/responsibility are not required for authorship (e.g., pseudonymous authors). So, it’s no problem if AI isn’t accountable or responsible, because our human authorship norms already don’t seem to require it.

Another case where human authorship norms don’t seem to require responsibility or accountability for a work is this: imagine Gettier discloses many years after his famous paper that it was a result of his dropping his typewriter down the stairs and, lo and behold, the stairs hit all the keys needed to type up his paper! There is a sense in which he’s responsible for the paper, but that seems to fall short of the thicker sense most people are talking about wrt responsibility in authorship norms. And yet, I think that in a case like this, Gettier could still be considered the author of the paper. It would be strange if Gettier disclosed this to the editor and the editor goes, “Sorry, we can’t publish this paper, even though it will would utterly an entire field of philosophy.” Or, “We can publish this paper, but we cannot credit you as the author.”

Ofc, lots of this is personally tentative. We are obviously now in a strange new world!

Alex Bryant
1 year ago

I think it’s worth having this research-specific discussion, and I’ve seen a lot of productive discussion about how to think about students’ use of generative AI tools as well. Beyond this discussion, though, I want to flag some upstream considerations as well.

The question being posed in this post is, roughly, “when is it ethical to make use of generative AI tools in philosophy research?” As Justin mentions, there are lots of ways an individual researcher might make use of these tools, and Brian W. has also pointed out that there are appealing ways these might be used in collaborative work, e.g. for editorial work, etc.. I think it’s worthwhile pursuing the ethics of these practices in the narrower way this post is framing the discussion (e.g. discipline-specific norms of academic integrity, etc.). This said, I think it’s important at these very early stages of explicit generative AI tool use to not lose track of the larger ethical considerations that might lead on to think genAI usage is immoral.

At a more general level, there’s a question of when using these tools is morally permissible, full stop. For some, the global environmental and social impact of the development and public dissemination of these tools will lead to the conclusion that using the standard versions of them is impermissible. It could also be that one or another model is objectionable because of the training data it uses, or some other model- or company-specific issue. Maybe it turns out only local model usage is morally permissible, but then we’re into considerations of access to resources–maybe the use of genAI is research is a problem of distributive justice, and so on and so on.

I think it’s useful to have this discussion about the ethics of the specific use cases, I just want to have a least one pin this thread about the broader questions about the ethics of genAI usage far upstream of these.

Kenny Easwaran
1 year ago
Reply to  Alex Bryant

And it’s important to remember that the environmental concerns and the intellectual property concerns are quite distinct.

Whatever environmental concerns there are around the use of large language models should be shared with any research activity that is similarly energy intensive, like watching videos of lectures, or driving to campus. (There are of course research activities that are far less energy intensive, like Googling and reading .pdfs, or riding a bus or bike to campus, and ones that are far more energy intensive, like flying to conferences).

Intellectual property concerns are very different, and don’t have easy comparisons to any other research activities. They also might be different between different AI services.

To the extent that there are other social impacts, it’s worth considering whether those are greater or less than those of using social media (either for promotion of one’s own research, or learning of the research of others) or importantly different in some other way.

Bruce P Blackshaw
Bruce P Blackshaw
1 year ago
Reply to  Alex Bryant

If using GenAI is impermissible for environmental concerns, then using search engines is also likely to be impermissible (the most recent research I’ve seen claims the impact of a GenAI query is similar to a Google search, and that earlier claims of an order of magnitude difference are incorrect).

David Wallace
David Wallace
1 year ago

To put this in perspective: An AI query costs about a kilojoule of energy. Driving for a few minutes costs about ten thousand kilojoules.

Will Fleisher
Will Fleisher
1 year ago
Reply to  David Wallace

I worry about this for two reasons. First, as a general point, the fact that a bad thing is less bad than another bad thing doesn’t make it not bad. Driving a gas powered vehicle is canonically a bad thing for climate and other environmental reasons. That many of us are essentially forced to do it gives a reason to think it’s perhaps permissible. But using generative AI is not a requirement for survival in our society, so it doesn’t get the same justification.

Second, these claims about power use of individual queries are complicated, and most of the sources are back of napkin type calculations. However, when we zoom out and look at the changes to the power infrastructure being caused by the AI boom, things look quite different. New power plants are being brought online, coal plants are being kept in service, and there are massive infrastructure projects being undertaken by utilities (the costs of which are being passed on to energy consumers).

That suggests that, however we carve up individual causal responsibility, AI is driving a huge increase in power consumption, and is drawing in huge material resources to support it. Meanwhile, we are in an all hands on deck situation for the green energy transition, meaning that these resources are needed for other things, and we desperately need to cut our power usage, not grow it.

Hence, I don’t think these “less than driving a car” types of arguments are very persuasive.

David Wallace
David Wallace
1 year ago
Reply to  Will Fleisher

All fair. I don’t think the energy comparison is decisive (and I agree that growth of data centers is a major issue for energy politics), I just think people often discuss these energy costs with no sense of the actual, quantitative levels. (I recall people in the pandemic seriously suggesting that the energy costs of Zoom calls needed to be allowed for in comparing the environmental consequences of in-person vs. remote conferences.)

Kenny Easwaran
1 year ago
Reply to  Will Fleisher

I do think it’s important to look at the actual numbers here though. I think the absolute worst upper bound for the impact of AI on CO2 emissions is if every single dollar spent by these companies were spent on diesel fuel to power generators. Their spending seems to be about three times their revenue. So a rough upper bound for the CO2 emissions caused by a $20/month subscription to a AI system is spending $60/month on gasoline – not great, but well in line with a lot of discretionary harms that a lot of people cause (and this is a worst-case upper bound).

The most relevant comparison is probably online streaming video – both passive watching (whether it’s TikTok or YouTube or Netflix or Hulu) and active videoconferencing with something like Zoom. I’ve seen various estimates that suggest that the impacts of an LLM query are comparable to the impacts of somewhere between a few seconds and a few minutes of streaming video use. I don’t have a lot of trust in specific details of what precise ratio it is, but somewhere in this order of magnitude seems very plausible.

Training must be included
Training must be included
1 year ago
Reply to  David Wallace

Does this include the energy invested in the training process? Though I have no recipe for how to account for that in a query, I think that it must be included, given the expectation that it would shift the calculation nontrivially.

Kyle S. Hodge
Kyle S. Hodge
1 year ago

Some of the items on the list seem permissible for research, attribution optional.

  • chat or brainstorm with it about your ideas, as you would with a colleague
  • make stylistic edits to your writing to enhance comprehensibility
  • format bibliographic information
  • draft an index
  • draw out the structure of and check the validity of an argument you’ve made
  • search through articles and other resources to see which will be useful for a specific project
  • raise objections to an argument you’ve written
  • explain technical work in relevant empirical fields
  • ask about the implications of your thesis for other questions
  • have it analyze a paper you’ve written and provide detailed outlines of other papers you could easily write based on the ideas of the first

The rest seem permissible, but only with attribution.

  • create examples, counterexamples, and thought experiments [specifically thought experiments]
  • co-author substantive text with it, as you would with a human co-author (some of the resultant passages are written solely by the AI, some by you, some a mix)
  • write abstracts of papers

There are many aids to my research that I don’t acknowledge, even though they make material contributions to the efficiency of my research and writing process (e.g. Zotero, JSTOR’s side panel recommending related articles, etc.).

Right now, I would resist using the technology to make substantive contributions to research, but if such contributions were made, it seems professionally obligatory to clearly acknowledge that that is how those contributions originated.

It is natural to expect the assessment of work which relied heavily on the use of AI to be held against the human author in some sense. If I coauthored an interdisciplinary paper with a physicist which involved deep knowledge of theories of physics and complex computations, I would expect readers to presume that I was not responsible for what was substantive about the physics contributions. It would be more impressive if I could do both what a physicist can do and what a philosopher can do, which it would be reasonable to attribute to me if I was a solo author of the same paper. The case will be similar, and perhaps slightly more extreme, with AI attributions. One major concern is how likely readers will be to assume good faith on the part of the author, i.e. that the author only used substantive AI contributions in the places they acknowledge, something it is harder to do with human authors.

In general, I think the broad goal should be to avoid misrepresenting one’s own philosophical capacities, and this is something of which I think most professionals acting in good faith have a reliable inner awareness.

Amy
Amy
1 year ago
Reply to  Kyle S. Hodge

Can you explain why you would think these are impermissible without attribution:

  • create examples, counterexamples, and thought experiments [specifically thought experiments]
  • co-author substantive text with it, as you would with a human co-author (some of the resultant passages are written solely by the AI, some by you, some a mix)
  • write abstracts of papers
Kyle S. Hodge
Kyle S. Hodge
1 year ago
Reply to  Amy

I think examples and counterexamples are fine (since they simply draw your attention to things you were not otherwise aware of), but thought experiments are a bit more elaborate and better reflect one’s capacity for thinking philosophically. Similarly with co-authorship. Lacking attribution, one is explicitly or tacitly representing oneself as having engaged in philosophical activity that one did not actually perform.

I’m a bit more ambivalent about abstracts, admittedly. Maybe it is just vibes, but I don’t like the idea of submitting or receiving something which is totally the product of AI, even if I, or the author, carefully reviewed and fully endorsed the content. If abstracts are merely devices of convenience, then perhaps it is permissible. What do you think?

Kenny Easwaran
1 year ago
Reply to  Kyle S. Hodge

To the extent that one ought not misrepresent one’s capacities, I think there’s an interesting question about whether putting forward an AI-assisted paper as one’s own work is such a misrepresentation. Is it important that one’s CV represent one’s capacities in an unaided setting? Or is it fair to have a CV that represents one’s capacities given standard institutional support?

Of course, many people think that the purpose of a CV or publication is to do something other than represent one’s capacities to employers, such that misrepresentation is not obviously a problem.

Kyle S. Hodge
Kyle S. Hodge
1 year ago
Reply to  Kenny Easwaran

I think this is a fair question. What do you make of my analogy with a physicist coauthor? Supposing everyone writing on philosophy and physics had their own physicist assistant on the basis of institutional support, could the philosopher take credit for everything they produced in concert just because it is standard institutional support to have access to a physicist for any work one wishes to do?

Certainly, in a setting where I have access to a physicist, I can reliably produce work in physics and physics adjacent domains. But there is a meaningful sense in which the institutional support is not just facilitating something I can do on my own, but is adding something to me for which I have no present capacity.

For my part, the physicist would be entitled to attribution and recognition not merely because of their substantive contribution that I could not have made myself at that time, but because of their voluntary assistance in my work (which is why students who support research faculty deserve recognition). Only the former pertains to AI/LLMs.

Kenny Easwaran
1 year ago
Reply to  Kyle S. Hodge

As it happens, I do have a physicist at home (well, a physical chemist) that I talk to about work things on occasion. If that contribution rises to a certain level, it’s absolutely important for his sake that I credit him! (I’m actually a co-author on one of his papers for the converse reason.) But I think that as far as my employer is concerned, there’s no worry about mis-representing my capacity if they hire me and expect me to keep my current capacities (especially since in all our previous hires, we’ve been able to move together and maintain our household abilities).

Old Hack
Old Hack
1 year ago

I’ve done some combination of all of these things for the authors I work with as an editor—though like many of my colleagues, work has slackened since the arrival of these AI tools. For this reason I’m predisposed to dislike AI, but it does seem to me that these sort of uses all have their counterparts in pre-existing arrangements.

Michel
1 year ago

For my part, my personal guarantee is that everything I do is 100% LLM-free.

Until and unless the primary use of these tools is something other than cheating on assessments, I will not revisit my position.

WWW
WWW
1 year ago
Reply to  Michel

What do you mean ‘the primary use’? Other than maybe ‘most salient use for me, psychologically’, it’s hard to see any explicit definition you could give in which that would be true

Nicolas Delon
Nicolas Delon
1 year ago
Reply to  Michel

In terms of either user numbers or volume of queries it’s clearly not true that the primary use of LLMs is students cheating on assessments. What standard would you use to determine when it’s appropriate to revisit your position (which is a fine one to have; you shouldn’t be under pressure to use AI if you don’t want to)?

Kenny Easwaran
1 year ago
Reply to  Michel

As long as you refuse to use the LLM for anything other than cheating on assessments, that will remain its primary use for you! But for other people, who are using it to help draw useful diagrams, to help proofread their papers, to help them read articles outside their specialty, and so on, the primary use might be these other things.

Bruce P Blackshaw
Bruce P Blackshaw
1 year ago
Reply to  Michel

If you’re a student, that’s an admirable stance. For everyone else who is using it for a myriad of other things and has no assessments, why does it matter that some students are using it for cheating? Students use Google to find ways to cheat as well.

Michel
1 year ago

Because of the culture that informs its use. I worry about the twin problems of contagion and erosion; if everyone is using it, then it starts to seem fine or maybe even necessary to use it oneself. And, based on what I’ve observed with students, I worry that the more people use it, the less clear they become about the boundaries of proper and improper use. I do not wish to contribute to this cycle. So: I will continue to uphold the strict standards I apply to my work by doing it all myself.

Which I do better than the chatbot does, anyway. Yes, I’m sometimes slower, but my scholarly output is just fine as it is.

Paul Wilson
Paul Wilson
1 year ago

Tertiary tools are not generally cited in scholarly work:

  • Philosophical abstracts, yearbooks, book reviews
  • Dictionaries
  • SEP and Routledge encyclopedias
  • PhilPapers
  • Google Scholar
  • etc.

Did anyone cite AltaVista or Google? Or their graduate students’ underpaid library research?

How is a LLM index of 10,00,00 plus books and papers sucked down from LibGen and other shadow libraries so radically different?

Besides, why not avail yourself of bibliographic and summative technology? Compete against those who do?

Or use index cards – and good luck with that.

Gorm
Gorm
1 year ago
Reply to  Paul Wilson

Paul
We must live in different worlds – I would certainly cite book reviews (that I discuss), as well as SEP and Routledge encyclopedias. As well, I cite Google Scholar searches in paper, when I cite citation data. I do not understand why you are under the impression that it is permissible NOT to cite these sources. I was an editor for two different journals.

David Wallace
David Wallace
1 year ago
Reply to  Gorm

I don’t think this is what Paul Wilson had in mind. Of course if you get citation data from Google Scholar, you should cite Google Scholar; if you discuss things stated in the SEP article on free will, you should cite that article. But normally: if you used Google Scholar to search for articles on (say) the ethics of surrogacy, and then discussed some of those articles, you would cite the articles but not the Google scholar search; if you used SEP to get up to speed with contemporary issues in free will and then engaged with papers you found that way, you’d cite the papers but not the SEP article.

Daniel Weltman
1 year ago
Reply to  David Wallace

I think most of the people who suggest that you cite an LLM are not suggesting you cite it if it suggests a paper that you read, but rather that you cite it if it e.g. writes a section of your paper or comes up with an argument you use.

Amy
Amy
1 year ago

Note that academics have been paying undergrad research assistants, copyeditors, translators, using librarians, etc for as long as they’ve been writing articles—without attribution.

Kenny Easwaran
1 year ago
Reply to  Amy

Undergrad research assistants are often attributed in the acknowledgments section of a book’s preface! But it’s notable that many of these other people weren’t.

Amy
Amy
1 year ago
Reply to  Kenny Easwaran

often and probably they should be, but it doesn’t seem like misrepresenting one’s work as one’s own not to acknowledge them.

Nicolas Delon
Nicolas Delon
1 year ago

Holding students and teachers or scholars to the same standards is a seductive but ultimately implausible approach for at least one important reason.

The primary good that a class—students and teachers combined—seeks to deliver is learning (also credentialing that is ideally conditional on said learning). The primary good that research seeks to deliver is knowledge or scholarship.

LLMs can enhance the output of scholars or make the process of knowledge production more efficient. Bracketing the important proviso of proper attribution, barring scholars from using AI seems based on a profound misunderstanding of what the constitutive aims of scholarship are; i.e., not credentialing, not citations, not status, not purity of the soul, but knowledge, understanding, and the like.

In contrast, reliance on LLM can clearly (though it does not systematically) undermine the constitutive aims of learning. The reason we hold students to certain standards regarding AI should be that, if they substitute AI use for the valuable process of learning, they are turning coursework into a pointless credentialing exercise. On the other hand, a professor using AI to create course activities, handouts, or exam questions might thereby enhance the value of the product or experience they can deliver to students. It would be antithetical to the constitutive aim of learning to prevent professors from delivering greater value to students just so that they can feel good about prohibiting students from using AI in ways that detract from the same constitutive aim. A teacher may use AI for exactly the reason that a student should not: so that the student can better learn.

For this reason, I can’t see how applying the same standard to students, teachers, and scholars is consistent with the respective aims of their separate (or joint) endeavors.

Kenny Easwaran
1 year ago
Reply to  Nicolas Delon

Exactly.

If I’m the manager of a fire station, and I want to require every firefighter to go to the gym to lift weights at least once a week, I’m going to have very different opinions about their use of machinery to help lift things when they’re in the gym and when they’re in a burning building.

Chris
Chris
1 year ago

For my part, I have been highly resistant to using AI/LLMs in my research. Now, it is interesting to ask an AI to give me comments on something that I’ve written, even though its comments are usually pretty shallow and give me little that is useful. I don’t have any objection to using AI in this way. But when it comes to generating a text, I can’t shake the feeling that it is cheating. Moreover, I like writing papers. I don’t want to outsource the part of the job that I like.

However, I’ve recently decided to loosen up and try using AI to write my abstracts. Why? Because I hate writing abstracts. The amount of time I spend writing abstracts does not correlate to the amount of satisfaction that I get from the task or to my sense of the importance of the task. (Perhaps some people really like writing abstracts. Maybe for them it is a challenge, like writing a haiku. To those people, I say: Good for you! I nonetheless still hate it.) And I really don’t think that the abstract is the make-or-break part of my essay that gets it published.

But maybe using AI to write the abstract is unjustifiable? My willingness to use AI here does not rest solely on my disliking the task. I also think abstracts are not substantive enough to care about. I view them to be about as consequential as the bibliography. And if we can use AI to generate our bibliographies, then why not abstracts too?

WWW
WWW
1 year ago
Reply to  Chris

Nope – not unacceptable. It’s just up to you to decide when an AI-drafted abstract is appropriate for your paper.

David Wallace
David Wallace
1 year ago
Reply to  Chris

I think it’s extremely unwise to have this attitude to abstracts. They might not be that crucial to whether you get published but they’re critical to whether you get read!

That said, if you think AI writes better abstracts than you, go for it. (But I also feel that about all other bits of the paper.)

Richard Hanley
Richard Hanley
1 year ago

I suggest adding one more item to Justin’s list, one where I think AI could be very helpful. It’s adjacent to: “search through articles and other resources to see which will be useful for a specific project.” Call it “a comprehensive check to see if you’re reinventing the wheel.” In my view it happens too often in professional philosophy that someone publishes a “new” argument or analysis or account or whatever that is not so new; they are just unaware that someone else has beaten them to all or most of it. Not apportioning blame here… but having AI do the drudge work to avoid it seems to me highly desirable.

Cassandra Veritas
Cassandra Veritas
1 year ago

Just to share some practical observation from using these systems: paradoxically, the more “advanced” a chatbot is, the less reliable it often is at tedious, high‑precision tasks like checking bibliographies, line-by-line proofreading, or indexing. (I’m also worried about the “stylistic edits” suggestion mentioned in the article, since these models don’t distinguish between what’s merely stylistic and what’s making a substantial claim.) These models are optimized to produce plausible continuations. Verifying each token against a source simply takes too much energy for that. The output can look exactly like what a careful human would say, while skipping the actual verification.

Of course, this can be overcome by very careful prompting and detailed instructions. But at least from my experience by the time I’ve specified constraints and audited the results, the time savings are almost negligible.

A more worrying consequence of this is that the skills it takes for “prompt engineering” to coax reliable work from LLMs seems very different from the kind of skill-set to make a good philosopher (hopefully the definition of what a good philosopher won’t change much in the age of AI).

Gorm
Gorm
1 year ago

Thank you Cassandra.
If I may quote you: “These models are optimized to produce plausible continuations.”
That is so nicely said. Wake up everyone.

Kenny Easwaran
1 year ago

One thing I’ve found in some of my recent use is that these models are often quite “lazy”! I’ve been experimenting with various LLMs to see how well they can solve crossword puzzles. I’ve found they’re far, far beyond my ability in terms of looking at a clue and guessing the correct entry – I upload the .pdf of a crossword puzzle, and the LLM says “oh, that’s nice, here’s the answers to 1/3 of the clues, and I’ve already figured out what the four theme entries have in common”. But the ones that it gets wrong very often have the wrong number of letters in them, and it has a very hard time looking at the image of the grid and identifying where the black and white squares are.

I spent some time using the LLM to write an interface for it, so that it can have a precise representation of the grid in a machine-readable form, and can query the grid to identify what letters are present from the crossings. But it’s only about 85% reliable at putting new answers into the correct squares of the grid.

And the “lazy” part – I’ve sometimes discovered halfway through that it didn’t even write down all the clues when I uploaded the .pdf! It just wrote down 2/3 of the across clues and five of the down clues and figured I wouldn’t care that it didn’t get around to doing the rest! And on other occasions, when I’ve given it big lists of phrases and asked it to see which ones have fun wordplay opportunities with homophones, even if I remind it to check every single word in every single phrase for homophones, after it does the first five or six, it gets “lazy” and starts only checking one word for homophones.

What I really need for these purposes is a traditional symbolic computer program that does all the boring bits of going through the list and isolating each word, and then calls the LLM for one small question at a time, rather than asking the LLM once to do all the items in the list.

Kenny Easwaran
1 year ago

> (hopefully the definition of what a good philosopher won’t change much in the age of AI)

I do think it will change though – it has changed with every relevant change in technology and social norms. Now that we have easy access to printed editions of lots of other people’s work, a large part of being a good philosopher involves being familiar with that literature, and adequately citing it. This obviously wasn’t an important part of being a good philosopher in most other centuries.

Similarly for the importance of different forms of clear expression, in formats like 20-page papers, 1-hour talks, Q&A periods and the like, that were less relevant for other eras.

Daniel Weltman
1 year ago

I’m not very interested in the permissibility question. I think a lot of stuff that one ought not to to do is still permissible in some minimal sense. It’s permissible to go to the park and tell children that someday their parents, their friends, and they themselves will die, but you shouldn’t do it. If that’s what you really think you ought to be up to, I don’t have anything to say about why you must stop. I do however think it would be better if you stopped. I feel similar about most AI use in philosophy.

My feelings stem mostly from seeing the sorts of ways people defend AI use, the sorts of results people seem proud of achieving via AI use, and concerns about what the world will look like (and to a large extent already does look like) given widespread acceptance of copious AI use. It also stems from worries about what things will look like once we have a generation of philosophers using AI who didn’t learn about philosophy sans-AI.

Although once in a while someone will flip one from side to another, I don’t think my concerns will be very compelling to those who use AI. So I think mostly what is going to happen is that we will end up with another divide in the profession, like the one separating those who think it’s very important to cite a lot vs. those who think you can write a 10,000 word paper with 2 or 3 citations and that’s no big deal; those who think philosophy must be empirically informed vs. those who think it can or must be done entirely from the armchair; those who think philosophy must be historically informed vs. those who think everything between Plato and Rawls is mostly a catalogue of silly arguments hardly worth entertaining; those who think morality is relevant to political philosophy vs. those who don’t; and a dozen other divides I’m sure you can name (many of which are three-sided or four-sided instead of two-sided).

I can tell you what part of the AI divide I am on and am likely to remain on, for the foreseeable future if not for my entire life. But I’m not sure there is much I can say to get someone to cross from one part of the divide to the other.

Kenny Easwaran
1 year ago
Reply to  Daniel Weltman

A relevant question to consider here is whether the existence of these divides is good, bad, or neutral. There are lots of reasons why it’s good for our intellectual communities to contain some sub-communities doing things one way and others doing them another. What it means that these are sub-communities is that there is less communication between them than within them, just like with distinct academic disciplines. As long as there is some communication across these boundaries, and some abilities of recognizing when some work done on the other side would be worth engaging with on this side, I tend to think the good outweighs the bad in most of these cases.

Jimmy Lenman
Jimmy Lenman
1 year ago

AI is an increasingly toxic and culturally corrosive technology which should be deplored and shunned.

And we should be very careful about what we choose to deem permissible. It’s one of those areas where whatever we decide to make permissible we end up making compulsory. It’s like performance enhancing drugs in sport. In a competitive pursuit if your rivals are allowed to use them you can’t afford not to use them too if you are serious about staying in contention. You might reply that philosophy is not as brutally competitive as athletics is. But especially for early career academics launching themselves on an ever more desperate job market I’m afraid it often is.

Nicolas Delon
Nicolas Delon
1 year ago
Reply to  Jimmy Lenman

Do you think this of all AIs or just LLMs? Do you object to using AIs like AlphaFold to predict protein structures too? Should we pass on scientific breakthroughs that could save lives because they were AI generated, instead leveling down the playing field so all humans can compete fairly? I’m not sure what a good equivalent of protein folding would be in philosophy, but there’s a point where a finding is important enough that I don’t care who produced it.

J_B
J_B
10 months ago
Reply to  Nicolas Delon

The technology is here. Therefore it will be used and it will reshape the field. This is one of the most irrefutable lessons of history.

Michael Gorman
Michael Gorman
1 year ago

This is a complicated discussion, and what I want to say right now applies to only some aspects of it.

It strikes me that sometimes, what people say is that AI use can be good in philosophy because it allows us to produce more. But I’m not sure that philosophy should be thought of as a productive activity. The point of being a philosopher isn’t, it seems to me, to produce philosophical products, but instead to *do philosophy*, sometimes alone and sometimes with others–by which I mean other humans.

Cars go faster than people, and forklifts lift more than people, and yet we still have the Olympics. That’s because we value human activity as a good in itself, apart from its products.

For reasons along these lines, I don’t want to use AI, and the more others’ work was generated with the help of AI, the less interested I am in it. I think philosophy is a human activity for humans, and I think we should preserve it that way, in as pristine a form as we can, even if it’s less productive or efficient than AI-powered philosophy.

Daniel Weltman
1 year ago
Reply to  Michael Gorman

A similar point (or maybe another way of putting your point) is that I think one of the main benefits of doing philosophy is that it is a process of inquiry through which we discover who we are. The process is also partially constitutive of who we are. The way that process looks when it involves LLMs, and the conclusions the process arrives at, and the sorts of people we find ourselves to be and (partially) become when we involve LLMs in this process, are worth contemplating, specifically as versus what we look like when we do philosophy without LLMs.

David Wallace
David Wallace
1 year ago

I think this whole discussion illustrates a deep divide between philosophers as to what academic philosophy research is for.

Justin could have asked: what are the ethics of using AI in the search for cancer treatments, or in hurricane forecasting, or in looking for high-temperature superconductor candidates, or in mapping structure formation in the early Universe or the evolutionary patterns of dinosaurs. But I think most people would agree that the issues here begin and end with how effective AI is in those research projects (and perhaps externalities like the energy cost of AI development): there wouldn’t be some sense that somehow the research is intrinsically less valuable because it wasn’t done by unaided humans. The point of this research, and the point of funding it publicly, is to deepen our collective understanding of the subjects being researched; the criterion for whether a given method or tool should be used is just whether it leads to better research outputs (and, again, possibly external ethical constraints, as in animal experimentation).

One view of philosophy research is that it is essentially like scientific research. If so, there would be no special ethical problem in using AI in philosophical research, and there would be no reason to think the reasons to be cautious about AI in teaching also apply in research. (No more is there an inconsistency between restricting the use of graphing software in Calculus 101 and permitting it freely in physics research.) I hold this view.

But there is an alternative view where the purpose of philosophy research is less about the output, more about the human activity of producing that output. If you hold that view, as some people on this thread clearly do, then it makes much more sense to restrict or even proscribe AI in philosophy research. (Though it is less clear why public money ought to be spent paying for that research.)

Justin Weinberg
1 year ago
Reply to  David Wallace

“I think this whole discussion illustrates a deep divide between philosophers as to what academic philosophy research is for.” I agree! AI, in both its current form and with its imagined future capabilities, raises interesting questions for us about why we are doing philosophy. For reasons I can’t go into now—today’s my first day back at teaching, and I’m busy with teaching-related stuff—I think it puts some pressure on “essentially like scientific research” views. (This was the subject of a talk I gave earlier this month. Perhaps I will put up a post about it when I have a chance.)

Eric Steinhart
1 year ago

Yes, there should be a post about this.

Eric Steinhart
1 year ago
Reply to  David Wallace

Right – if philosophy research is more like research in science (or math), then using AI to do it is acceptable and maybe obligatory.

But I don’t think philosophy research is like research in science (or math). And, more generally, I don’t think philosophy is much like science at all.

Unfortunately, lots of philosophers do think philosophy is like science. Or is the hand-maiden of science, or something like that. This view is becoming disastrous for us.

So what is philosophy like? It’s more like art.

If that’s right, then there are reasons to limit AI use in philosophical research. But then philosophical research would change into something very different. We’d be producing far more creative work in many media. We’d be far more engaged with culture. And we certainly wouldn’t be writing the way we write.

Regardless of our beliefs here, AI might push philosophy in a more artistic direction.

David Wallace
David Wallace
1 year ago
Reply to  Eric Steinhart

“Unfortunately, lots of philosophers do think philosophy is like science. Or is the hand-maiden of science, or something like that. This view is becoming disastrous for us.”

I think it depends who “us” is. I think the philosophy I do, and the philosophy good people in my subfield do, is like this: I think it’s meaningfully and increasingly integrated into the scientific enterprise in productive ways. But plausibly not all – maybe not even most – good philosophy research is like that. And maybe that other style of philosophy needs to be thought of in fundamentally different ways.

Eric Steinhart
1 year ago
Reply to  David Wallace

Relative to your work, I think this is a very interesting issue. I’m familiar with some of your papers, and they often seem to be highly scientific — you’re just doing science. Are you also doing philosophy?

For instance, your paper “Gravity, Entropy, and Cosmology: In Search of Clarity” just reads like excellent science. Do you think there’s any philosophy in that paper? Why?

I don’t think I’d endorse the trivial thesis that any reasoning or argumentation is philosophical.

I’m very curious what you think about this in terms of your own work.

David Wallace
David Wallace
1 year ago
Reply to  Eric Steinhart

I think a large part of philosophy of physics, including most of my work, is interdisciplinary: it’s on the boundary between philosophy and physics, and there are no sharp (non-institutional) ways to draw the line. (And the existence of any distinction between physics and philosophy is of comparably recent vintage, 18th century at the earliest I think.)

Modern philosophy of (contemporary) physics draws on ideas from fairly mainstream philosophy – probability, emergence, necessity, functionalism, underdetermination, realism, the mind-body problem, metalinguistics – but also draws heavily on the technical content of modern physics. Different bits of it draw to differing degrees from philosophy and from physics (for myself “in search of clarity” is about as close as I get to physics proper; my work on structural realism or decision theory is about as close as I get to philosophy proper).

I think fairly similar things could be said about philosophy of mind and of cognitive science; my work is about as philosophical as Dennett’s.

In that particular paper of mine I’m trying to point out a widespread misconception in both the physics and the philosophy-of-physics literature, and explain a better way to think about the issue. The ‘tell’ that institutionally it would struggle to count as physics is that it doesn’t carry out any novel concrete calculation or make any novel empirical prediction: it’s analyzing the conceptual content of extant physics, which the norms of 21st century Anglosphere academia treat as foundations of physics or philosophy of physics, or perhaps as physics pedagogy. I couldn’t have got the paper into Physical Review, but I could probably have got it into Classical and Quantum Gravity; the choice to publish in BJPS instead was partly for institutional reasons (physicists are less concerned with where things are published when deciding whether to read them; it’s more useful on my CV to have a BJPS paper than a physics paper).

Eric Steinhart
1 year ago
Reply to  David Wallace

Ok, sure, I blame society too. But that’s not satisfactory.

Given that your work really does straddle the science / philosophy divide, it’d be great to hear you give some conceptual attention to that divide.

And perhaps thereby answer the question: is philosophy a continuation of science or not?

Of course, that deserves an article, but it would be one I’d happily read.

David Wallace
David Wallace
1 year ago
Reply to  Eric Steinhart

I wasn’t intentionally blaming society for anything.

As to the broader question: I don’t think it can seriously be disputed that up to the era of Newton, Galileo, and Descartes, there was no particularly coherent distinction drawn by scholars of the time between physics and philosophy. By the start of the 20th century there definitely was such a distinction; nonetheless, leading physicists like Einstein and Bohr drew deeply on philosophy. By the postwar period those connections between philosophy and physics were much more attenuated.

In light of this history: one might think that the ‘natural philosophy’ of the 17th century encompasses both philosophy of physics and physics proper, but can be separated from a more humanistic conception of philosophy that includes (say) ethics and aesthetics. (And one might then ask where, say, causation or metalinguistics lie.) Or one might defend a science/philosophy distinction that classifies philosophy of physics as part of philosophy (and perhaps tries to read that back into the 17th century, so that, say, the Scholium to Principia is philosophy of physics while the rest of Principia is physics). Or one might just adopt a Quinean continuum view.

I am moderately interested in that question (and lean towards the Quinean view). But ultimately answering it does not play to my comparative advantage and I am more interested in continuing to work on the first-order questions that interest me, and on connecting those first-order questions to work by other scholars both in physics proper and more mainstream philosophy.

Put another way: the question ‘is philosophy of physics really philosophy, or just physics?’ strikes me as fairly analogous to the question ‘is biochemistry really biology, or just chemistry?’ – that is, not of zero interest, but not something one needs to answer in order to do worthwhile biochemistry.

Eric Steinhart
1 year ago
Reply to  David Wallace

Well, you did talk a lot about institutions, and their social forces driving you this way and that. So that’s blaming society.

And sure, maybe the question ain’t all that interesting. So back to the OP:

Is what you’re doing just wide open to AI automation?

David Wallace
David Wallace
1 year ago
Reply to  Eric Steinhart

The operative word is “blaming”. I don’t regard it as particularly bad or blameworthy that it’s socially contingent which work gets done in which department.

As to the question: no idea. But if it turns out that my academic work becomes doable by AI better than I could do it myself, I would seek to do something else with my time rather than defend the value of my research being done by *me* rather than an LLM.

(And: if AI can do interdisciplinary work in philosophy of science better than humans, it can probably do science better than humans too. At which point, the world will change radically, in ways that are hard to predict.)

Eric Steinhart
1 year ago
Reply to  David Wallace

You wrote: “I would seek to do something else with my time”. What? That’s the exactly question at issue here. We are all concerned with this issue.

Would that “something else” be philosophy? And, if so, how would you conceive of it? Especially since that “something else” would be resistant to AI automation.

Given the rapid progress of AI, I submit that this question is urgent for you, as for us all.

I think we would all be interested in your answer.

David Wallace
David Wallace
1 year ago
Reply to  Eric Steinhart

If AI develops to the point that it can do scientific research more effectively than humans (which on balance I would bet against happening in the short- to medium-term), that is pretty much the definition of a technological singularity.

I have no real idea what things would be like within a few years of this happening, and neither does anyone else. Insofar as I feel like speculating, the interesting questions are things like “will there still be any such thing as work or the economy?” and “will we all die?”, rather than parochial questions of what I in particular will do instead of philosophy of science.

Kenny Easwaran
1 year ago
Reply to  Eric Steinhart

I think art is a useful comparison here! Again, there’s a divide between those who think that there is value in artworks, as a product that many people can contemplate and engage with and get aesthetic understanding or pleasure or other goods from; and those who think that the value of art is primarily the process for the creator.

You don’t have to think that aesthetic value is all that much like scientific value to have the former point of view. But if you have that point of view, then something that helps you produce more of it, that enables audiences to get more value out of it, would be good.

Ben M-Y
1 year ago
Reply to  David Wallace

I agree (and think it’s very perceptive and helpful to note) that this discussion nicely illustrates the deep divides that already exist around the question what academic philosophy research is for. And I agree with other commenters that it would be good to have a discussion about this. (The paper of his that Justin refers to may be a nice starting place!)

I would like to add, though, that I do not see the discussion of the aim of philosophical research as separate from some of the other very important considerations that have been adduced in this thread. This is worth making explicit, I think, because they can appear unrelated. Yet I think it would be a mistake, both practically and theoretically, to separate them.

One issue is about consistency (or non-hypocrisy). I am sympathetic to the claim that we shouldn’t engage in practices that we prohibit our students from engaging in. As has been mentioned, all of this should be contextualized relative to appropriate aims. But I think the teaching-research distinction is often overblown. And I think it will seem especially artificial given the view that philosophical research is less about output and more about activity. So, concerns about what we require of our students and what we require of ourselves may, in the end, turn on one’s view about the aim of philosophy research.

Another issue is about social impacts, especially ways in which practices shape the future. I am of the view that the main benefits of using generative AI tools (at least in philosophy, as opposed to other areas like cancer research) are ones of efficiency. These tools allow us to do things we can already do (check the lists above) faster and easier. Crucially, they also allow us to do them in isolation–that is, alone and without the input of other human beings. So, use of these tools reduces social interaction in the context of philosophical research. This is one tradeoff that needs to be weighed against the efficiency gains. Perhaps this will be especially worrisome to those who take the activity view, but it seems possible that it may be a worry even on the product view–for example, it may be the case that more human-to-human interaction is more productive, in the long term, than less human-to-human interaction.

Use of AI tools also reduces opportunities for paid work and the opportunity to be socialized into various research practices. These are benefits, say, to graduate research assistants, that will no longer be had by any human beings if we offload these tasks onto AI tools. The loss of such benefits should be weighed against the efficiency gains. Once again, the underlying view about the aim of philosophy research may shape how one views these tradeoffs. But many if not all of the same tradeoffs will be there to consider no matter which view one has of the aim of philosophy.

There’s more to say, but I’ll leave it at this. I think this is a good and important discussion to have. I just want to highlight that it is also a complicated one, and it may be tempting to cleave off considerations that at first glance appear irrelevant but are in fact very much relevant.

David Wallace
David Wallace
1 year ago
Reply to  Ben M-Y

“These tools allow us to do things we can already do (check the lists above) faster and easier. Crucially, they also allow us to do them in isolation–that is, alone and without the input of other human beings. So, use of these tools reduces social interaction in the context of philosophical research. “

That’s fair – but AI is scarcely the first time this has happened. Using word-processing to type a paper or book yourself rather than getting a secretary to type up your handwritten manuscript; downloading a paper from a journal rather than calling up the bound volume from the library stacks and photocopying it; (in the sciences) doing a calculation via computer rather than via a team of assistants laboriously iterating an algorithm, etc. There are genuine tradeoffs in each case, but in none of them would I want to go back.

Ben M-Y
1 year ago
Reply to  David Wallace

Also fair, up to a point. But I’m not sure that you’re comparing apples to apples. Many of the things on the lists (e.g., brainstorming or coming up with counterexamples) arguably have benefits for those engaged in the activity, on top of any benefits in terms of products and employment. The examples you give (calculating an algorithm or typing a handwritten manuscript) seem like pure busy work, with all benefits being in terms of the end products or in terms of employment.

Will Fleisher
Will Fleisher
1 year ago

I worry that widespread use of AI for philosophy research has the potential to create significant negative epistemic effects, even if some people produce better research using the systems. Basically, I’m worried that algorithmic monoculture and homogenization effects will lead to subtle biasing of philosophical research in a way that won’t be noticeable right away but might still be really bad, taken as a whole over time.

The worry is that AI systems are biased in various predictable ways. They make mistakes that are systematic rather than noisy. Of course, that by itself doesn’t distinguish them from humans. But if everyone is using the same few models, models which are built using basically the same architecture and training methods, then the biases influencing the research will typically be the same in every paper. Which biases a particular paper displays will thus itself be systematic, rather than noisy.

The worry is that research will be subtly guided in ways that are influenced by the AI system’s biases, and that this will be very hard to tell when just looking at an individual paper. Any individual paper may be no more likely to be biased or mistaken than usual – though to be honest, I would bet against this if we are only talking about current models – but the overall effect could be pernicious.

The point here isn’t just that research would have more prejudice, though it might. Rather, having different biases and perspectives is an important part of getting the benefits of diversity and disagreement within a research field. If we smooth out all the differences in perspective, we will lose these benefits.

I’m not saying this by itself shows that using AI for helping to write paper is thereby impermissible. But it is a serious pro tanto reason against using them, I think.

Nicolas Delon
Nicolas Delon
1 year ago
Reply to  Will Fleisher

Very interesting. Thanks for sharing. Tangentially, this reminded me of recent essays by Jac Mullen on parallels between the genesis of writing (as a state coerced means of accounting and legibility) and the rapid spread of AI. Mullen ends on a relatively optimistic note—just as the origins of writing didn’t constrain what we could use it for, the current algorithmic homogenization doesn’t have to constrain how we could harness AI for decentralized, creative work. The emphasis is on *rapid*, so I share your concern.

Here’s the first essay; he’s since followed up with a couple more on his Substack:
https://substack.com/@jacmullen/p-162377104

Daniel Weltman
1 year ago
Reply to  Will Fleisher

I think this is a serious consideration and one that is important enough to suggest at the very least extreme caution when using LLMs (caution far in excess of what I see most people exercising). But, similar concerns apply to most reasoning, so I wouldn’t want anyone to think that thinking solely with one’s brain renders one immune to this sort of stuff. Humans in various contexts are biased in various predictable ways too, after all.

I do think we have ways to help work around this problem when it applies to humans (e.g. make sure the field comprises people of diverse backgrounds) that we probably don’t have when it applies to LLMs, so ultimately this is an anti-AI consideration. But without due attention to how humans work, it can be an anti-human consideration too.

Paul Wilson
Paul Wilson
1 year ago

ChatGPT-5 (prompted by Kelly Truelove of TrueSciPhi.ai) responds to Justin: https://open.substack.com/pub/truesciphi/p/philosophy-with-stochastic-machines

74
0
Click here to commentx
()
x