Publishers Want To Sell Companies the Right to Train AI on Your Books: Should You Consent?


Should authors consent to have their publishers grant licensing requests by firms and projects to allow them to train their generative AI on their books?

[Laura Owens, untitled (2016)]

That question was suggested by Elliott Sober (Wisconsin), who is curious what philosophers think of the issue.

It’s worth noting that not all publishers are asking authors for consent. As I reported here last July, Taylor & Francis, which owns Routledge, decided to provide Microsoft (who is a partner with Open AI, the makers of ChatGPT) with “non-exclusive access to advanced learning content and data [that is, the books they publish] to help improve relevance and performance of AI systems” without bothering to seek permission from authors.

But some publishers have been asking for permission. Should you grant it? Should a publisher’s policies on the matter affect whether you publish your book with them? Your experiences with this issue and your thoughts on it are welcome.

guest

16 Comments
Oldest
Newest Most Voted
Inline Feedbacks
View all comments
Hmm
Hmm
1 year ago

Insofar as the threat of replacement by AI really does apply to those in the philosophical profession, consenting to have such programs feed on our original writings strikes me as a perfect example of trees voting for the axe.

Arch Lancaster
Arch Lancaster
1 year ago

Suppose you worked for years and got a paper in the Phil Review (or some other top five) and you were then told that the publishing agreement required allowing some AI to be trained on your paper. Would I say yes? I would! The AI can train itself silly on that paper; the alternative would be to not get an earned top 5. Now, that said, if the issue was merely that you had to agree to have the submission trained on AI in order to submit the paper in the first place… then I’d still say ‘yes’ (if I had a good reason already to submit there). Those however are the only two conditions under which it would make sense to say ‘yes’.

Helen Dunn Frame
Helen Dunn Frame
1 year ago

Hell No.

Noah
Noah
1 year ago

Should we allow the blood-sucking publishers, who reap all the economic value of our work for free and gate the information we produce behind paywalls, to also sell this work to AI companies for the generous compensation of $0.00? Obviously not, independently of all considerations about AI.

etum
etum
Reply to  Noah
1 year ago

Seems quite odd to voluntarily give your work to a blood-sucking publisher who reaps all its economic value, and to do so for free. Perhaps you are getting some value out of it after all?

Meme
Meme
Reply to  etum
1 year ago

Boy, I wonder if Noah wasn’t talking about the other ways publishing might be valuable, and that he had economic value in mind when he said… “economic value.”

etum
etum
Reply to  Meme
1 year ago

I was thinking more of the “for free” part.

Love
Love
1 year ago

If it were for a free and open platform like a public ai library, commitment to democratization of knowledge a shared power – a thousand time yes!

Nate Sharadin
1 year ago

I’m very happy for any of my work to be used without any restriction by any entity that can find a use for it.

Tom Cochrane
1 year ago

Last month Cambridge requested AI rights for my 2018 book and after some deliberation I agreed. My reasons were that the book is already in the public domain, and the agreement specifies that any use of my work would be cited. (It also mentions royalties but I won’t hold my breath). Apart from that, I figured that I’d prefer an AI is able to accurately report what I’ve written than that it bullshits about it.

For reference, here are the most important terms of the agreement:

DEFINITION OF AI SUBSIDIARY AGGREGATION
For the purposes hereof, “AI Subsidiary Aggregation” shall mean a new type of subsidiary use where rights in the Work (in whole
or in part) are granted to a third-party licensee by Cambridge in order to permit the licensee to include, excerpt or adapt the Work
in an AI content aggregation engine or LLM (Large Language Model), for public or private distribution, including where that engine
or LLM contains work published by both Cambridge and other publishers.

2 REMUNERATION FOR AI SUBSIDIARY AGGREGATION
Remuneration to the Author for use of the Work under an AI Subsidiary Aggregation licence shall be a royalty rate of 20% and paid
on a case-by-case basis, allocated to the Author from (and where applicable, as a pro-rata share of) the Net Revenue derived from
such licence.

3 MORAL RIGHTS
Nothing herein shall affect, reduce or amend the assertion of moral rights by the Author pursuant to the Original Agreement. Any
AI Subsidiary Aggregation licensing will be conducted by Cambridge in compliance with such assertion and the Work shall be cited
accordingly.

Derek Bowman
Reply to  Tom Cochrane
1 year ago

“I’d prefer an AI is able to accurately report what I’ve written than that it bullshits about it.”

Assumes facts not in evidence.

Reinhard Muskens
1 year ago

So that the AI can reuse your ideas without any citation?

Marketeer
Marketeer
1 year ago

Not to be a huge downer, but I’m not sure how much it matters in practice (which is what I care about). Chances are very good that one’s book will end up on Libgen, Z-Lib, or some such place, which AI companies are relentlessly scraping.

Kenny Easwaran
Reply to  Marketeer
1 year ago

I think they were scraping that sort of site heavily a couple years ago when they were getting started. But now that they’re big enough that they’re starting to sue each other, I believe they’re trying to be a bit more careful. Illegally-scraped YouTube videos were in all the original training sets – but if OpenAI uses illegally-acquired YouTube videos in their next model, you can be pretty sure that Google will sue them (and similarly if Google uses any Apple TV or Netflix shows in the next version of Gemini, or if any other publisher is teamed up with one of the AI companies rather than another).

David Austin
Reply to  Marketeer
1 year ago

A long thread at https://social.kernel.org/notice/AqJkUigsjad3gQc664 provides evidence that web-scraping by AI-bots generally is a serious problem because of the load on websites scraped. It’s said that the AI-bots typically ignore robots.txt files.
[One posting in the thread says,

Anthropic has been the worst offender by far; everyone else seemed to at least rate limit, but Anthropic seemed determined to destroy the internet.

Amanda Askell, Anthropic’s chief ethicist, was interviewed by Lex Fridman among others from Anthropic. More here and here. Her PhD (2018) is from NYU. Perhaps she could be asked Sober’s question.]
Nepenthes is a tar-pit for AI-training bots. (See the 404 Media account.) It is being discussed across many fora for software engineers. (I found the long thread cited above in searching for information about Nepenthes.) However, one of the warnings from the developer of Nepenthes says that using it at a site will almost certainly result in a site’s being removed from search engine results. I’d think that would make using it unattractive for many site maintainers.
Cloudflare has an option to block AI-bots. See, e.g., https://arstechnica.com/tech-policy/2024/09/cloudflare-lets-sites-block-ai-crawlers-with-one-click/.
But AI-bot scraping seems to remain a serious problem.

Peter Suber
1 year ago

MIT Press asked its authors this question in Nov 2024. I’m a philosopher with two books there, tho neither one is in philosophy. After I sent my answer to the press, I posted it as a Mastodon thread. You won’t need a Mastodon account to read it.
https://fediscience.org/@petersuber/113443473594224752

Last edited 1 year ago by Peter Suber