In this era of AI, “can we still rely on take-home writing assignments to assess student learning? And, should we allow students to use ChatGPT in order to complete such assignments? My answer to both questions is ‘yes’.”
That is Carlos Zednik, assistant professor of philosophy at Eindhoven University of Technology and director of the Eindhoven Center for Philosophy of AI.
He has observed AI-assisted philosophical writing in hundreds of students, developed a version of prompt-grading as an assessment method, and gathered some evidence that shows that “prompt grades appear to be sufficiently valid as a measure of philosophical knowledge and argumentative skill, at least insofar as they track more traditional measures such as essay grades.”
His conclusion is that philosophical essays can still be valuable pedagogical tools if they are assigned and graded by professors who understand the technology their students are using.

Grading Prompts to Measure Student Learning
by Carlos Zednik
Ever since its release in November 2022, university instructors have worried about ChatGPT’s impact on higher education. The situation is particularly worrisome in disciplines like philosophy, which traditionally rely on take-home writing assignments to promote and assess student learning. Students are expected to write in order to acquire philosophical knowledge and argumentative skill, and teachers are expected to read their students’ written products in order to measure their learning progress. However, now that large language models can be prompted to generate argumentative essays, ethical analyses, and reading-responses, commentators have begun to question the value of asking students to produce these kinds of written products—and of asking teachers to grade them. Is the college essay really dead?
Let’s get something out of the way first: There can be no doubt that generative AI technology has reached a level of sophistication at which students can use it to produce excellent pieces of philosophical writing. Long gone (on AI time scales) are the days when AI-generated essays could be reliably distinguished from mid-level undergraduate work. Moreover, recent releases of “reasoning models” by OpenAI, DeepSeek, and others—whose internal processes can be visualized in “chain of thought” mode—reveal sophisticated patterns of argumentation, analysis, and knowledge-integration that rival those of advanced graduate students, and perhaps even of professional philosophers. Although these systems’ written products are fallible and prone to hallucination and bias, they are arguably no more imperfect than their human-produced analogues.
Just as students would appear foolish not to take advantage of machines that can do their homework for them, teachers would appear to be selling themselves short by tediously commenting and evaluating hundreds of pages of botshit content every semester. So, can we still rely on take-home writing assignments to assess student learning? And, should we allow students to use ChatGPT in order to complete such assignments?
My answer to both questions is “yes”. Indeed, for just over two years, I have permitted unrestricted use of AI in my own Bachelor and Master’s-level philosophy courses at Eindhoven University of Technology, without worrying that students are failing to acquire relevant knowledge and skills, and without feeling that my efforts are lacking in value. On the contrary, by allowing students to use generative AI when they write, I believe to be providing them with a uniquely appropriate kind of preparation for their future lives and careers in an AI-assisted world. At the same time, by asking students to keep a detailed record of their AI-use, I have a privileged glimpse into a new but increasingly commonplace way of working, creating, and thinking.
Whereas traditional assessment methods in university-level philosophy courses center on the written product—the completed essay, report, or analysis—I now instead focus on the writing process—the way in which students interact with generative AI systems such as ChatGPT to complete a piece of philosophical writing. A student’s interaction with generative AI centers on their use of prompts: inputs that are designed to maximize the correctness, relevance, effectiveness, beauty, desirability, or (more generally) appropriateness of the AI system’s outputs. Of course, “prompt engineering” is among the buzziest of buzzwords on YouTube, TikTok, and LinkedIn—all of which feature lengthy tutorials on prompting frameworks such as RISEN (Role, Instruction, Steps, End goal, Narrowing). If mastered, such frameworks can help students prompt ChatGPT to produce text that satisfies the constraints of any assignment description (or, for that matter, to complete almost any other task that would traditionally be completed by a human working on a computer).
I have no intention of evaluating students’ mastery of prompting frameworks—I teach philosophy, not prompt engineering. That said, what regularly gets lost in discussions of generative AI and its impact on the classroom is the fact that good prompting is not merely a technical skill, captured in the principles of a domain-general prompting framework, but that it also requires domain-specific knowledge to precisely define and narrow the context that constrains an AI system’s outputs. Indeed, specifying context in prompts is critical for generating content: ChatGPT produces exactly those outputs that it estimates to be the most likely to occur within the context of its inputs. Insofar as likelihood is a proxy for appropriateness, the more narrowly the context is specified, the more appropriate the generated outputs are going to be. Specifying the right context for a particular task—writing, drawing, video-making, computer programming, tax-filing, or restaurant-reservation-making—is tricky, and may require a human prompter to appeal to prior knowledge, deploy domain-relevant skill, and to critically, creatively, and repeatedly engage with a system in order to evaluate, select, and tweak the content being generated.
Consider two illustrative examples. Although influential AI artists such as Refik Anadol use state-of-the-art image generators such as StableDiffusion or MidJourney to produce compelling art, they do so through an iterative process of prompting, selecting, cropping, editing, and re-prompting that requires a considerable amount of experience, skill, and aesthetic sensibility. Similarly, although AI-assisted vibe-coders allow generative AI systems to take care of the syntax of a computer program while they focus on defining the task specifications, supervising the logic, and ensuring the correctness of the program’s semantics, these latter kinds of activities still benefit from programming experience, domain knowledge, and coding ingenuity. The point is not that AI systems cannot be prompted to generate art or computer code with relatively little effort, but rather that (assuming similar levels of prompt-engineering mastery) domain-experts are likely to generate better final products than novices. The reason for this is that experts are better able to define and narrow the context within the prompts that are used to generate content.
AI-assisted philosophical writing is no different. Whereas the use of domain-general prompt-engineering principles may allow an ill-informed and disengaged student to generate a passable piece of philosophical writing solely on the basis of their prompt-engineering chops, that piece will compare unfavorably to one that is generated by a student who suitably narrows the context by introducing relevant philosophical knowledge and argumentative skill into their prompts.
Let us look a bit more closely at the process of AI-assisted philosophical writing as I have seen it unfold in the prompts of several hundred Bachelor and Master’s students in Eindhoven. Almost all students use generative AI to research ideas as well as to summarize lecture slides, news articles, and assigned readings. They prompt ChatGPT to brainstorm possible solutions to moral problems, and to come up with counterexamples and objections to philosophical claims. They may of course ask the system to generate arguments in support of a particular claim, or to provide empirical, textual, or logical evidence in support of a particular premise. In their interactions with generative AI, students also often introduce technical course concepts from the lectures or readings, and ask the AI system to reflect on, define, or apply these concepts. They might also prompt a system to analyze a case study from the perspective of a particular ethical theory, or to apply the distinctions or frameworks championed by a particular philosopher. Of course, after having done all of these things and more, students prompt the system to generate complete sentences, paragraphs, or sections which they go on to copy-paste, review, tweak, and ultimately submit.

I don’t consider these interactions to be evidence of cheating. Rather, I take them as demonstrations of the kind of knowledge and skill that is required to build a compelling philosophical argument or ethical analysis, albeit with the help of an AI-powered computer. A student’s ability to produce high quality work depends not only on their mastery of RISEN, but also on their ability to deploy philosophical knowledge and argumentative skill in their prompts. By inspecting these prompts for evidence of such knowledge and skill, teachers can assess their students’ and assign a meaningful grade.
How effective is prompt-grading as an assessment method? As we report in a current preprint, the initial results are promising. For one, students who demonstrate philosophical knowledge and argumentative skill in their prompts produce better philosophical writing (i.e., they receive a better essay grade) than their peers who do not. Moreover, they produce better philosophical writing than students who only or primarily employ domain-general prompt-engineering principles, and do better than students who who merely use AI to improve linguistic “surface features” such as spelling, grammar, word choice, and style. Thus, prompt grades appear to be sufficiently valid as a measure of philosophical knowledge and argumentative skill, at least insofar as they track more traditional measures such as essay grades.
For another, different teachers grade prompts similarly. Whereas I initially was the only instructor daft enough to sift through 100-page ChatGPT interaction logs (a process that now takes me approximately 15 minutes per student, which I deem to be acceptable), the efforts of additional graders resulted in sufficiently low between-grader variability. To this end, efforts were made to ensure consistency between graders by defining a prompt-grading rubric, explaining to students what kinds of things would be rewarded during the prompt-grading process, and by applying a standard prompt taxonomy to help graders classify and evaluate prompts according to their relevance and quality within the context of philosophical essay-writing. Thus, prompt-grading appears to be not only valid as a measure of philosophical knowledge and argumentative skill, it also appears to be reliable.
Given this preliminary evidence, I recommend taking seriously the idea that a student’s philosophical knowledge and argumentative skill can be assessed by evaluating the quality of the prompts they use to complete take-home philosophical writing tasks. This form of assessment can be implemented by requiring students to submit complete interaction logs that contain all prompts and AI-generated responses, and by asking teachers to evaluate these logs systematically, using a grading rubric and prompt taxonomy. A student’s interaction log can be graded, and that grade can be combined with a grade for the final essay or report to yield a combined measure of writing process and written product.
Of course, I do not believe that prompt-grading should be the only measure of student learning. The ability to contribute to in-class discussions, or to produce philosophical writing without external aid, continue to be effective measures of relevant knowledge and skill. Moreover, I also do not believe that prompt-grading is a perfect measure. In particular, it is not entirely immune to the worry that AI technology might allow students to pretend to know things they do not. Recent research on meta-prompting—training AI systems to generate prompts—suggests that enterprising students may use AI not only to generate philosophical writing, but to also generate the prompts with which to do so. That said, although I cannot predict the future, I am cautiously optimistic that even meta-prompting would benefit from the insertion of relevant knowledge and skill that could in principle be evaluated.
In conclusion, reports of death of the college essay appear greatly exaggerated. Certainly, ChatGPT poses a challenge. As is always the case, however, the challenge comes with an opportunity. In this case, the opportunity is to let students actively explore the limits of this powerful new tool, all the while teaching them to use it reflectively and responsibly.


This prioritizes one set of skills over another. As a student, I would be upset that I am not being taught the skillset that is lost. Undergraduate programs (and even some grad programs) fail to teach students how to write well and how to write good philosophy. And now we’re just abandoning teaching that skill to teach some other skill.
“As a student, I would be upset that I am not being taught the skillset that is lost.” I wholeheartedly agree. Tell your prof(s) that you want to actually work and be acknowledged for it: https://certifiedaifreeskillsandknowledge.org/
Note the the point of the whole “grading prompts” idea is precisely to value student’s work. The work being valued is effective, domain-relevant, and insightful prompting. Why should it matter whether a student expresses their knowledge and skills in a traditional essay, as opposed to an AI interaction log?
I understand the concern, and it is true that AI-assisted writing may promote different skills than traditional writing. However, to me it is not so obvious why the latter should be prioritized or valued over the former. We do not lament the decline in abacus-using skills, now that we can rely on digital calculators.
More generally, which knowledge and skills are valuable is context-dependent. The knowledge and skills that are needed to be productive, responsible citizens depends on the society we live in, and the technology that is available. The prevalence of AI technology not only requires teachers to adapt their assessment methods, but also to reconsider their learning objectives: Which knowledge and skills will outgoing students need to flourish in the society of the future?
But using AI to alleviate some of the heavy lifting in a philosophy essay is not a skill needed to be a productive and responsible citizen. Being able to think independently of such technology strikes me as much more valuable.
Every assignment prioritizes one set of skills over another.
I absolutely wouldn’t want this to be the only sort of assignment students have – for one thing, I suspect there are a lot of skills of evaluating writing that students have a hard time developing from merely this high level interaction with it, which they would develop more easily with more traditional writing assignments.
But a choice that no philosophy class should ever have any assignments of this sort is also a prioritization of one set of skills over another.
That’s the hypothesis, but what is the experiment to test it?
It seems something like an oral interview with the student is needed to verify their understanding. And understanding is the thing we’re worried that AI use will interfere with (the process), not so much writing quality with AI (the product).
Pre-ChatGPT, the product was a reasonable proxy for the process, but that’s no longer the case. A perfect essay can still be generated (incl. appropriate, informed prompts along the way) without actual understanding behind it. Sure, cheating has always existed, e.g., paper mills or hiring ghost writers, but AI makes is so much easier and at scale. It’s not just a difference in degree but also a difference in kind.
Carlos, was that kind of testing done? Interesting approach, nonetheless!
Why should an oral interview be the only (or even the best) way to verify a student’s understanding? The point of the above is to argue that another way of verifying understanding is to inspect and evaluate a student’s prompting behavior. The experiment that was conducted in support of this claim was to ask experienced graders to measure the degree of philosophical knowledge and argumentative skill that students deploy in their prompting, and to correlate that measure with an assessment of the same students’ essay-writing. As we report in the following preprint, the correlation is strong and positive: https://osf.io/preprints/psyarxiv/cne9j_v3 . In other words, prompt-grading is (at least) no worse a measure of philosophical learning than essay-grading.
I didn’t say it was the only way, only that there needs to be a reasonable way, and it wasn’t obvious what your experiment was to test your hypothesis.
Even if you rely on experienced graders to report back to you, you might still want a more reliable baseline. A direct assessment with the student seems much more reliable than another proxy measurement. You agree with that, right?
There could be legitimate reasons you haven’t done that yet, e.g., you will later, or that it’s too resource-intensive (though it seems you spent a lot of time to set up your project). But denying that it’s a good, if not best, way seems to be denying the obvious.
Now I’m curious what your IRB/ERB looks like, which I assume you needed and obtained from your university. Were there concerns or commentary about your methodology?
Prompt-grading as described here is direct assessment. We are directly observing student behavior (namely, their prompting behavior), and assessing it.
The correlation we found was between prompt-grades and traditional essay-grades, within-subject. You are right that an additional baseline (e.g. with oral interview grades) would at least be possible, and perhaps also beneficial, even if it is difficult to do for reasons of scalability. It is perhaps worth mentioning that (unlike much other research on AI in education) the data collected here is very “naturalistic”: these are real students doing real work receiving real grades. Given that our class sizes are large (ca. 150 or more), we typically rely on written work rather than oral work.
Yes, we received ERB approval. I don’t recall that any methodological concerns were expressed, but will check with my co-author and ask him to elaborate here if relevant.
I hope you (and others) don’t take these critiques personally, even if they can be pointed. Since it’s still so early with AI in education, anything still seems possible, and it can be worth experimenting with AI pedagogies.
But as I assume everyone knows, the skepticism surrounding AI in education is very strong. This means you have an uphill fight — the burden of proof is on you — and need to very careful with your methodologies and to not overstate findings or conclusions, if you want to guard against such criticisms.
If there can be a sound, effective AI pedagogy for philosophy classes (which doesn’t exist yet), I’m sure we all would love to know about it. Many of us would quickly use it so to not swim against and drown in the current of rampant AI cheating.
Even so, many others will continue to resist AI for other reasons, e.g., environmental cost, labor exploitation, etc.
Good luck.
Certainly not taking anything personally, and I do appreciate your reflective comments!
I agree that the burden of proof is on those of us who are deciding to use these tools. That’s exactly why we have collected, analyzed, and reported on the data in the preprint, even if it is not through peer-review yet. I am all for being cautious, but I am also in favor of trying things out and seeing what works.
This is the sort of thing that can only be genuinely believed by someone who thinks that good writing is merely ancillary to philosophical thought.
Also, that example prompt was deemed to be a sufficient demonstration of knowledge for a master’s-level course? Dear lord. That’s maybe a passable demonstration for a first-year critical thinking course. “Student understands that thesis requires justification”? “Student formulates a thesis in response to the essay question”? Yeah, at the graduate level, I should sure hope so. These are literally high-school-English-class essay-writing criteria. If this is what’s considered sufficient to be an MA student at Eindhoven, I fear for its future accreditation.
🙂 Thanks for your concern! That prompt was by no means sufficient. Rather, it was just one of a long list of prompts in this particular student’s interaction log. Typical interaction logs in this dataset were between 20-100 pages in length, and contained anywhere from 5 prompts to 50. The evidence for any particular student’s knowledge and skill is cumulative.
Might I submit that knowing that one’s thoughts should be structured and one’s claims supported is not good evidence of being able to structure one’s thoughts or support one’s claims?
Am I missing something, or are students being graded in large part on their ability to rephrase the rubric in the form of a question?
I am inclined to agree that knowing that (thoughts should be structured…) is different from knowing how, though this is of course a matter of philosophical controversy. That said, I think it is important to be careful about the relevant unit of analysis here: who or what is “one”, the bearer of the thoughts? If you take seriously the ideas of embodied/embedded cognition theorists, then I don’t see a problem of attributing the relevant “knowledge how” to the system comprised of the student and their AI-assisted device of choice. Similarly, “I” don’t know how to take square roots; but “I plus the calculator” sure does.
I don’t quite understand the point about rephrasing the rubric. The rubric says things like that students need to demonstrate domain knowledge and argumentative skill. The rubric does not specify what that knowledge and skill is; the course material does.
“Similarly, “I” don’t know how to take square roots; but “I plus the calculator” sure does”
Quoted without comment.
But isn’t this itself already a comment? I suppose you disagree with the idea? Great! Care to elaborate?
I don’t mean this snarkily; I genuinely don’t think it would be productive or collegial to do so.
One worry about this is students prompting AI to generate prompts. You say that “I am cautiously optimistic that even meta-prompting would benefit from the insertion of relevant knowledge and skill that could in principle be evaluated.” But if you aren’t evaluating these things, then you’re just letting students cheat on the assignment such that they learn nothing, right? So are you evaluating these things? If you start evaluating these things, do you worry about students prompting AI to prompt AI to generate prompts? If so, would you be optimistic that meta-meta-prompting would benefit from being evaluated? If so, would you start evaluating it? If you start evaluating it, do you worry about students prompting AI to prompt AI to prompt AI to generate prompts? Etc.
I suppose I am also interested in learning why you are optimistic that meta-prompting would benefit from the insertion of relevant knowledge and skill that could in principle be evaluated, and whether you are also similarly optimistic about whether meta-meta-prompting and n-meta-prompting for any given n would likewise benefit. If the process gives out at some point, what strategies will you employ? Is there any reason those strategies won’t work right now at the lower level?
(I suppose I am also interested in what you even mean when you say meta-prompting would benefit from blah blah blah. Aren’t we supposed to make sure students benefit? I don’t care if in principle I can make them better at meta-prompting unless making them better at meta-prompting makes them better at something I care about for its own sake. Indeed I have the same worry about prompting, which others have expressed already.)
The standard response to these legitimate concerns has been to say that they are a sign of paranoiac distrust of students or obsession over inevitable cheating loopholes. Sadly, what you are suggesting just happens to be true. Many students will simply get A.I. to work on itself. I bet many of them would also not see it as cheating. The ones who are already relying on A.I. to help them with their homework need to learn the absolute basics, not only about philosophical writing, but about the integrity of good scholarship. I don’t see how that is achieved by letting them play with these machines when the most reliable method is right in front of us and has been since the invention of paper and pencils.
Certainly, meta-prompting is a serious concern here. I have little first-hand experience with it (so much to do, so little time), so the assumption that it too would require relevant knowledge and skill is little more than a hunch at this point.
That said, yes if students go down this path I would expect them to submit their prompts, and yes I would evaluate them. That’s the deal: I let you use whatever you want, as long as you are transparent about it and allow me to (try and) assess your philosophical knowledge and argumentative skill. I don’t care at which level of meta-ness that knowledge and skill is expressed.
You say “if students go down this path I would expect them to submit their prompts.” What if they don’t behave according to your expectations? What if they just lie to you and tell you that they came up with this stuff themselves? Do you have any mechanisms for preventing this from occurring? If not, do you worry that some subset of your students (potentially a large subset) are not learning anything?
Right. To my knowledge there is no reliable way to expose this kind of cheating, much like there is no reliable way of exposing the more familiar kind of cheating of asking an AI system to write the essay you were asked to write unassisted (AI detectors don’t work). I am aware of this issue. I do think that there is more deceptive energy involved in pretending you wrote the prompts to generate the essay, than in simply pretending you wrote the essay in the first place, so am unsure why anyone would take that route. But of course, if someone really wants to cheat, they will find a way to do so.
The most I can do is to disincentivize. Rather than impose prohibition, my way of doing so is to be at once maximally permissive (you can use any AI you like) as long as you are also maximally transparent (you should tell me what you are doing).
JFC. Never a clearer instance of lipstick on a pig.
Haters gonna hate, I suppose.
What reason do you have to think that they’re not feeding the scaffolding to AI, too? The thing is: prompt engineering isn’t hard, and knowing that essays ought to be structured isn’t the same thing as knowing how to structure them.
Maybe it doesn’t matter, but I don’t see that they’re learning what they used to learn.
Note that it was explicitly stated that I do not assess prompt engineering. I assess philosophical knowledge and argumentative skill, both of which is transmitted in the relevant course. Also as stated in the above: it is easy to underestimate how difficult good prompting actually is. Generating a high-quality end product (be it an essay, a computer program, or an artwork) requires the infusion of domain-relevant knowledge and skill.
RE: knowing that vs. knowing how: See above response to @Kelly Weirich.
RE: learning what they used to learn: See above response to @Student.
But many students don’t aim for a high-quality end product. They’ll be content with a B- that requires almost no effort.
This is true. Perhaps then our overall standards for the final product should go up, assuming AI is used. So, an “almost no effort”-level essay that can be produced using minimal prompting with no relevant knowledge- and skill-infusion should not receive a B-, but rather a failing or almost-failing grade. As the assisting technology gets better, so should our expectations on the final product.
I see that you’re saying it, but I think it’s likely that they’re also using AI for creating or refining the prompts. And if you’re providing rubrics, you’re doing a lot of the legwork for them.
I think it’s great to experiment. But I’m mindful that one of the problems with the rise of “vibe coding” is that it’s much harder to teach the basics to first and second year students, which means that when they try to vibe code in upper division classes, they don’t have the understanding to do it well. The CS department at my institution dealt with this by moving all 1000-level courses to the proctored lab.
I agree that from a student’s perspective, knowing the rubric is a big help. It is perhaps worth mentioning that in the Dutch higher education system, rubrics are mandatory to promote grading transparency and consistency. Of course, rubrics can also promote a kind of “box-checking” behavior, but that is true whether or not students use AI assistance of any kind. For my part, I try to make sure that the rubric states clearly what is expected, without revealing too much.
The problem you raise with respect to vibe coding generalizes, and is well-known: if many of the basic tasks and entry-level jobs are off-loaded to the machines, who will eventually do the more advanced tasks and supervisory jobs? We certainly need students to learn the foundations, and if they cannot learn them by interacting with AI, I am happy to promote other kind of learning. Indeed, in my own courses I use many different methods, including hand-written exams and oral presentations. But I also use prompt grading. There is no reason to think that this has to be a one-size-fits-all solution.
My biggest worry about the claims here is the reliance on correlation between prompt grades and more traditional evaluation of writing. I would not be surprised if people who are already good at traditional writing tasks do better at having good prompting interactions than other people. The question I’m most interested in with this kind of assignment is whether there is any evidence that students who initially do badly at this get better at it, and are then able to transfer that skill to other tasks.
Great question! Thus far we have not measured the extent to which prompts can promote learning (much in the way unassisted writing can promote learning), but only whether they can be viewed as evidence for learning. The correlation suggests that it can. Whether or not this is surprising, the practical value here is the finding that there exists a reliable signal that can be measured.
Perhaps going a bit in the direction of the question that most interests you, in the accompanying preprint at https://osf.io/preprints/psyarxiv/cne9j_v3 we do distinguish between different types of student prompters: students who use certain types of prompts (or, prompting strategies) rather than others get better essay grades. In particular, students who use the AI as a partner in developing and articulating ideas do better than students who use it only to do internet research and language improvement. But note that here too we still only have evidence for correlation, not causation.
Along the lines of my recent guest post here, I would encourage you to explore assigning theory-driven prompting strategies, which you can do even without customization (though that is better). It looks like you evaluate prompts mostly for structure, not for content. A nice complement would be to have students supply the GPT with lexicons (terms and definitions) and grammars (relationships between terms) so they are give the GPT a theory to work with. For example, if they are writing on Ross’s “the right and the good” they would need to provide the GPT with definitions of fidelity, reparation, gratitude, beneficence, and non-maleficence (the right), virtue, knowledge, justice, and pleasure (the good), provide the grammar for how the right and the good connect to one another. Then you can evaluate the substantive content of the instructions (are they correct, complete, suited for the purpose), how well the structural parts of the prompts make use of the theory, and how well the ultimate output reflects a thoughtful, faithful, and comprehensive use of the theory part of the prompt.
Yes I saw your post a few days ago. Very nice! When we grade prompts, we do look at both structure (in the sense of argumentative structure imposed by the students) and content (in the sense of theoretical knowledge introduced by them). The latter typically occurs when students introduce definitions or conceptual frameworks from class, or when they advance their own claims/ideas which they then either wish to justify or elaborate with the help of the machine. So I think what we are doing is very much in line with what you describe, except that we have not looked systematically at students’ use of customizable GPTs. We have observed students doing this occasionally, in which case we requested insight into the GPT to see how it was set up, so that we can apply the grading rubric to it.
The author considers meta-prompting as a form of cheating but is “cautiously optimistic” that students who use it would still be developing the relevant skills. What they overlook is that traditional cheating would be even more tempting under this assessment. Traditional cheating is getting your friend, or acquaintance, who is much better at the subject than you to write the essay for you (or paying an essay writing mill to write it). Because writing a unique essay of decent quality is a significant investment of time, even for a competent practitioner, good-old-fashioned cheating was never widespread. However, writing a high-quality essay prompt is NOT a significant investment of time. It could presumably be done in a couple of minutes by a skilled practitioner. So, students who want an A+ on an essay prompting assessment, yet lack the domain specific knowledge to write a good prompt, can simply solicit a prompt from an acquaintance who is great at writing prompts for this subject. And, where such acquaintances are absent, we can expect “prompt-writing mills” to emerge, selling A-level prompts for a few dollars.
The truth of the matter is that these ways of cheating are just too tempting for students when the main reason for going to university is credentialling. Yet developments in AI are putting pressure on the credentialling business model that has dominated higher education for decades. The solution is that universities need to shift their business model towards offering more intensive, high-quality learning experiences and adopting assessment methods that make cheating impractical. This is where oral exams come in. As I argue here, they are very effective at disincentivizing AI misuse, plagiarism, and outsourcing, while still letting students experiment with AI tool under evaluative feedback (in the form of an assessment of how well they have taken intellectual responsibility for their work).
That’s an interesting reflection, thank you. Certainly, domain-experts will find it easier to engineer prompts to write a paper, than to actually write that a paper. Insofar as this makes the job of ghostwriting/paper milling much easier, I agree that the proliferation of AI could have the effect of making these services cheaper and thus more widespread.
I don’t agree, however, that the use of prompt-grading would worsen the matter. For the ghostwriter/paper-mill, there is no significant difference in either cost or labor to sell the prompts as opposed to selling the paper that was generated using those prompts. To the contrary, one could also argue that the ghostwriter would even charge more, since the cheating student would be purchasing two things (prompts + paper) rather than just one. So it could also be that grading prompts in addition to grading papers (indeed, this is what we did here, prompt-grades are only one part of the total assignment grade) would be more more “ghostwriter-proof” than the traditional method.
Of course, this says nothing about problems associated with “credentialling” or the unique advantages of oral exams. I like your thoughts on those.
Hi Matthew, I would generally agree with your solution, provided an instructor has the time to do all this.
But a comment on this bit:
If I may speak for those who prohibit AI:
It’s not that we want to prohibit it, but it’s that we see it as the lesser of evils. Yes, AI can be used responsibly, e.g., for deeper thinking, but will it? Seems reasonable to think that allowing for any AI use, no matter how tightly prescribed, can cause confusion or enough ambiguity or opportunity that it encourages irresponsible or proscribed use.
So, a prohibition makes it crystal-clear that any use is cheating, and the student would need to deliberately choose this path, which is a form of deterrence.
Also, even if your pedagogy is sound (as it seems on a quick read), you noted the high costs to the instructor to do all this. As you say, it seems to be an “unrealistic burden.” I would think that a viable solution must be a sustainable one.
But I don’t know many instructors who would have the extra time to do all this. Even if they have an extra 20 hours to spare per term (assuming 80 students per term, on average), it’s still much, much easier to just impose a ban and require students to do all of the work themselves, which is time-tested pedagogy.
If I were someone teaching in-person to only 30-40 students, perhaps I would try this, assuming I’m able to put aside other concerns about AI use, e.g., environmental, labor exploitation, etc. Anyway, it seems worth experimenting with different pedagogies to account for AI. We need all the help we can get!
Unfortunately, some of us have been devoid of the choice. Where I teach, blanket prohibition itself is now prohibited, so we’re forced to force the ambiguity on students.
Thanks Patrick, I have learnt a lot from reading your work on this topic.
Yes, I understand that, but prohibiting it risks making us and our courses irrelevant, so I think there must be a better way. My argument is: (1) Being able to take full intellectual responsibility for your work is the test of using it responsibly, (2) the oral exam gives students direct feedback of how well they have taken intellectual responsibility for their essay and also, via the grades at stake, motivates them doing so, (3) if undergraduates get the opportunity to take many courses with the essay + viva combination they will get the opportunity to refine their ability to responsibly use AI over time (e.g., if they overuse AI on one essay and get a poor mark for their viva, they can use the feedback of where their understanding of their own work failed to adjust how they use AI on the next essay they write).
Regarding the time commitment, for instructors without institutional support, one thing they can try is readjusting their syllabus to make room for the viva. For example, you might reduce the content in your course and leave the last two weeks of class time free for conducting vivas. Or, if you used to assign two essays, you might replace that with one essay plus viva (presumably conducting a viva and marking a second essay take about the same time). Obviously these options involve compromises that might be less than ideal, but in difficult circumstances we must be willing to make such compromises.
Having said this, as I argue in the piece, I think that ultimately universities need to respond to AI by investing in more intensive forms of instruction. Their standard business model is now under threat and people are increasing asking: If AI can complete the classwork set for students, what value do universities provide? The answer is to invest in more intensive, high-quality forms of education.
In an age of perpetual budget crises, I understand that many universities are not in a position to do this. But this just shows that they (or their government funders) have dug them into a hole with an unsustainable business model. For those universities that have the means, making these changes is essential. At my institution, all classes are capped at 48, which makes a viva doable but still demanding in a fully subscribed class. However, our leadership recognizes that more intensive education is a sensible way to respond to AI and are planning to introduce a new Freshman seminar capped at 20 that provides this. And we are a public university—although admittedly a well-funded one as the Singapore government has prioritized investing in higher education.
“How effective is prompt-grading as an assessment method? As we report in a current preprint, the initial results are promising. For one, students who demonstrate philosophical knowledge and argumentative skill in their prompts produce better philosophical writing (i.e., they receive a better essay grade) than their peers who […] only or primarily employ domain-general prompt-engineering principles, and do better than students who who merely use AI to improve linguistic “surface features” such as spelling, grammar, word choice, and style. Thus, prompt grades appear to be sufficiently valid as a measure of philosophical knowledge and argumentative skill, at least insofar as they track more traditional measures such as essay grades.”
Yes, but are they better or worse than students who do not use AI at all during their academic career? If they are worse, I am afraid we are doing them a disservice by allowing them to use something that will make their thinking and writing worse.
This is a good question, and perhaps the most important one of all. On a narrow interpretation, within the context of the particular courses and data we have analyzed, we found that there was no significant difference in overall essay-writing performance between AI-users and non-AI-users. Indeed, the average essay grade for both groups was surprisingly close (less than 2 percentage points). Of course, the essay grade does not account for other factors that might be relevant in your conception of “better” or “worse”, such as time taken to complete the task, or ability to transfer the learnings to other domains.
From a broader point of view, however, I don’t think we can answer your question definitely at all. It all depends on what you mean by “better” or “worse”. If the standard you are applying is the knowledge and skills that were needed to thrive in a pre-AI society, then it may well be that students who use AI too much now will be “worse” off in the future. But if the standard is instead the knowledge and skills that students will need in post-university AI-assisted careers, then I would certainly argue that learning to use AI effectively and reflectively in university is “better”. So, what this all comes down to is the question about learning objectives: what do we want students to learn? See also the previous reply to @Student.
The point is that we are talking about why studying philosophy, not about acquiring specific or technical skills. *If* there is something to be learnt through a philosophy education, this is certainly not how to quickly write a good-enough paper. Rather, I suggest that what we want students to learn is the ability (and patience) to sit comfortably with difficult ideas and try to find a solution which is not going to be easy nor ready-made. The use of AI, with its quick outputs, seems to me to be extremely counter-productive in this sense.
I think you might be missing an important point here, which is that the aim of all this has nothing to do with “quickly writ[ing] a good enough paper”. If that were the aim, I would just be grading generated essays and call it a day. Instead, the aim continues to be teaching students to think, reflect, and apply philosophical knowledge that they have acquire through reading, thinking, discussing, and listening. The point is just that I think (and the data seems to show) that we can determine whether this aim has been satisfied by inspecting and evaluating students’ prompts. Student prompting is just another behavior (like traditional essay-writing) that we can observe in order to assess their learning.
Thanks for your answer.
Just to be sure I am understanding correctly:
—You appear to disagree with my idea that the frustration of not finding a solution is an important part of philosophy and that we need to teach students to endure it and even welcome it
—I am not sure I understand how you can avoid an infinite regress (e.g., someone prompting AI to generate a prompts that shows that they understood what you want them to show, after having fed them your paper)
I don’t disagree with the “frustration” point. But it is a mistake to think that using AI to generate good essays is trivial. As stated in the piece, I think it is common (in philosophy particularly!) to underestimate how much labor and expertise actually goes into generating something good. Just ask AI artists or expert software developers, who know that crafting good prompts takes considerable expertise and labor. The idea here is to try to measure this expertise and labor as and when the student expresses it. Grading essays in a more traditional way tries to do exactly the same thing!
And also, remember that I am not saying anything how or whether students can learn by using AI–that is another conversation altogether. I am only talking about how or whether teachers can assess by looking at the way in which students use AI. And the data says, tentatively, “yes”.
Your final question about “infinite regress” seems to refer to the worry about meta prompting. I do acknowledge that this is a worry, and have stated both that I don’t have an easy fix for this but that I do have a hunch that this too is likely to require more expertise and labor (on the student’s part) than we might initially assume.
Carlos, you said:
Not all AI art, right? Some art is very simple, while others have all kinds of details and involve sophisticated techniques.
Same with philosophy: some writing can be very technical and complex, while others are one-sentence aphorisms that stay with you for the rest of your life.
Without nailing down a definition of art or philosophy, I think the most we can say right now is that AI-generated works can resemble art and philosophy. But whether it is art or philosophy is still an open question.
And judging whether something has aesthetic value (to you) seems much, much easier than assessing whether some writing has philosophical value. That may make a difference in our attitudes on AI art vs. AI writing.
This is an interesting thought, thanks Patrick. I am personally unsure about the appearance/reality distinction.
We seem to agree that AI-generated philosophy, like AI-generated art, can resemble the human thing very closely–the appearance is similar. We could even speculate that the similarity runs deeper, to similarity in all relevant intrinsic properties. So then what you seem to be proposing is that reality–whether something really “is” art or philosophy–depends on its relational properties, such as the social context in which it is created, how much effort is expended, or whether or not the author is a person. That may well matter.
I am unsure, however, whether this all clearly speaks against counting either AI-art as art or AI-philosophy as philosophy. Again, for one, I do think more human effort and ingenuity is required than most commentators appreciate. In that sense, using AI to do anything is not all that different from using any other technology. So in the cartoon you included, I don’t see a substantial difference between “pressing enter” after writing a prompt and e.g. running a printing press after you have set all the pieces and patterns to be printed. I would be happy to see the results of either process being called “art”. Perhaps philosophy is similar.
But I suppose the interesting test case would be one (mostly counterfactual) in which the AI does the art or the philosophy completely autonomously. Here I think we really do arrive at the question of what “art” and “philosophy” actually mean, since it would involve answering the question of whether relational properties matter for the identity of either, and if so, which ones. I cannot answer that question, of course.
Again, thanks for the thought-provoking reply!
But why do any of this? You can also assess learning via blue book examinations and participation in class discussions. Why put all this effort into finding a way to save essay writing if, at the end of the day, they’re not actually writing the essays? Teach your material, and then test them on it. That’s how you know how well they learned it.
Agreed. First, as I suggest in the piece, there is no need to think of this as the only viable kind of assessment. Blue book exams–great! Class discussion–also great! My point here is to say that prompt-grading is just one additional assessment method that you might want to consider in addition to the others, since it appears to be both valid and reliable.
But, let’s dig deeper: every assessment method has its advantages and disadvantages. Blue book exams are obviously good at testing the “biological” student but have limited real-world applicability. Very few people in non-classroom settings still think, write, learn, and argue with pen and paper. They do so in front of a computer, sometimes with the help of the internet, and sometimes with the help of AI. If this is the future of creative work, why measure their ability to do it with antiquated technology? Why not measure it (and, I repeat, it seems to be possible!) using the new technology directly?
And of course, the students are “actually writing the essays”. Is Refik Anadol not “actually doing the art”? Using an AI system to create something good, and original, and convincing requires more than just very simple prompting. Students still need to work hard and use their ingenuity to get the machine to do what they want it to do, and prompt-grading measures this labor and ingenuity.
“Teach your material, and then test them on it”–that’s exactly what I am doing, just that the test is through the medium of prompting rather than the alternative media of handwriting or class discussion. All of these media afford assessment and have their own advantages and disadvantages. Everyone can just pick their favorite.
As I posted in another thread, even in a coding company with wide incentives, most people choose to never use AI. I suspect that in philosophy, the vast majority of faculty will also choose to avoid AI. I think your finding is very interesting, but I guess people will mostly ignore it because of their choice to avoid AI in research and in classroom.
Yes this is interesting. I suppose what matters more than whether faculty choose to use AI is whether the students do.
FWIW, I have been looking at student prompts for approximately three years (we started doing this just before the initial ChatGPT release). At first, in the context of a small, optional assignment, all students were allowed to use generative AI in whatever way they wanted, and although we looked at their prompts we did not grade them. Here, AI-use rates were very high, over 70%. When in later years we announced that we would start grading prompts, AI-use rates dropped down to approximately 15%–my guess is that students were cautious both with respect to the quality of AI outputs and with respect to the novel assessment method. At present, we see AI use rates of approximately 30-35% (even with the announcement that prompts would be graded).
So, overall, my guess is that, whatever faculty do, (a) students would use AI highly if there is no immediate consequence of doing so (such as being graded on the prompts), and (b) students are using AI at increasing rates (even when we grade prompts) as they become more familiar with the technology.
Of course, the “real numbers” might be higher, since there might always be cases of AI-use that we do not detect.
Yes, what you observed aligns with the report shared by my school. What I don’t know (or don’t remember) is the comparison on students’ AI use and its effect between a course with AI guidance like yours and a similar course that is AI free and uses only blue-book and oral methods
“Blue book exams are obviously good at testing the “biological” student but have limited real-world applicability. Very few people in non-classroom settings still think, write, learn, and argue with pen and paper. They do so in front of a computer, sometimes with the help of the internet, and sometimes with the help of AI.”
This idea of “real-world applicability” makes little sense. In the “real world” it is rarely the case that anyone needs to construct a complex piece of philosophy at all. The point of teaching philosophy is not to prepare people for the real world, but to enrich them and enable them to think, read, and write critically. Learning how to do it with pen and paper isn’t outdated, because making them use pen and paper isn’t for the sake of making them better pen and paper users in their future life. It’s intended to make sure they produce nothing but what they themselves have learned. Doing so, you’ll be a heck of a lot better in any other setting involving a computer.
“If this is the future of creative work, why measure their ability to do it with antiquated technology? Why not measure it (and, I repeat, it seems to be possible!) using the new technology directly?”
If “the future of creative work” refers to academia, then I would say: we don’t have to let it be the future of academic work. We actually control that, because we teach it. If you refer to creative work in industries outside of philosophy, then see above.
For sure, the fewest individuals will need to “construct a complex piece of philosophy” in their post-university lives. Also, certainly “the point of teaching philosophy is…to enrich them and enable them t think, read, and write critically”. I agree whole-heartedly. But I don’t see why what I am doing is not a way of measure exactly this, just that unlike others, I look for evidence of thinking, reading, and writing ability in the prompts, rather than some other medium. In what you have written thus far, I don’t see any direct engagement with this idea at all. Don’t you think it is possible to measure this capacity in the prompts? Why not? How is our experiment faulty and our data misleading? Please do explain.
In the old days, there were websites where students could pay someone to write their essays for them. This would usually result in papers that were quite good relative to most undergraduate work, and the academic dishonesty behind them was sometimes hard to detect. On many occasions, you as the instructor could prove manifestly that the work wasn’t theirs, for example by tracking the student down for a face-to-face conversation where they turned out to be wholly at a loss to explain anything about the essay. However, in other cases, even when the deception was obvious from your point of view, there was no concrete and reliable evidence beyond your own judgment that you could pass on to an office of student conduct to justify a punishment. This was especially so if the student had devoted even a modicum of time and energy to hiding the most obvious signs that the work was not theirs.
Imagine that these websites had stuck around, but changed. Instead of charging money for paper-writing, they would now offer their services for free. Instead of taking significant time to complete these essays, they would now return full assignments within minutes at the most. Accordingly, imagine that these websites quickly became common knowledge among students, and also became extremely popular, with perhaps a majority of students coming to regularly use them to do much of their work.
Soon, instructors would have to deal with a situation where perhaps most of the assignments submitted to them were completed at least in large part, and often entirely, by people other than their students. Although the instructors, not being stupid, could tell that this trend was occurring, given the massive change in the work they were reading, they would not be able to penalize most of the dishonesty, given the near-absence of reliable, objective detection in all but the most egregious cases. They would thus have almost no other way to keep their students from taking credit for work they didn’t do, except to make instantly-ignored pleas for students to stop, or to resort to such drastic measures as ceasing to assign take-home work, which would make it much more difficult at best to teach students paper-writing skills, and which would create the risk that students would simply avoid their courses for others that would be more lax.
I would hope that, in such a situation, the entirety of the academic world would have immediately recognized this for a danger of the utmost severity to all of education. Students cannot learn to think, read, or write by getting others to think, read, and write for them, any more than someone can learn to throw a baseball by using a pitching machine. Allowing students to take credit for work that isn’t theirs would thus completely defeat the purpose of virtually any course, and would be deeply unfair to the few students who have enough integrity to follow the rules, and do the work themselves. This would threaten to turn all of education into an absolute sham.
I would hope that, in such a situation, institutions would immediately have done all they could to address this problem. They would have made strict rules with firm penalties against using paper-writing services in any way, and repeated these rules with emphasis to students. They would have offered instructors extensive training on how to detect the use of these services, and guidance on course policies that could effectively prevent the practice. They would have made it so that all classes, including online courses, would have mandatory in-person testing accounting for a significant portion of the final grade, so as to prevent instructors who use such measures from running the risk that their classes won’t fill. All of this would no doubt still have failed to contain the problem, but it would at very least be doing all that could be done about it.
In reality, after the birth of ChatGPT and related services, we’ve ended up with a problem that comes with the exact same effects and dangers. There are now computer programs which are in many ways functionally identical to paper-writing services, insofar as they will produce A-quality essays for students, but which also require no money, no understanding of the material or scholarly abilities, and almost no time; and there are no objective and reliable means for detecting or punishing the use of these programs in most cases. Education is rapidly turning into a sham as a result, as students come away with better and better grades while learning less and less, and instructors are forced to spend more and more of their time commenting on work their students didn’t do.
However, in reality, the problem has been met with the opposite response. Institutions are actually falling over themselves to support and encourage the use of AI, making it even more accessible to students, and often trying to get instructors to use it as well. The instructors who take a strong stand against this are given little to no help in dealing with the problem. And some instructors have even gone so far as to encourage their students to make use of programs that are functionally identical to paper-writing services, and teach students how to do so. Certain instructors have even joined the bandwagon themselves, letting computer programs compose their lectures, design their assignments, and even write feedback for students, all while happily collecting their salaries for this work they didn’t do. (Some of these instructors might even indignantly stammer out excuses for this behavior in reply to this very comment.) A few of these instructors, in an act of superlative wishful thinking, have convinced themselves that they are still teaching their students in all this, insisting that students will still have to practice some skill or other in order to get the computer program to do their work. They are, and will likely remain, in deep denial about the fact that the students can simply get the program to do that part for them as well.
We’re all cheerily sailing along towards a future where college becomes a bizarre ritual in which students send in AI-generated work, and instructors return AI-generated feedback, in an endless loop of vacuous unreality. Before too long, either students or institutions will realize it’s not worth paying instructors for this, and higher education will vanish, but not before they start selling keyboards for students whose only buttons are CTRL, C, and V.
Am I exaggerating in all this? Less than I’d like to be.
Thank you for putting this so well.
Consistent with what you say, I would suggest that the continued efforts on behalf of philosophers to reconcile our teaching/grading with AI may, ironically, have the effect of killing academic philosophy even if higher education as a whole doesn’t vanish, and even if there weren’t instructors letting machines do their jobs for them. Those who manage the institutions will lose the little tolerance they have left for the humanities departments which exist just to give grades to essays about which it is impossible to tell if any conscious being authored it. Even if the admins are reluctant, the students will vote with their feet.
Honestly, I never know what to do about these kinds of doom-and-gloom predictions. Perhaps I am too much of an optimist, or perhaps I am just more willing to adapt and less tied to “the way things have always been”.
Interestingly, I have personally never felt that there was a better time to be an academic philosopher. Never has there been more need and desire, within academia, industry, and society generally, to understand what exactly is happening in the context of these technological innovations. Many philosophers are addressing this need by teaching students to be responsible and reflective users and developers, but also by teaming up with developers and regulators to shape the technological and societal environment of the future.
Certainly, not all philosophy need to concern itself with AI directly, but of course even that kind of philosophy does not spin frictionlessly in a void: it always occurs (is practiced and taught) in a societal and technological context. At present, that context is colored by the developments in AI, just like in the past it might have been colored by the invention of the printing press, the questioning of religious traditions, or the transformation of political and economic institutions. Those philosophers who engage these issues actively–and there are many ways to practice this kind of engagement–are likely to stay relevant, and to position themselves to effect positive change in society. Those who ignore these issues, simply wishing them to go away, will have as much of a hard time in the marketplace of ideas as they will in the marketplace of economics.
I have not tied my concerns about AI to the idea of keeping to “the way things have always been”, as though that is an independently desirable outcome. That is an unfair characterization. Things were better in academic philosophy before the AI explosion *for a reason*, and that reason is what has driven most of the objections raised against your approach.
Suggesting that the development of AI technology is on par with any other past historical context in which philosophy has been done also strikes me as clearly false. The printing press didn’t quickly and relatively painlessly spit out completed assignments for students, did it? Nor did political or economic transformations. That seems like a relevant difference.
You yourself have admitted that the possibility of students prompting AI to prompt itself is a problem. But that’s not just a problem; it’s an enormous problem. With the precarious situation many philosophy departments are already in, don’t you think university admins will be less than pleased to find out that we have no way of really proving our worth anymore? That we are paid to let our students do what they want with these tools, even if that means we can’t know exactly what it is they’re doing?
We have a way around this, which is to show that we are teaching our students how to think in an AI-less context. I suggest we take that option seriously, given that we don’t yet seem to have any good way to hold students accountable for their “proper” use of AI.
I am sympathetic to the concern that this round of technological innovation seems different in many ways, because it specifically targets cognitive labor. But I also don’t think it is completely unprecedented. Here is perhaps a more compelling comparison: the invention of writing. Recall that philosophy (at least in the Western tradition) was originally an oral discipline. With the invention of writing, it was possible for the first time to capture thoughts in a stable way over time, thereby making them objects of reflection in their own right (isn’t that what philosophy is, to a large extent?). Certainly this was a revolution, and it fundamentally changed not only what we as a species could do, but also how we think and reason. Presumably, philosophy as it was practiced and taught before the invention of writing differed profoundly from the way it was practiced and taught afterwards. Was everyone happy about this transformation, at the time? Not sure, but I can certainly imagine some of the more conservative visitors to the Agora complaining about the lackluster oratory skills of “writing-based philosophers”. Despite their complaints, however, writing has become absolutely central to the discipline of philosophy, and I don’t think that anyone is complaining in retrospect.
Why not think that AI could have a similar transformative effect, and that this effect could be just as positive as the effects of writing? Why is your prior so high on the effect being negative, not only for philosophy but on society more broadly? What’s your reasoning? What’s your evidence? It seems a bit like highly speculative futurist dystopia to me, whereas here I am trying to give you actual data to suggest that we can in fact observe patterns in human behavior that closely resemble the kinds of thing that we want to see, namely critical thinking, applying knowledge, structured argumentation, and so on.
As for metaprompting–yes, I know that this is an issue, and I have stated my hunch about here before. But I don’t have any data to confirm it, yet.
As for the precarious situation of philosophy departments: yes, departments that don’t adapt may be in trouble. That said, I also don’t know of any departments that forbid reading and writing and instead focus on oratory.
This is a helpful analogy in many ways. But your discussion failed to acknowledge that human assistance for essay writing comes in degrees. Having an essay mill produce the entire thing from scratch is just one extreme. Other forms of human assistance include writing the essay yourself and then having professional editors work on the document, proposing hundreds of improvements that you can either accept or reject. Or, chatting about the topic with a friend who is a subject area expert, having them help you work out what you want to say, and then lead you to the best arguments you might use to defend your position, after which, you write the essay yourself. Or, imagine some combination of these things: You have the conversation with the subject area expert that clarifies your thesis and helps you find the best arguments, your admin assistant takes notes during this conversation and writes a detailed essay plan based on what you indicated you wanted to do, you review and revise this detailed essay plan and then pass it on to professional editors who transform your bullet points into well-crafted paragraphs, which you then review and revise until you have a final essay that you are happy with.
A difficult question here is where we draw the line for acceptable levels of human assistance. Obviously, wherever we draw the line, it must count using an essay mill as unacceptable. However, it would also obviously be wrong to hold that ANY form of human assistance is unacceptable. Asking my friend if she thinks this line of argument I am pursuing makes sense, or getting my brother to proofread my final draft are things that everyone should recognize as acceptable. It is the cases in between (getting substantial help from professional editors or subject level experts, etc.) where things become more contentious. Luckily for those of us who taught pre-AI, we were largely able to get by without resolving these contentious cases because these uses were rare (mainly because of cost barriers). However, now that AI assistance is readily available to every student, and can be used in a range of ways, analogous with different level of human assistance, we are forced to resolve the contentious cases.
Given this, I want to make two points. First, banning any degree of human assistance (including discussing your argument with a friend or getting someone to proofread your final draft) would be a very silly and unjustifiable response to the problem of essay mills. Likewise, banning any use of AI in any part of the essay writing process is a very poor response to the problem of students submitting essays entirely written by AI.
Second, a very attractive way to resolve the contentious cases is to appeal to the notion of intellectual responsibility. Getting assistance from another human, or from AI is ok so long as it does not stop you taking full intellectual responsibility for the essay you produce. As I explain here, taking full intellectual responsibility means that when you submit your essay, you are able to explain and defend every significant choice you made—the thesis you advanced, the evidence you selected, the counterarguments you considered, the conclusions you drew. If you can’t explain why you structured your argument the way you did, or why you chose one interpretation over another, then you haven’t taken intellectual responsibility for your work. Given this, I think the lesson to draw from your analogy is that testing how well our students are able to take intellectual responsibility for the work they produce should be a central part of higher education.
What you say at the end of your reply seems to me to be perfectly correct, and I would consider my prayers answered if institutions adopted an approach to countering AI-powered academic dishonesty that revolved around ensuring that students can take intellectual responsibility for their work in this way, such as by requiring them to account for what they’ve written in face-to-face or at least live conversations in all their courses. I have long wished I could set up a system where I have conversations like this with all my students in my own classes, and have tried to think of ways to make it work, but I still haven’t come up with one. I have between one and two hundred students per semester and no teaching assistants, which means that I would have to schedule all or most of these meetings outside of class; I could not handle the demands this would place on my time, and I question whether it would be fair for me to ask this of students.
Without such a system in place, I don’t believe there is a practicable way to allow some use of AI without forbidding all use of AI as a matter of course policy. Detecting whether students have used AI in writing an assignment in any way whatsoever is already next to impossible in cases where they make any effort to disguise this. Ascertaining whether students have used AI in ways that are consistent with intellectual responsibility, in a way that would allow me to differentiate this from cases that are intellectually irresponsible, seems to me to be entirely infeasible under circumstances like mine. Again, the only real way to do this is through live conversation, but since probably most students would make some use of AI under such a policy, and since I can’t figure out a practicable way to have live conversations with any significant number of them, I’m left at an impasse.
I would agree with your point that this is analogous to human assistance, if not for the fact that AI will do something which pretty much no human will ever do, namely write work that is virtually guaranteed to be A-quality for free in mere moments and on demand. For this reason, I see no need to stop people from, say, bouncing their ideas off their friends, because the odds are vanishingly small that their friends will respond by giving them a polished complete outline of their paper, or a set of well-composed topic sentences, or a full paragraph, which they can reasonably assume they can then profitably incorporate into their own paper without making anything more than superficial revisions. By contrast, a student would have to make quite active efforts to discuss their paper with ChatGPT without getting any of this sort of material in response.
In short, students can very rarely make use of the input they get from other human beings without having to do some work to understand this input in their own terms, relate it to the subject of the assignment, figure out how to work it into their paper, confirm that this input isn’t based on mistakes, and so on, all of which will usually require them to have a decent grasp of the material, and practice important writing and critical thinking skills. This is why I don’t think there’s any analogous need to police contributions to paper-writing from other humans in most cases, even though it could lead to the same outcomes in exceptional instances.
Hi Abraham. You say:
Curiously, this is exactly what we argue to be true for AI. Prompting is not trivial. Generating good essays is not a one-click thing. It requires exactly this kind of back and forth. And this is what we measure…
Prompting is in fact trivial when you get ChatGPT to come up with prompts for you, which you have conceded you make no effort to stop your students from doing.
Where does your confidence with respect to the triviality of meta-prompting come from? Are you an expert meta-prompter? One of the recurring tropes of science fiction is that of the infallible and omnipotent machine. But we know that this is not an accurate picture, at least not with the kind of data-driven AI that predominates today. So why think–or rather, insist without argument–that AI-driven prompting would be so easy and effective, trivially and effortlessly giving students what they want?
Yes, I acknowledge the limits of my approach, at least in its current design. I have stated my hunches in this respect, and declared them to be such. In contrast, what we get here is a level of confidence in one’s own assumptions that we otherwise only know from machines.
Where does my confidence come from? See for yourself.
I went to chatgpt.com, and entered the query, “Come up with a prompt that I can give to an AI to write a good paper for a class on the philosophy of artificial intelligence.” The AI proceeded to do precisely what I asked, just as readily as it would have written an ordinary philosophy paper for me if I had asked for that instead.
Just for kicks, I then input the query: “Write a sequence of other prompts similar to the first one, and make these prompts build on each other in such a way as to make it look as though I’m revising and refining my original prompt.” The AI then proceeded to do precisely that.
If I were doing this for your assignment, I would give these inputs to ChatGPT on my phone, and then feed them into ChatGPT on my computer as inputs, and then give you the resulting logs, as well as a paper answering the final, “refined” version of my prompt. I might rewrite them in slightly different words, introducing some imperfections in grammar and word choice along the way, to make them look somewhat more authentic (though I also might not bother, knowing that you wouldn’t penalize me even if they were entirely inauthentic).
Did this require me to exercise any critical thinking or writing skills? No.
Does this take any other sort of ability that we should be teaching our students? No.
Is it unrealistic to think that most students can, and many will, use this method to complete your assignment? No.
If you don’t share my confidence in all this, then you can perform the experiment yourself; if you’d like a larger sample, you can do the same thing many times over.
I’ll quote from the original piece:
Of course you can generate a paper with little effort, and of course that paper may even appear to be quite decent. I never denied that. But how do you know if it is really good, in the sense that perhaps only students who paid attention in my class (and in this sense, are domain-experts) would actually know? When I grade prompts I–like you, hopefully, when you grade essays or conduct oral exams–look for very specific things.
Hence, if a student diligently writes down most of what is said in class and then includes it in the process of AI meta-prompting, they’ll probably have a decent chance of fooling you?
Probably, yes. If you permit just a little bit of science fiction: if you “download” the contents and capacities of a student’s mind onto a computer and let it do that student’s homework, then yes the professor would be fooled. I don’t see how that is problematic specifically for this kind of assessment as opposed to almost any other, technical and philosophical worries about personal identity notwithstanding.
This is exactly my worry with your approach. A student can have plenty of “contents” in their mind, and they can have the capacity to do something with those contents, but nevertheless fail to understand it or do anything with it. And yet, at the same time, they could produce a very nice paper if they just get ChatGPT to do something with those contents for them.
I would like to measure the extent to which my students have actually developed their capacities.
I take it that you’re now arguing that, in fact, students are unlikely to be able to get ChatGPT to generate good papers, or good paper prompts, unless they’ve been paying attention in class, and understanding the material well.
This is not true. In order to see as much, repeat the experiment I described above several times. If you wouldn’t consider the output you get from ChatGPT to be A-quality work if it came from a human student (even granting that it’s not perfect work), then your expectations for undergraduate essays are unreasonably high, unless perhaps you’re at an extremely elite university, which few of us are.
I take it that your idea is that ChatGPT’s response, although it’s overwhelmingly likely to be good in a generic way, might overlook points from your readings or lectures, and thus be less than good in the context of your class specifically.
But here we face a question: did you ask your students, in the assignment, to address those particular points from class and the readings? If you did not, then it would seem unreasonable to detract from their grade for failing to do so, if they give an exemplary response to the assignment that you actually explicitly gave them. If you did, then students can quickly give ChatGPT the names of some of the articles you read, and perhaps a few hazily-remembered details from your lectures, and ask the program to address these concepts, and the program will oblige them. This perhaps requires a bit more of them, namely that they read some part of the syllabus, attend some lectures, and retain some memory of them, but that is not much.
Another one of your responses to this immediate and devastating criticism is to say that there’s too much uncertainty here for us to judge whether or not students can and will trivialize your assignment, with your full permission. But again, you yourself can play around with ChatGPT for five minutes and see the manifest evidence that this is so.
Honestly, despite the immediacy and devastating nature of your criticism, and the plethora of manifest evidence you presume to exist, I’m not particularly concerned. Again, prompting is not trivial, nor is meta-prompting. At some point, the work required to fool the interrogator would likely be more than to just produce the real thing. Call me foolish, but I prefer to work the problem, propose a solution, evaluate its effectiveness, and communicate it transparently. If that doesn’t convince everyone, I am OK with that.
Thanks for the exchange!
Let’s sum up.
You point students to the tools they can use to complete your assignments without doing any real work themselves, and explicitly permit them to use those tools in this way.
You respond to the quite conspicuous problems with this approach with denial in the face of plain evidence to the contrary which you thus far refuse to even acknowledge. (Have you still not just asked ChatGPT to generate some paper prompts for you, and looked at the output? If you have, how can you still deny that ChatGPT can complete your students’ assignments for them with their practicing any semblance of the skills we’re supposed to be teaching them? If not, why not?)
You also suggest that the matter is hopelessly uncertain in spite of said unaddressed plain counter-evidence, and claim that those who disagree with you simply does not understand the technology, when you can in fact see for yourself in under a minute that the only claims they’ve made about the technology are quite correct.
Thank you for the exchange.
(In closing, let me apologize to DailyNous’ moderator, whom I’ve been overworking.)
That does sound difficult. As I noted in an earlier comment, the ‘mass education on the cheap’ model that many universities have pursued over the last 40 years is unsustainable. To stay relevant in the era of AI, universities (or the governments that fund them) need to invest in more high-quality, time-intensive forms of education. Unfortunately, such reforms are unlikely at many institutions, which will instead continue to resort to austerity measures such as even bigger class sizes with less resources. The result will not be renewal but collapse, as institutions that refuse to change will simply disappear.
This is a nice rhetorical piece, even if the analogy is flawed and the understanding of the technology is somewhat lacking. The analogy would work only if we do nothing about the challenge posed by generative AI, and keep grading take-home writing assignments as if nothing were amiss. The understanding is simplistic because effectively prompting a machine for a high-quality result takes much more care and expertise than you admit.
The approach I am describing here is nothing like the paper mill analogy. I am not (only) grading the final product, but (also) the process by which it was created. This process is one in which students must demonstrate their ability to think, reflect, and apply knowledge. It also happens to be a process by which they engage with AI, but I don’t see how that is relevant as long as the former is true as well. So the analogy would work only if, when buying the service from the paper mill, the student would have to very precisely guide the service into what they are supposed to write, how they are supposed to write it, and then iteratively refine their drafts. Paper milling does not work this way, but prompting does.
Those who have an adequate understanding of the technology understand that, just as students can use AI to write their papers for them, they can also use AI to write paper prompts for them.
You’ve suggested in reply to this point that we should allow students to do this as long as they tell us they’ve done so. Let me quote one of your previous replies on this subject: “I let you use whatever you want, as long as you are transparent about it.” Thus, you’re apparently aware of the fact that students can use AI to perform this task as well, and quite fine with their doing so.
You seem to be under the impression that, in order to use AI in this way, students will still need to use the skills we’re supposed to be teaching them. This is not true. Again, just as students can ask ChatGPT to complete an ordinary paper-writing assignment without using those skills at all, they can also ask ChatGPT to complete a prompt-writing assignment without using those skills at all. There seems to be nothing to stand in the way of students acing your assignments using only the CTRL, C, and V keys on their keyboards, all with your wholehearted approval.
In brief, your proposal for how to deal with AI-powered academic dishonesty is to cease making any efforts to stop it, and let a crisis become a catastrophe.
See the exchange about meta-prompting above. But just to make sure the record is clear: it’s not academic dishonesty if the academy allows it, and there might be perfectly good reasons for doing so. Of course you might disagree with this purported goodness, as you do in the thread above, but you cannot just deny the conclusion and leave it at that.
It’s true that we can eliminate all academic dishonesty in one stroke by permitting students to take credit for work they didn’t do without restriction. We could also banish crime from the world forever by eliminating all laws. The question is whether the effects that would follow are desirable.
You suggest that there “might be” good reasons for an approach which permits students to use AI to complete their work rather than do so themselves. What are these good reasons, specifically? Are these uncertain too, in much the same way that it’s uncertain, in your view, whether ChatGPT can generate A-quality prompts for papers, despite the fact that entering a single query on chatgpt.com demonstrates that this is indeed the case?
Look, whether or not higher education risks dying, it is *even more important* how well our society functions as a whole.
Even if we just zoom in a tiny aspect in our workplace, the work slop caused by AI use is very financially costly and has a non negligible effect on our economy. It is thus important to guide people on proper use, and to deliver results that are actually good. Remember that these AI slops are produced by people many of whom have received proper (aka AI-free) education in the past.
Exactly. People like to complain about AI slop, and certainly the AI itself is to blame for much of it, but so are the people who use the AI poorly, unreflectively, and irresponsibly.
Well I agree with your new comments, but will take a break from this debate. I think it is going nowhere partly because people are facing *very* different circumstances. Some of us are (much) luckier than others. For example, unlike Abraham, I never have to teach more than 20 students without TAs, and can afford to assign complicated assignments, and in general have loads of free time to play with AI. These parameters are just too different. So I will leave it be and shut up.
There’s a certain skill involved in making wishes from a genie. You can’t just wish to become the best basketball player in the world; if you do, the genie might simply kill off everyone who plays except for you. If you want to enjoy the sort of future as a star athlete that you’re really envisioning, you’ll have to be able to avoid this sort of monkey’s paw situation. Note two things, however. First, the skill of wording your wishes with sufficient care is still very different from skill in basketball. Second, the former skill will be fairly trivial to learn if, in fact, you have a genie who is reliably able to tell what you really want, and willing to give it to you even if you overlook certain lawyerly loopholes in your request.
The same applies to AI. The skill of using AI well is trivial to acquire, especially when someone has already gone through the trouble of writing out the exact questions that you need the AI to answer, which professors very serviceably do when they write their assignment prompts; in that sort of situation, you’d have to go out of your way to mess things up. The skill of using AI well is also entirely different from the skill of writing well and using critical thinking skills well, which is what we’re supposed to be teaching our students.
You might reply that you have to do those things to come up with an input that will induce the AI to give you the output you want, but of course you don’t; you can just ask the AI to come up with an apt input, too. There is no sense in following Prof. Zednik in attempting to cast doubt on this point or minimize its fatal significance.
You might reply that these aren’t the skills we need to be teaching students so that they can fare well in our brave new AI-dominated world; I would consider this a concession that pro-AI academics are simply renouncing any effort to teach students the deeply important and deeply human skill of thinking for themselves and putting the thoughts they form in the process into words.