Harness AI to Help Your Students Learn Basic Logic (and more) (guest post)


Michael Rota, professor of philosophy at the University of St. Thomas, has been using AI in his teaching, and is trying to make it easy for you to do so, too.

His experiments (previously) have led to the development of a free app you might want to try out.

In the following guest post, Professor Rota talks about the app, called Personify, what he uses it for, what his students think of it, and how it might be used for tasks besides logic instruction.


Variation on a diagram by Hans Reichenbach. Click the image for more.

Harness AI to Help Your Students Learn Basic Logic
by Michael Rota

I’ve been using generative AI to help my students learn basic logic in my two sections of introductory philosophy this semester.

Whereas in the past I gave paper homework assignments and graded them by hand, I now deliver the same sorts of homework assignments through Personify (https://personifyai.app), a free web app I helped develop that allows students to get immediate, personalized feedback as they work through course assignments. And the AI does the grading.

Although built on top of ChatGPT, Personify AI is much more reliable than ChatGPT alone, because as the AI evaluates the student’s answer to a question, it is armed with a good answer to that question provided beforehand by a human being. So the AI doesn’t have to solve the problem correctly itself, it just has to guide the student to the instructor-approved right answer. (And faculty have full control over what that instructor-provided answer and explanation is.)

Two of my colleagues here at the University of St. Thomas are also using the exercises, and we gave an anonymous survey to three of the sections, asking this question:

Compared to doing similar assignments in a traditional format (e.g. pen and paper, or non-AI software), how much did you prefer or not prefer using the Personify AI tutor?

The possible answers were “Strongly Prefer Traditional”, “Prefer Traditional”, “Indifferent”, “Prefer Personify”, and “Strongly Prefer Personify”. Of the 79 respondents, 1% said they strongly prefer traditional, 4% said they prefer traditional , 9% said they were indifferent, 44% said they prefer Personify, and 42% said they strongly prefer Personify.

The students seemed to like the immediate feedback and the tutoring experience provided by the AI:

“it taught me to do the questions rather than telling me the right answer right away.”

“I like the instant feedback it gives you since it helps me quickly realize my mistakes and fix them rather than getting everything wrong and not knowing until later.”

“I liked that it was easily accessible and provided interactive feedback which was better than just a ‘correct’ or ‘incorrect’.”

You can try out the student experience of the logic exercises by clicking on this link (no account required):

I recommend trying a question or two from some of the later assignments (assignments 15-18) to see how it performs on harder problems.

Here’s a representative student chat (shared with permission) on a question most of my students used to miss when I included it in my old paper logic homeworks:

If you’d like to use the assignments in your own course, you can click on the next link (below) to import a copy of the course, which, in addition to letting you preview the student experience, will also allow you to (a) share the assignments with your students, (b) use the Grades feature to see their scores and view submitted chats when desired, (c) add, delete, or edit the questions and the instructor-provided answers, and (d) delete whole assignments or create new ones. For example, a given instructor might want to use some but not all of the assignments I’ve been using. To import the course, click here and then create a free Personify account by authenticating with any Google or Microsoft email account you already have.

You can email me ([email protected]) with any questions. Also, within Personify there is a little chat icon in the lower right (in a purple circle), which will allow you to send a message to a real human being at Personify, if you have customer service questions. And there is a help icon on the lower left which links to a few pages of documentation on how to use the product.

The app can also be used to good effect for reading engagement questions, to prepare students to discuss the reading in class, and to incentivize them to do the reading in a way that doesn’t require the professor to do more grading. I created some assignments on Aristotle’s Nicomachean Ethics, Books I – II, for example. A colleague in San Diego has created questions on almost all of the reading assignments she assigns. I also use Personify to deliver a practice logic test, so the students have better information about what they should expect for the logic test I give in week 4 of the semester, and so they get tutoring on the problems they get wrong on the practice test.

In case you’re interested in creating your own questions and answers, here’s a little more context on how the product works: you the professor can input a question, and then input an instructor-provided answer or set of evaluation criteria. When the student is taking the assignment and answers the question, a message is sent to ChatGPT which includes the question, the student’s response, the instructor-provided evaluation criteria/answer, and instructions to ChatGPT to evaluate the answer as correct or incorrect by referring to the instructor-provided answer. ChatGPT is also instructed to Socratically guide the student to a correct answer if the answer is evaluated as incorrect.

For questions that don’t require a specific answer, the instructor can say things like this: “Any plausible answer that [state condition] should be counted as correct.”

Personify assignments are meant for low-stakes assignments, typically to promote learning outside of class (whether through homework, reading engagement questions, or practice tests.) Feel free to experiment with this and reach out if you run into any problems.

guest

18 Comments
Oldest
Newest Most Voted
Inline Feedbacks
View all comments
David
David
1 year ago

The representative sample is a student simply blurting out answers, which might be guesses, and the AI bludgeoning them with extremely leading questions or even giving the answer away itself until the student gives the right answer. Seriously, in the response the AI says “We cannot confirm the argument as sound” and then turns around and asks “What can we say about the soundness?” This might be helpful for very early problems to ensure students have a rudimentary comprehension, but the pedagogical use to me seems more or less limited to that.

Getting these leading nudges *feels* better than getting ‘correct’ or ‘incorrect,’ which makes sense of the student reports. But it is not actually better at promoting learning for more independent problem solving (learning to solve problems independently often feels quite unpleasant). Much better is immediate feedback that the answer is incorrect and the student trying to identify where they themselves went wrong and seek guidance if they simply aren’t able to identify the mistake, at which point it’s still best to have an instructor who does not jump right in with feeding the student the answer.

Daniel Weltman
Reply to  David
1 year ago

Yes, if the representative sample is anything to go by, this seems like it might be teaching students how to wheedle the answer out of the AI, rather than teaching students about logic or whatever. I would be interested in seeing tests of how students perform in a course using this AI compared to how students perform in a course without it. That would help alleviate the concern that this is not promoting learning.

Certainly polling students on whether they prefer this kind of feedback does not strike me as any evidence at all. You could poll students on whether they want to do problem sets, and if they strongly prefer not to, this would be no evidence that problem sets are bad for learning logic. It would just be evidence that students don’t want to do work.

Michael Rota
Michael Rota
Reply to  Daniel Weltman
1 year ago

This is the first semester we tried out these exercises (2 colleagues in my department used them as well), so we don’t have much evidence. But we do have a little: one of those colleagues last taught the course in Spring 2023. In the Spring 2023 course, the class average on the logic test was 80.1%. In his Fall 2024 course, this colleague tells me, the methods used in class were the same, the material taught was the same, and the test was the same, given at the same point in the semester. The only difference in terms of pedagogy was that in Spring 2023 paper homeworks were used, and in Fall 2024 the Personify assignments were used. The class average on the logic test in the Fall 2024 course was 84.2%, which was a .36 standard deviation increase in this case.

sahpa
sahpa
Reply to  Daniel Weltman
1 year ago

Also agree that evidence is needed, but I wouldn’t immediately cast suspicion on this kind of back and forth. After all, what do we think teachers/TAs are doing with struggling students in office hours if not something like this — the leading questions, the ‘giving it away’ (from our expert’s POV), etc.?

Micheal Rota
Micheal Rota
Reply to  David
1 year ago

I hear what you’re saying, David — perhaps the AI gives away the answer too easily. But perhaps not…if the problem is too hard (but still counts for a grade), some students will just use ChatGPT to cheat. And they can do that – cheat – on almost any assignment we give, including traditional homework. But we still want to promote learning outside of class. That’s what this product is primarily meant for. (It doesn’t replace the hopefully more effective human instruction in class.) So when I compare what this student would have gotten with the logic homework on my old way of teaching to what he’s geting now, I see it as an improvement. Let’s call the student Sam. If Sam had been in my class last year, he likely would have just gotten this question wrong and turned in the homework. Then later I’d talk about the problem in class, and I’d either just explain it (in which he would be given the answer) or I’d ask the students to explain it (in which case probably a student other than Sam would give the answer). So either way, Sam’s going to initially get it wrong and then be given the answer. The ideal scenario in which a skilled human Socratic tutor works one-on-one with a Sam would indeed be much better, but with 30 students in the class, I can’t do that, much less do that for every point of logic I’m covering. So the goal here is to help students get more out of the homework experience, as compared to the status quo.

David
David
Reply to  Micheal Rota
1 year ago

I’m not sure it really is an improvement in the case of Sam. The pedagogical situation you are bemoaning can be more effectively addressed by other means (the worry about cheating too). Before I turn to other means, I want to point out that what the LLM is doing here can basically be done just the same by using a model answer directly. The hardest kind of case for getting students without motivation to learn are problems that are effectively multiple choice (which the problem in the post is). If you have a mechanism to give them immediate feedback, you will thereby allow themselves to guess their way to the right answer through process of elimination. This is what the LLM enables. The putative advantage of the LLM is it can explain why the right answer is the right answer. But so can a model answer placed in the ‘answers’ section of a textbook, or shown to the student automatically by a quiz in an LMS like Canvas. In Canvas (and perhaps others) you can even have it give pre-written hints/nudges for incorrect answers which accomplishes more-or-less what the LLM is accomplishing here.

Really, though, there are superior (if still imperfect) solutions to the problem of students who have low motivation and to LLM cheating. My favored way is to make learning the material necessary to do well in the course in a way that will be clear to the students. For example, in my logic classes, student grades are almost entirely determined by their performance on in-class, proctored tests which are pass/no pass. Passing requires solving almost all of the problems on the test fully correctly. I make very clear that the only way to prepare to solve the problems they must solve on their own for the test is to solve the problems on their own for the homework. I then reassure them that if they become stumped on the homework, they can come to office hours and show me their work and I can give them guided tutoring to get them back on track. The homework grading is handled using carnap.io which will give instant correct/incorrect feedback to students each time they try to submit their answer, so they will know when they get the correct answer without having to wait to come to class or for me to explain it (for this reason, I do not explain or post homework answers).

For reading engagement, you can also use in-class assessments that can’t be cheated. In philosophy classes, I teach my students how to annotate readings and require them to bring hard copies of the annotated readings to class and then at the start of class, they complete a short quiz that tests them on comprehension of specific points in the reading where they can consult their annotated reading. The quiz questions are the kind that can only be answered reliably if the student read carefully (i.e. I avoid ‘big picture’ questions that you could get from a summary). The quizzes are all multiple choice, so grading them is very quick.
As far as in-class feedback, a class size of 30 actually makes one-on-one tutoring perfectly viable, you just need to flip the classroom to some degree and dedicate some class periods entirely to solving problems. Students who get stuck can then raise their hand, you can float over and try to give them a nudge in the right direction. I do this regularly in my logic classes of that size. It works great.

There are many pedagogical methods to address these kinds of problems that have at least some empirical research supporting their efficacy that go beyond student and faculty satisfaction surveys. My worry about rolling out these LLM solutions is that we are doing so without sufficient reason to think they are actually pedagogically sound (as Daniel Weltman highlighted). It is especially worrisome to use satisfaction ratings/feedback to evaluate them because we know that learning is often uncomfortable for the learner (and we know that instructors like to do less grading). This is an even more acute worry for any implementations that don’t involve relatively clear-cut correct answers. LLMs are prone to hallucinations and drift in ways that do not give us reason to trust that they may in their tutoring end up giving students misleading or outright incorrect guidance. They need serious vetting before we come to rely on them.

M G
M G
Reply to  David
1 year ago


Can you give some more details about these reading quizzes? I’m thinking of moving more in this direction, and I’m looking for ideas.

David
David
Reply to  M G
1 year ago

Sure! They’re ~5 questions (the total score is out of 5 points, sometimes I have a more complex 2 pointer and use formats like matching or multiple answer). I give students 10 minutes to complete them at the start of class. I use specifications grading, so the pass threshold is 4 points out of 5, though sometimes for harder readings I will bump the threshold down. The questions I pick are a mixture of (1) testing for careful reading: e.g. did students catch a crucial qualification, can they easily match X objection to Y argument (2) ensuring comprehension of concepts important for any activities/discussion I have planned, so these will be more oriented towards testing basic comprehension by having students fit a novel example to a concept or the like.

Going over the answers in detail can take most of a class period, so you have to make a judgment call about how to handle that and how to integrate discussion. I am currently trying an A/B weekly structure where the first day all reading is due and they take and we go over the quiz, and the second day we engage in activities & discussion.

This approach was inspired by readiness assurance tests in team-based learning (https://learntbl.ca/what-is-tbl/ensuring-student-readiness/). I do also incorporate other elements from TBL, so I have students in term-long teams and they take the quiz first individually and then as a team.

The other thing they like is that I let them get credit for their team’s performance contingent on not free-riding. I’ve approached that in a couple of ways so far and both have worked. One is to let them show me their annotations at the start of class. In this case I insist that they flag according to the method I teach in the first week inspired by Concepción’s “Reading Philosophy with Background Knowledge and Metacognition.”

The other approach I use in courses where I’m using team work more heavily. In those courses, teams draft constitutions with member expectations at the start of the term. At the end of the term, the team votes anonymously to pass or not pass each team member in terms of meeting the expectations. Non-passing members don’t receive credit for their team’s performance and are judged entirely on their individual performance. There’s a mid-term check-in where teams have to discuss member’s standing in relation to team expectations (so no one gets caught entirely off guard).

Michael Rota
Michael Rota
Reply to  David
1 year ago

Your courses sounds fantastic, David. I too have come to the conclusion that making most of the class grade come from in-person exams is a good way to deal with the increased ease of cheating on homework and other take-home assignments. So that’s what I do, but, like you, I also find value in including assignments that promote learning outside of class. In the blog post I neglected to mention that I’m using this not in a logic class, but in an Intro to Philosophy course, where I just want my students to understand basic propositional logic and be able to apply it. So I’m giving them questions on validity and soundness, argument forms, and eventually questions that ask them to look at an argument in ordinary English and extract the form. So by week 4 in the course we’ve worked up to questions like this (in assignment 18 in the linked course tutor):
“Consider the following argument:

I won’t get fired! Because I either mailed the letter yesterday or two days ago. And if I mailed the letter yesterday, then it will arrive soon enough. And if I mailed the letter two days ago, it will also arrive soon enough. And if it arrives soon enough, I won’t get fired.

Let F = I will get fired.
Y = I mailed the letter yesterday.
T = I mailed the letter two days ago.
A = The letter will arrive soon enough.

Now add any implicit steps needed to make all inferences valid in virtue of form, and write out the argument in logical order, as a series of propositions, using those letters and the logical operator words. Number your steps. Use “Therefore” for any inferences. In every line that contains a “Therefore”, say what inference form and what steps are being used (e.g. modus ponens from (1) and (2)).”

Before I graded these by hand (both to provide feedback and because I find that attaching a grade incentivizes students to do the problems). Grading these more open-ended questions took a lot of time, because some of them have different possible solutions, and because students can get the answers partially right and partially wrong in all sorts of ways. Now, I don’t need to spend time grading those homeworks. And the students get lots of practice, plus accessible feedback that speaks to the precise errors they’re making. The rate of hallucination is much, much lower than in an ordinary conversation with ChatGPT, because with Personify the AI has been given the right answers (or evaluation criteria).

JDRox
JDRox
1 year ago

Huh, I guess I was more impressed with the “representative sample” than you guys were. But yes, one can spam the AI until it leads you to the right answer. Of course, that won’t help you learn, and so students who do that should expect to do poorly on the exams/quizzes. But of course, there’s not much we can do to help students who don’t want to learn, is there? Still, this concern has something to it. I like Harry Gensler’s (RIP) LogiCola for related reasons. It penalizes every wrong answer such that to complete a homework assignment one basically has to have a certain level of proficiency (that level is adjustable by the student, and they’ll receive different amounts of credit depending on the amount of proficiency their homeworks demonstrate). Sadly, LogiCola doesn’t work well on Macs and is probably going to fall into disrepair now that Gensler is gone…

Nathan
Nathan
1 year ago

Let’s imagine that you’re right (contrary to the reasonable concerns of David and Daniel) and there’s significant instructional value in these types of Personify assignments. Is that added value worth the environmental cost, compared to the (again on assumption) lessened instructional value of traditional instruction? Given how resource intensive this technology is you really need the value difference to be enormous to be worth it all things considered.

Daniel Weltman
Reply to  Nathan
1 year ago

And of course the worry that the end result of relying on, legitimizing, refining, and otherwise expanding the use of these technologies is unemployment for people who would otherwise be philosophy instructors or TAs. Maybe it’s pointless to try to fight against this sort of thing via refusal or whatever, and maybe the benefits outweigh the drawbacks, and so on. But it is worth reflecting on whether this sort of thing is step one on the road to putting 95% of us out of business.

(I think the question gets more salient if one thinks that the eventual end state is not one where people have in fact been rendered pointless because their jobs can in fact be done better by AI, but one where worse AI has replaced better people because the few use cases where AI was good enough made it look attractive enough for those in power to mandate its use in cases where it’s clearly not good enough.)

Kenny Easwaran
Reply to  Daniel Weltman
1 year ago

The point of logic classes is not to produce employment for philosophy TAs. The point of logic classes is to teach students. If we can’t think of anything that TAs can do to improve student instruction once students have access to LLMs, then that is an indictment of our own imagination, not of the LLMs.

Daniel Weltman
Reply to  Kenny Easwaran
1 year ago

I think things can have multiple points. I think one point of many required courses at many universities is to secure employment for philosophy TAs. At the university where I did my PhD, the Philosophy department has wisely managed to get two ethics courses into the curriculum for a lot of the undergrads no matter what their major is. There are many reasons this is good: exposing future engineers and doctors to philosophically sophisticated ethical reasoning might help them think about ethics more clearly, studying some humanities is good for everyone no matter what their major, etc. But one of the reasons this was done was so the grad students would have something they could TA for (and, sometimes, teach). And it’s good to care about this sort of thing. Departments that don’t care about this sort of thing now are departments that might not exist in 50 years.

Plus, my worry is not that there’s nothing else a logic TA can do for a course once the robot takes their former job. I’m sure we can think of plenty of useful things for them to do. My worry is that the people whose job it is to destroy our existence for the sake of saving the university money will not be impressed to learn that your logic class can get 1.3 times better if you have the TAs do new stuff and have the robots take over the old TA jobs. Instead they will be impressed that your logic class can be equally good if you fire all the TAs and replace them with robots. Their goal is to fire as many expensive humans as possible, not to make instruction as good as possible. Indeed they’d probably accept robot logic TAs that lead to a course 0.8 times as good as the human course, given the savings. And probably they’d be thrilled to employ robot instructors that lead to a course 0.63 times as good, given the economies of scale (now they can sell the course online! to as many people as they want!) etc. In this area, many are willing to race to the bottom to save money on TA and professor salaries.

Anyways, my main worry is not about employing logic TAs. For all I know, you can replace them with robots in a way that promotes learning. Perhaps logic is a pretty rote robotic process in the first place and AI is a great teacher for it. (I wonder whether AI can handle advanced logic stuff, like modal logic, but whatever.) My worry is that the people whose job it is to replace us with robots will not make fine distinctions between subfields in philosophy, and on the basis of failing to make these distinctions they will not be particularly compelled by the claim that a robot can’t replace an ethics TA even if it can replace a logic TA. And I am less sure a robot can replace an ethics TA.

Kenny Easwaran
Reply to  Nathan
1 year ago

Is the environmental cost of a student using a Large Language Model for an entire semester more or less than the environmental cost of a single hardcopy book, or of riding the bus to class?

I’ve found estimates ranging from 1-10 g CO2 emissions per query of a Large Language Model. I’ve found estimates of about 1 kg of CO2 emissions for a printed book. I’ve found estimates ranging from 30 g to 500 g of CO2 emissions per passenger-mile for riding a (non-electric) bus.

So 1kg of CO2 emission gets you at worst 100 queries of an LLM, or a book, or at best 30 miles on the bus. (This is about 1/8 of a gallon of gas, or about 5 miles in a Prius.)

It’s good to be thinking about carbon emissions, but you want to actually *think* about them. An LLM-heavy class that encourages students to make dozens of queries a week probably produces nearly as much emissions as asking the students to be physically present in class. If you’re at the point of reducing emissions where you’re worried about these emissions, you should also be considering eliminating the in-person meetings of class (unless the students at your university all walk or bike or use electric vehicles).

For what it’s worth, the global average is about 5,000 kg of CO2 emissions per person (Europe is close to the world average, Mexico and India are lower, China, the US, and the Middle East are higher.)

David
David
Reply to  Kenny Easwaran
1 year ago

This just seems like whataboutism. First, nothing in Nathan’s comments suggests they don’t try to reduce emissions where they can (feel free to point a finger at me for requiring hard copies of readings).

Second, Personify (and LLMs in general) do not *replace* sources of emissions like buses or paper books (pdfs are the replacement for the latter that’ve been around for a while). LLMs are being added *onto* the emissions generating infrastructure we are already failing to keep under control and what they do seem likely to replace (Google searches, writing emails and so on) are all much lower in terms of emissions. As we have learned the hard way: once we have widespread dependence on technology because we have woven it into our society, it is very, very hard to extricate ourselves from it. This is an important asymmetry: it is not easy to do without things like search engines or buses (or worse, cars) because of how our society is arranged, but it is *quite easy to do without LLMs* and so it is perfectly natural in a blog post about LLMs to raise concerns about their emissions without also bringing up the many, many other areas where we should also be reducing emissions.

EHG
EHG
Reply to  David
1 year ago

It is not whataboutism. It is just providing valuable context about the actual environmental cost of LLMs. It seems extremely valuable to know how large the environmental cost actually is when deciding whether or not to avoid LLMs due to their environmental cost. Since most of us do not have a great sense of how much a grab of CO2 emissions is, it is also useful to compare against the costs of other tools used in philosophy classes.

Adam
Adam
1 year ago

Confused…it does not seem to me that, in the example, the soundness of the argument hinges on whether there has been life on Mars, but rather on whether there having been life on Mars is entailed by 2+2 not equalling 5, so the argument is definitely unsound (unless you consider that, in a coherently structured universe, any true fact entails all other true facts—which seems to me outside the scope of what is being driven at in the example question).

Relatedly, I would not trust ChatGPT to do anything where precision and subtle distinctions are the point…