News for & about the philosophy profession

Trouble in the US Non-Academic Job Market and What It Means for Philosophers (guest post)

“I’m here to sound a warning bell for those who are now in the position I was several years ago.”

In the following guest post, Noah Gordon, who earned his PhD in philosophy from the University of Southern California, shares his experiences on the non-academic job market, thoughts on our current economic and technological situation, and advice for those considering leaving academia.

There are lessons in his reflections, as well, for philosophy departments interested in supporting the pursuit of non-academic work.


Trouble in the US Non-Academic Job Market and What It Means for Philosophers
by Noah Gordon

Disclaimer: This piece is based primarily on my own experience on the US non-academic job market, my evaluation of news related to this market, and my evaluation of many other anecdotal reports.[1] While I have tried to assess my claims in light of rigorous, reliable studies and data, and sometimes reference such work below, my attempts to do so have revealed that such resources are lacking. My sense is that radical changes in this job market have been occurring only fairly recently (in the last few years), are currently diffuse (highly dependent on specific roles and industries), and that it will take some time before these changes are captured by the slow process of economic measurement in academic journals, government reports, and so on. I am not an economist or a labor theorist, and I strongly encourage you to do your own research and analysis before relying heavily on the conclusions and analysis here.

[If you are only interested in my concrete advice, I advise skipping to the last section and then reading back for justification.]

One Philosopher’s Experience on the US Non-Ac Job Market Post-COVID

I finished my PhD in philosophy at USC in May 2023. I considered it a success: I contributed to the philosophical dialogue with some publications and presentations, met some good people, and got to spend years studying some fascinating esoteric topics. I did not go on the academic job market. Instead, I had a plan to spend some months preparing before starting applications to non-academic, white-collar jobs in the US. Primarily, this meant doing research on relevant industries and positions, brushing up on and obtaining online certifications for relevant technical skills, doing personal projects and freelance work that would be attractive to employers, and updating and polishing job app materials, like a personal website, LinkedIn profile, and generic resumes. I had a sense of what types of positions I wanted to search for: primarily something like ‘Technical Writer’ or perhaps that strange-sounding title ‘Ontologist’, which frustratingly goes by way too many different names on online job boards (you’ll also see titles like ‘Taxonomist’, ‘Semantic Data Modeler’, ‘Linguistic Engineer’, ‘Knowledge Graph Manager’, and many other variations). I had some loose geographic preferences within the US: West or East coast, if not remote.

Although I was clearly a bit unfocused, I felt that my chances of securing some or other decent job in the specified market were pretty good. I had some white-collar work experience outside the confines of academia, unlike many other philosophy grad students. I had an undergraduate major outside of philosophy in a relevant field to the positions I was interested in, specifically, in IT. And I had (and this is the tasteless part where I brag about my academic credentials), all the academic insignia which are sometimes purported to be valuable on the non-ac market, including a high GPA and a higher degree from a fancy name-brand university (though not an Ivy-league one).

The results of my on and off applications over the past 1.5 odd years are a hodgepodge of unstable and generally lowly-compensated freelance and contract positions with poor working conditions. Even initial phone-screens have been few and far between over hundreds of applications, and that’s before you even get to the many layers of interviews white-collar positions in the US now require (generally, after an initial recruiter phone screen, you will have at least one interview with HR and one with the Hiring Manager, however this is a minimum and often there are many more rounds of interviewing).

While I would love to complain more about my particular circumstances, that’s not the point of this piece (it’s just the hook). Instead, I’m going to give my analysis of the broader conditions behind my experience and some things those within philosophy with an eye on the non-ac market (including current and prospective grad students and undergraduate philosophy majors) should be thinking about. Especially if you are a current philosophy graduate student thinking about the non-ac market, even as a plan B, I strongly encourage you to think about these issues carefully. Academic philosophers are, bless their hearts, mostly blissfully unaware of the non-ac job market, and many of them will only have experiences to share about previous students who went non-ac. However, this particular configuration of troubling conditions on the non-ac job market are quite recent (post-COVID, especially 2022 and later), so those experiences may no longer be representative. In short, you probably won’t hear it from them and you should take any older experiences with a grain of salt. I’m here to sound a warning bell for those who are now in the position I was several years ago.

The Situation in Tech and Three Deleterious Trends for US White-Collar Work

The job market in tech in the US is undergoing a major contraction. From a high-point in 2022, job-listings on major job boards for positions like Software Developer and IT Operations are way down:

Source

Source

Meanwhile, mass layoffs by tech companies began in earnest late in 2022 and are ongoing. Some commentators are, or at least were, convinced that this is merely a correction due to some combination of tech “overhiring” during the pandemic, copycat behavior between businesses, and some very general macroeconomic conditions around high inflation and low consumer demand. I see little evidence to think that this is the case, and stronger evidence that this is not the case in the excellent economic performance and profitability of these companies throughout the relevant period. As one representative example, Meta recently laid-off 5% of its workforce despite reporting large increases of 48% in income and 59% in profit in 2024.[2] In other words, even taking the premises about the macroeconomic conditions at face-value, “overhiring” does not seem to have hurt the performance of these companies in a way that explains the layoffs.

I believe that instead the tech companies are direct witnesses to, and early adopters of, three concerning labor trends that predate COVID, but have been greatly exacerbated and enabled by it. The tech companies, I believe, are going to function as test cases for slower, more risk-averse, and less technologically competent industries. If the success seen in the tech industry continues and these trends intensify as I expect them to, you can expect the situation in the tech industry to spread to many other sectors in the coming months and years.

Remote Work and Offshoring

US companies have experimented with offshoring white-collar work to cheaper labor markets for decades. My understanding of the history of offshoring for white-collar work is that its success has been mixed, and that it largely has not had noticeable negative effects on the US white-collar market.[3]

However, things seem to be different post-COVID.[4] The pandemic effectively functioned as a nation-wide accelerator and test case for making many white-collar positions fully remote. The infrastructure for doing all sorts of positions remotely was developed rapidly, and employers and employees alike became familiar and competent with relevant processes like replacing in-person meetings with video calls.

Many employers found that a good number of white-collar positions could be done at a level that was at least close to the performance they saw for in-person work.[5] The same tech companies that are laying off thousands of workers are increasing their hiring in cheaper labor markets like India and Mexico. The inevitable reasoning becomes “If this work can be done fully remotely from the US, why can’t it be done from Krakow or Lima? We could cut our salaries immensely and still offer top-market rates in those areas.”

These offshoring efforts are now more likely to succeed because of the increasingly mature infrastructure for remote work, and also the increasingly skilled labor pool in cheaper markets. Of course, over the long run economic logic dictates that things should even out in a globalized labor market, but there is a lot of pain coming for those in current high-cost labor markets in the meanwhile.

AI and Automation

The tech businesses that are downsizing are the very same ones currently building frontier AI systems explicitly designed to automate a massive amount of white-collar work. This could just be a coincidence, but I suspect that the front-row seat decision makers within these businesses have to the progress of AI development has given them a sense that they will soon, if not already, be able to do more with fewer human workers.

The effects of AI automation on labor are a subject of massive debate amongst economists and others. The economic consensus appears to be that automation has, in the past, led to the creation of new types of work that have made up for the loss of previous human work. It remains to be seen whether this will also be true given the fundamental differences between AI and previous technologies.

What is increasingly less a matter for reasonable disagreement is that AI systems that can perform the cognitive component of most knowledge work at or above the level of human experts are indeed coming in the near future (read: within the working careers of anyone under 40). Philosophers should not stick their heads in the sand on this point. I was much more of an AI skeptic just 6 months ago, so let me share some of what has changed my mind.

  • The position I just described is increasingly the consensus of AI researchers and scientists working in the field.[6] I am not talking about business people and “thought leaders” who are incentivized to spout high claims about AI to sell their products and narratives. I am talking about sober-minded people with verifiable credentials and research accomplishments who see the progress on these systems every day and have only weak incentives to exaggerate.
  • My personal experience in areas that I am competent in has made me more confident in AI progress. When GPT 3.5 was released in November 2022, its performance on some teaching materials for a course I was TAing were not impressive. Already by the next semester with the release of GPT 4, AI models were able to produce papers better than my average undergrad student. In the past 9 months or so I have worked on several freelance projects helping to train and evaluate AI models on philosophy tasks, and the progress is alarming. The newest models can accurately answer questions about which conditions to impose on possible-worlds models to guarantee specific principles about counterfactuals within Stalnaker’s semantics, and can identify an accurate summary of minute points within a paper by Selim Berker about normative principles from amongst ten plausible-looking options. I also freelance doing what’s called ‘technical content writing’, which involves writing those annoying blogs you might see on a business’s website if you Google ‘what is AI explainability’. Frontier models can now produce a draft in minutes of a quality that takes me nearly a full day of work to equal.

Though I still think at the moment of writing a reasonable case can be made for a more pessimistic view on the timeline towards such AI systems[8], this case is rapidly becoming less plausible. And let me be clear: if your view on AI progress is informed primarily by interacting with older AI systems, talking points about “stochastic parrots”, or funny pictures of AI systems being unable to count the number of times ‘r’ appears in ‘strawberry’, then your views are out of date and epistemically unjustified.

Even if you are not convinced by my brief case for optimism about progress towards AGI, you should bear in mind the following points. First, the effect of AI on the job market in the near-term depends more on employer perception of AI progress than the reality of it. And second, AI progress will be uneven and even if some areas of white-collar work remain resistant to AI automation, others are quickly falling. Later I will give some reasons to think that the areas quickest to fall will be the ones that philosophers are most qualified to work in.

Gigification and Contractification

Academics are aware of the increasing use of adjunct labor on the academic job market. In parallel, labor theorists talk about the rise of the “gig economy”. The proportion of work done outside the confines of a traditional employer-employee relationship in the US is increasing. In the traditional model, you work full-time for just one employer. Your position has no set end date, and it is reasonable to expect in many cases that your employment situation will be fairly stable. Your employer provides health insurance and other benefits like retirement contributions.

By contrast, in these alternative work arrangements, your situation is far more precarious. Worryingly, it seems that, particularly post-COVID, many positions in the US white-collar market that used to be secure full-time positions are being converted to these alternative arrangements. This is particularly noticeable for entry-level positions of the kind that philosophers entering the non-ac market would be competitive for (you are fooling yourself if you think that your PhD in philosophy will vault you into the hallowed halls of middle management). Let me highlight some of the issues you are likely to face in two of these alternative arrangements:

  • Freelance / Independent Contractor: These are arrangements where you accept tasks or projects on an ad-hoc basis. Some of the larger platforms that manage this work include Upwork, Remotasks, and Mechanical Turk. While the flexibility afforded by freelancing is nice, you will receive no benefits, including no health insurance. Moreover, the work tends to be inherently unstable, with projects coming and going on a dime. You often have no way of communicating directly with the client that commissions the work, leading to numerous problems. And the larger platforms are inhuman labyrinths of logistical routing that sprout kafkaesque horrors at every turn.
  • Contract Positions: These are arrangements where you are rented out by one company (the contractor) to do full-time work at another company (the client), usually for a set time period. There are a number of reasons why the client might want a contractor rather than a regular employee, but the bottom line is that they aren’t prepared to hire you as a full-time employee and are not committed to your long-term success at their business. You may be treated as a second-class citizen at your primary place of work, even (or especially?) if the client is large and wealthy. For example, you might be denied benefits like parking or have limited access to internal resources and opportunities. The client is likely to decide at any time to switch contracting suppliers or decide the whole project is going to be scrapped; after all, that flexibility is a primary reason to use contractors rather than regular employees. Meanwhile, the contracting firm is your official employer responsible for legally required benefits like health insurance. Existing as a mere skeletal intermediary with no particular expertise besides filling seats with cheap labor as quickly as possible, these companies can be roiling cesspools of incompetence, indifference, and general sleaziness.

There is good reason to think that, within the relevant market, the proportion of work done using these alternative arrangements will continue to increase post-COVID.[9] The reason is that the same infrastructure built up to support remote work also provides the needed basis for more easily converting regular positions to alternative work arrangements.[10] If it can be remotely by one person, there is a decent chance it can be chopped up into tiny pieces and done by dozens of online workers that you don’t need to give health insurance to.[11]

Hopefully you manage to cobble together enough meaningful freelance and contract experience to become competitive for regular positions. But either way, there is a good chance you may spend years on the job market dealing with these conditions first. One takeaway: you should not be confident that you will avoid years of adjuncting merely by going non-ac.

The Overall Unemployment Rate: A Bad Objection

I will now pay homage to my philosophical training by anticipating an objection. However, I will dishonor this training by considering not the strongest objection, but rather the one I think will most quickly pop into the average reader’s head.

Objection: The overall unemployment number is near a 20-year low! Surely things aren’t as dire as you are implying.

There are many reasons why this is a bad objection. Let me highlight two of the most important.

  • I am concerned about the availability of decent The overall unemployment figure includes part-time workers, freelance workers, contract workers, and does not consider working conditions or wages. Someone driving for Uber a few hours a week while unable to find white-collar work in their desired field counts towards the overall employment rate. I would count towards the success side of this figure during almost my entire job search. The overall unemployment figure is not a good measure of the availability of decent jobs that philosophers would be competitive for.[12]
  • By many estimates, more than half of Americans are only partially literate. Less than 40% of Americans have an undergraduate degree. The kinds of positions that philosophers are likely to pursue and be competitive for are a sliver of the US job market. And while it’s easy to think that this might give you an advantage in applying for positions outside this sliver, you are more likely to be labeled “overqualified” or “an academic egghead”. What’s true at the general level may not hold for this particular market.

These Trends in Relation to the Philosopher’s Skillset

Worryingly, there is reason to suspect that the particular skillset of the philosopher renders us especially vulnerable to these trends.

  • The jobs that you are competitive for with only a degree in philosophy are mostly not insulated by regulation from any of these trends. There is no regulatory framework preventing jobs in, e.g., project management, grant writing, or academic administration from being either automated or sent overseas. By contrast, some fields require specific US-based certifications, such as in law, pharmacy, financial advising, and actuarial work.
  • The primary work outputs of the philosopher are fully digitizable and tailor-made for LLMs. These are verbal and written natural language outputs. These outputs can be transmitted remotely, and the training corpus for LLMs primarily consists of written natural language inputs, rendering them especially competent with this facility. Contrast cases: engineers, lab technicians, healthcare workers, architects, etc.
  • Training in philosophy does not equip you with a deep body of specific or geographically-restricted factual knowledge. AI systems currently have problems reliably recalling or retrieving particular facts: they “hallucinate”. And non-US workers are less likely to learn deep bodies of geographically-restricted knowledge, such as detailed knowledge about the contours and history of specific energy markets in the US. A mastery of these sorts of bodies of knowledge may give you an advantage relative to an AI system with a limited context window or an offshore worker, but this advantage is one training in philosophy does not afford.
  • Training in philosophy does not equip you with expertise in more difficult or obscure software programs. Most philosophers are certainly capable of quickly learning some fairly simple software systems that are commonly used in jobs within the relevant market, such as a content management system or a client management system. But even for an intelligent philosopher, it would take much more time to learn, and prove to employers that you have learned, more difficult or obscure systems like computer-aided design software. The issue is that AI systems will much more quickly be integrated with the former kinds of software programs for a number of reasons, including the greater market size for AI agents capable of interacting with them, the greater availability of training materials, and the more robust development ecosystem around them. And the larger the technical moat, the smaller the supply of offshore labor with the relevant skills. Those with proven skills in these programs will be less susceptible to replacement by both.

Sadly, it seems to me that someone coming onto the non-ac market with nothing but a BA or PhD in philosophy is among those most vulnerable to the deadly trio of offshoring, automation, and gigification.

Your Options and What Not to Do

So what’s an enterprising philosopher to do? Let’s assume you are sticking to white-collar work and staying within the US. Here are your options as I see them.

Play the Lottery

Getting a decent non-ac job as a philosopher using the traditional playbook (e.g. marketing one’s experience writing a dissertation as “project management”, or attending a short coding bootcamp) is still certainly not impossible, and there are many cases of philosophers entering this market for the first time even post-COVID and succeeding. The good news is that compared to the academic market, the size of the non-ac market is such that you’ve got practically unlimited chances, and the only costs are your time, energy, and sanity. If you opt for this route, I strongly recommend investing heavily in networking, or ‘nepotism’ as it is spelled in some dialects, which is undoubtedly the highest value currency in any job market.

Stay Academic

Every philosopher has been warned up and down about the horrors of the academic job market, and mostly rightfully so. In the nearterm, it is only getting worse due to the debacle with federal funding. But there are at least two good reasons to think twice about abandoning the academic market if the situation on the non-ac market is as I have described it.

First, you may be overestimating your prospects of getting a position on the non-ac market that actually satisfies the reasons you don’t want to go on the academic market. In my case, for example, two of the biggest reasons were not being maximally geographically flexible and disliking teaching. However, the difficulties I’ve experienced have made me far less confident that I will be able to stick to my geographic preferences and get work that aligns better with my preferences in an increasingly tight US non-ac job market.

Second, academic philosophy is actually somewhat insulated from the market forces I have described. My hypothesis is that the primary economic value of the academic philosopher in the future will be as something like a social set piece. Philosophical research, to be clear, is strongly under the gun for automation[13] and offshoring, but this research has little economic value except to the publishing companies, who do not pay philosophers’ salaries. In other words, no academic administrator would react to the news that Philosophical Studies is filled with LLM-generated papers and that AI is now better than humans at coming up with counterexamples to the KK principle by deciding “well, I guess we can now layoff the philosophy department”. Philosophers will largely be left to their own devices to figure out how to deal with the automation of philosophical research. The real economic value of philosophers for universities, of course, lies in students wanting to take philosophy courses. And even though there is early evidence that AI systems are already remarkably effective as teachers, I don’t believe that this will greatly reduce the demand for philosophy courses run by humans. Most students are not going to want to sit at home on a computer with an AI-based education system overseen remotely by a few lucky philosophers in Bangladesh. They will want to come onto a campus and sit in a stately looking room with other students where an eccentric person who rambles on about justice makes them feel like they are in Dead Poets Society. At least, that’s my guess.

Look for Havens

The broad factors I have identified will not affect all white-collar positions equally. Jobs in the national defense industry, for example, will certainly not be offshored (though they might be automated). Unfortunately, I suspect many of these havens will require qualifications beyond a PhD or BA in philosophy to be competitive for, and if they do not, they may be overrun by people from other sectors anyway. Law, for example, seems to be a promising possible haven: the requirement of a US law license for many legal actions makes automation, offshoring, and gigification less likely, and the legal industry still has a strong culture of recruiting true entry level positions from law schools. However, if this prediction is incorrect, you may find yourself tens or hundreds of thousands of dollars in debt and a few years older but back in the same situation.

In terms of what to do positively speaking, then, I have little concrete advice other than to take the non-ac market very seriously if it is at all on your horizons. Try to do internships, freelance work, or other non-ac work during the summers, even if this cuts into your research or dissertation time. Aggressively and shamelessly network.

Things to Avoid

What not to do is a little easier to say. Do not:

  • Reason that because the academic job market is so difficult, the non-ac market must be easier and better.
  • Take for granted that because philosophers have succeeded on the non-ac market in pre-COVID times, that they will continue to do so.
  • Reason that because the general unemployment rate is low, that therefore you will get a decent job.
  • Reason that you will get a decent non-ac job because you are smart, competent, or have academic credentials.
  • Assume that you will avoid years of adjunct-like conditions (unstable employment, low wages, and poor benefits) by going non-ac.

Good luck out there.


[1] If you want to scare yourself with an endless supply of these anecdotal reports, I recommend browsing reddit.com/r/recruitinghell, while bearing in mind that there is a strong selection effect with those that post experiences there.

[2] Predictably but infuriatingly, top Meta executives responsible for this layoff nearly simultaneously increased their own salary bonuses by millions of dollars.

[3] Broecke 2024 has a good discussion of the history of offshoring and its potential directions post-COVID.

[4] Baranes 2025 appears to contain a discussion of the issue, but I haven’t been able to access the text.

[5] A relevant piece of counterevidence is the widespread partial RTO (return-to-office) mandates at US companies, generally at 3 days per week. My expectation, however, is that these are due to factors that will not persist. Companies may reason that they don’t plan to take the radical step of immediately cancelling all office leases and selling off their physical infrastructure, and so “we may as well as get them in instead of having the buildings sit empty”. Decision-makers may also prefer to have people RTO for non-monetary reasons, such as preferring the environment of a teeming corporate office. Such reasons may not hold sway for long. Overall however, I recognize that current evidence on the effectiveness of in-person vs remote work for the relevant positions appears mixed.

[6] The highest-quality recent survey of the relevant experts that I am aware of was done at the end of 2023. It found an aggregate prediction of a 50% chance of such AI systems existing by 2047. While this might seem like weak evidence for my claim, bear in mind the following: (a) the trend of these predictions is towards shorter timelines: in the same survey in 2022, the 50% figure was at 2060; and (b) in my view, incredible AI progress has been made since late 2023 and an updated survey would reflect this.

[7] There are some outstanding concerns about this particular result. For a good discussion, see here.

[8] For a recent well-informed counter-perspective on this topic, see here. For the record, I share the intuition of some AI researchers that current AI systems appear “brittle” in a way that indicates a lack of some fundamental component of human intelligence. However, my expectation is that overcoming this will not require some fundamental change in current AI architectures, that there are currently research directions in AI that seem promising for this issue, and that the vast sums of energy, attention, and money being thrown at AI development will be sufficient to fix the issue in a timely manner.

[9] See ch. 2 and ch. 8 in Countouris et al 2023 for discussion.

[10] Woodcock and Graham 2020 (ch. 1) give at least three factors contributing to the rise of alternative work arrangements that are strengthened by this phenomena: platform infrastructure, mass connectivity and cheap technology, and digital legibility of work.

[11] This reasoning doesn’t explain why contract positions have increased or will continue to increase. I am less confident in that prediction.

[12] There are also other oddities with this measure, such as not including people who’ve given up trying to find work within the number of unemployed. To be fair, this population seems to be fairly small.

[13] Any philosopher who believes that there is something distinct about philosophical inquiry that makes it uniquely immune to AI systems, such that we might live in a world where adjacent areas like math and legal research are automated but philosophy remains a stubborn holdout is, I believe, deluding themselves.


Related: Non-Academic Hires

Central European University Philosophy

Subscribe
Notify of
guest

60 Comments
Oldest
Newest Most Voted
Daniel Greco
Daniel Greco
1 year ago

This is a convincing, sobering piece. To add to the reasons for pessimism, I’d suggest that even in industries protected by regulation like law, it may still be that there will be fewer entry level roles, as work that in the past would have been done by a partner managing a team of junior associates can now be done by a partner working with AI. Some initial reports that suggest law firms may be moving in that direction:

https://www.abajournal.com/web/article/law-firms-reduced-the-pace-of-associate-hiring-shifting-to-new-talent-model-report-says

https://www.artificiallawyer.com/2025/01/09/will-law-firms-need-fewer-junior-associates/

I really don’t know what advice I’d give philosophers looking for alt-ac work right now. (Not that I would expect to be in a position to give good advice–I’ve never been on the alt-ac market.) I am thinking about what sort of labor market my children will face. My oldest child is really into chemistry, and I suppose I’m hoping that hands on science and engineering will be relatively insulated from the trends the OP discusses; my guess is that we’ll still need to run experiments, those experiments will still need to occur in places with high tech lab equipment, and they’ll still require human oversight, troubleshooting, and adaptation to unexpected outcomes that would be hard to automate. Even being able to recognize and describe unexpected experimental outcomes–in a way that you could then ask an AI for advice in handling them–will, I expect, require a lot of training. But even if all that’s right, it’s not particularly helpful for philosophers.

Eric Wilkinson
Eric Wilkinson
1 year ago
Reply to  Daniel Greco

I want to push back against this, because I know working lawyers. According to them, AI is functionally useless. It reliably produces bad and inaccurate law advice, which is espeically dangerous when the stakes are high.

The growing use of AI has actually created more work for junior partners, since other, lazier lawyers or people “representing themselves” submit documents written by AI that are incoherent or obviously wrong, so someone has to waste painstaking hours submitting responses that point out obvious mistakes. My brother, himself a lawyer, illustrated this to me by asking Google AI an incredibly simple legal question, only for it to give the wrong answer (it said every employer in Ontario was entitled to the same amount of sick days as federal workers).

I’m more familiar with law, but the closer I look at each industry where people claim AI is taking over I discover the same thing. Bosses want to replace their workforce with AI, but the AI is useless and the promises outstrip what it can do. As philosophers we should be less credulous about all this tech-hype (Student Philosopher’s comment above is helpful).

Daniel Greco
Daniel Greco
1 year ago
Reply to  Eric Wilkinson

I’ll push back against the pushback. Here’s Adam Unikowsky (I know him from college), a pretty high power supreme court lawyer. If your brother doesn’t know how to use Google AI to generate good legal advice, that may just mean your brother hasn’t spent time optimizing the right AI engine to do it. Unikowsky did have to train an AI on a ton of supreme court briefs to get it to produce the results it did, but haven’t done so, its output was extremely impressive, or so he claims:

https://adamunikowsky.substack.com/p/in-ai-we-trust

https://adamunikowsky.substack.com/p/in-ai-we-trust-part-ii

Noah
Noah
1 year ago
Reply to  Daniel Greco

One reason for the discrepancy may be due to the different implementations of AI use being described here. I would caution against judging AI capabilities from Google’s “AI overviews” feature. Since there are a ridiculous number of Google search queries being performed at any time, Google is probably using a cheaper model with lower inference costs to serve this feature. Moreover, the results are generated instantaneously — the model isn’t given time to think, unlike with the newest AI systems that take more time to think with more complicated queries.

I’m guessing that Eric would get better results asking the question to a frontier AI system like Claude 3.7. Even better if Eric were to upload some relevant documents into the context window, which would more closely mirror the implementation of these systems in professional legal contexts. This piece by Unikowsky describes using Claude 3 Opus and uploading relevant PDFs into the context window, which explains the better results obtained.

Let me add that this incident illustrates a broader reality which actually makes law a more promising haven than some others. My outsider perspective is that law is less technologically mature than many other fields, for example there’s still a shocking amount of reliance on physical paper (although my perception here could be outdated). I anticipate that there will be more resistance and more friction in the adoption of AI and other technological solutions that allow offshoring and gigification in the field of law than others.

Daniel Greco
Daniel Greco
1 year ago
Reply to  Noah

Yes. A good way to think of it, I think, is as follows. Is the investment a legal partnership needs to make in identifying and training the best AI model to do junior associate work more or less than the investment they need to make in hiring a junior associate? Hiring a junior associate is a *huge* investment. So even if training an AI model to do junior associate work is non-trivial–you can’t just fire up ChatGPT and start asking it contextless legal questions, but instead need to pay for a top of the line model, and then spend some time fine-tuning it for your specific legal needs–there’s still a lot of scope for it being cheaper than hiring human beings who draw salaries and benefits year after year.

Eric Wilkinson
Eric Wilkinson
1 year ago
Reply to  Daniel Greco

If the contention is that a Supreme Court lawyer can eventually coach the AI to do what he himself could do in less time, I concede.

I’m being a bit glib, but this is not the scary time-saving automation we are often threatened with.

Daniel Greco
Daniel Greco
1 year ago
Reply to  Eric Wilkinson

I don’t think this is minor. It’s not like this is a rare skill that only supreme court lawyers have. If a supreme court lawyer can do it, I would bet tons of senior partners can do it too, maybe in a few years when the technology becomes more widespread (Unikowsky has a CS background, so I’d bet he’s a bit ahead of the curve on this.)

DoubleA
DoubleA
1 year ago
Reply to  Daniel Greco

No offense to Unikowsky, but in my humble opinion he’s got a very optimistic opinion of Claude’s outputs. I’m not saying plenty of associates don’t produce documents like some of those outputs, fwiw. But I don’t think Claude is an “insane genius,” for example. It’s also the case that Supreme Court decisions are a special legal arena in many ways and performance in that arena may not generalize. AI may do better or worse in more commonplace environments, but I would caution against drawing broad conclusions from that domain.

A lot of the pablum stuff about the relevant standards for most actions are things that experienced attorneys already have saved and macro’ed for writing, fwiw. AI does have the capacity to greatly alter the legal profession, but I think LLM’s get a disproportionate amount of attention. AI can greatly reduce research costs and generate decent predictions, but a lot of the work in these models is not LLM based, it’s “old fashioned AI.”

For anyone interested in this beyond blogs and substacks, I recommend Kevin Ashely’s book, “Artificial Intelligence and Legal Analytics.” It’s seven years old, but much of the cutting edge now is in integrating LLMs with many of the models discussed therein.

ConcernedGrad
ConcernedGrad
1 year ago

I’m surprised the FrontierMath results are being used here as evidence, especially given the linked discussion.

E D
E D
1 year ago

Why didn’t you pursue an academic job, OP? Per APDA, USC places students into TT jobs at an absurdly high rate of ~70% plus.

HFDJHFD:LA
HFDJHFD:LA
1 year ago
Reply to  E D

They mentioned they didn’t like teaching and didn’t want geographical constraints.

hhk
hhk
1 year ago

A few thoughts: in going to the non-academic job market, I think two things are absolutely key to securing a job (because the job market is difficult).

First, knowing and narrowing the sort of job that you want and researching to make sure it is realistic with the current market. In this job market, casting a wide net results in no fish. When you narrow down, you’re able to build the skills you need for that particular job.

Second, networking is absolutely essential, as mentioned in the article. When you have a narrow job that you are looking at, networking becomes a lot easier. You can talk to people in the field and learn how to be a competitive candidate.

writtenbyanLLM
writtenbyanLLM
1 year ago

In other words, no academic administrator would react to the news that Philosophical Studies is filled with LLM-generated papers and that AI is now better than humans at coming up with counterexamples to the KK principle by deciding “well, I guess we can now layoff the philosophy department”. Philosophers will largely be left to their own devices to figure out how to deal with the automation of philosophical research. The real economic value of philosophers for universities, of course, lies in students wanting to take philosophy courses.

I don’t agree that philosophical research has no economic value to universities. After all, most research-intensive universities make tenure and hiring decisions primarily on research capability. The idea here is that research is indirectly economically valuable to universities: better research causes higher rankings, which causes more grants and more students (things which are directly economically valuable to universities).

Assuming that in the near future, LLMs are able to do philosophical research better than philosophers, this suggests that research by philosophers will no longer be as economically valuable to universities, which may very well result in mass layoffs. Optimistically, LLMs may research better than philosophers but not teach as well (as OP suggested), and in this case, philosophers may keep their jobs but there will be a drastic shift in hiring/promotion criteria from research to teaching. In any case, I don’t see at all how philosophers will “largely be left to their own devices”.

Noah Gordon
Noah Gordon
1 year ago
Reply to  writtenbyanLLM

I really doubt that the automation of philosophical research capability will make much of a difference here. The measures that administrators use to gauge this are so far off the ground of what philosophers use that the stats are basically meaningless anyway, and administrators will be free to cherry pick or otherwise make up claims to make it seem like they have top philosophy departments anyway. The main thing that matters for student recruitment is the large overall undergrad rankings, like the US news and world report college rankings, and the effect of the automation of philosophical research on these will be entirely marginal.

(On another note, I didn’t say that LLMs won’t teach as well, in fact I think they probably will. I think the value of colleges will largely become for the social experience, and the economic value of the philosopher will be determined in relation to that.)

Student Philosopher
Student Philosopher
1 year ago

As someone in tech, working in AI, with at least a third row seat to what’s going on, I can vouch for some of this but want to caution that this statement is unproven:
“ AI and Automation: The tech businesses that are downsizing are the very same ones currently building frontier AI systems explicitly designed to automate a massive amount of white-collar work. This could just be a coincidence, but I suspect that the front-row seat decision makers within these businesses have to the progress of AI development has given them a sense that they will soon, if not already, be able to do more with fewer human workers.”

In my experience, this narrative is more hype than reality. As someone working inside these companies, relatively little was being automated with LLMs. Some work was being enhanced, but no one’s work was replaced. Many efforts, even to automate some administrative work, failed.

It’s important to understand that business leaders stand to profit enormously in the short term from hyping AI, convincing others to buy it, and laying off workers to create the appearance of improved profit and increase the share price of their companies.

Many experts and practitioners in the field are skeptical that LLMs are capable of doing professional work to the same level as a human. Those sounding the alarm are alerting to the dangers of replacing an intelligent human with an unproven technology. Read work by Geoffrey Hinton, Gary Marcus, Timnit Gebru and even Meta’s Yann LeCun.

I urge that we do not fuel the propaganda of business leaders that stand to profit from outsized promises. Instead, we should demand that they prove the technology works before we take their word for it.

Noah Gordon
Noah Gordon
1 year ago

It’s difficult to separate the hype from the reality here, and my view is that there’s a mix of both and it varies greatly by specific professions and tasks.

More specifically, I believe that a lot of the hype about the current ability of AI systems is justified for some tasks involving writing code and natural language. The current systems appear to be pretty good at this and getting better rapidly. I could cite some more evidence but there’s a decent amount in the piece. They work less well when you have to integrate them with other systems, including for simple administrative tasks like filling out forms or sending emails. I would imagine that almost no work in accounting, for example, can currently be automated well both due to the systems not working well with, e.g., Excel spreadsheets, and due to the very high costs of a hallucination like making up a figure in a certain cell.

Of course, as you agree, the near-term effects on the job market are more about the hype than the reality. In the long-term, the reality will matter, and my bet is that those shortcomings will continue to be overcome.

Eric Steinhart
1 year ago

This is a great piece. But I disagree when you say “academic philosophy is actually somewhat insulated from the market forces I have described”.

AIs can be used for grading. Universities are building online asynchronous programs. The courses in such programs are canned, and, after construction, the only role of the faculty is grading. If AI can do that, there’s no need for faculty at all.

Pretty much everything you’re saying transfers from the PhD level to the BA, meaning that it’s going to be very hard for philosophy undergrads to get jobs. There will be fewer and fewer philosophy undergraduate majors. Departments will close.

Based on anecdotal evidence from AI articles here on Daily Nous, philosophers are strongly opposed to using AI. (But there may be a strong selection bias at work.) If the entire economy is screaming: Use AI! And philosophy professors are banning it, students will have one more reason to avoid philosophy. Departments will close.

You write that students want to “come onto a campus and sit in a stately looking room with other students where an eccentric person who rambles on about justice makes them feel like they are in Dead Poets Society.” Very, very few students do that. Most come to a shabby room on a run-down campus where the instructor is poorly paid and not very qualified or interested. There is very little social incentive.

Noah Gordon
Noah Gordon
1 year ago
Reply to  Eric Steinhart

You’re right, I was probably too glib and overgeneralizing from my own academic background in that paragraph.

Thinking a little more carefully, if we’re being honest there are two market draws for colleges right now for most students: credentials and the social experience. Perhaps education itself still has some modicum of sway, but I suspect it’s quite small.

I’d have to think a bit more about how these 3 factors interact with credentialism; I’m not confident making any predictions there as of yet.

I do think that the higher education market will need to lean in even more on the social experience to survive. Unfortunately, a lot of places that don’t have the foundation for that might not survive, but perhaps there will be growth from other organizations to compensate. It’s difficult to say.

Amod Lele
Amod Lele
1 year ago

One thing I don’t think this article doesn’t mention is the alt-ac (as opposed to non-ac) job market: positions in teaching centres, grant seeking, student life, and all those other parts of the university where hiring over the past couple decades has greatly expanded while faculty jobs continue to decline. Those areas tend to focus on people skills, which are harder for AI to automate, and they’re ones where a PhD is a much more meaningful credential than in the non-ac world (because it means you can relate to facutly and know more than the typical undergrad about how universities work). My impression is that US alt-ac is not in a great position right at the moment due to the mammoth cuts to federal agencies, but I suspect it will recover as the dust clears. That isn’t to say alt-ac is for everyone, but I think it is one of the more promising exit strategies for a lot of PhD holders.

Emily Mathias
Emily Mathias
1 year ago
Reply to  Amod Lele

This is what I was going to mention along with a stronger emphasis on the need to develop networking skills. Student Affairs jobs abound in the US. The downside is you may have to settle for a lower paying job – but it’s a job with upward mobility and appreciation for your degree and years of experience. It’s also a location for alternative funding for graduate students as well!

Nicolas Delon
Nicolas Delon
1 year ago

Thank you for this excellent if sobering post for which I hope you receive the compensation you deserve.

TAG
TAG
1 year ago

It appears that we have reached the arresting conclusion that sooner or later (but probably sooner than later) capitalism turns us into meat (or batteries). In light of this it is a bit strange that we still go on discussing the best survival strategies…

William Large
William Large
1 year ago

I think there are good philosophical reasons not to fall for the hype of AI.

Kenny Easwaran
1 year ago
Reply to  William Large

There are also good philosophical reasons not to fall for the hype of anti-AI!

We should all be paying attention to significant new developments in the processing of information and production of things like coherent text, images, and computer code – particularly when these developments are of the sort that just a decade or so ago most people thought were several decades away. We don’t have to believe these systems are producing something that makes sense to call “human-level intelligence” to think that these systems will radically change the lives of people who spend a lot of time producing text, images, or computer code (whether as part of one’s job or not).

The default assumption should never be that the next 20 years will be just the same as the last 20 years, but when you see these sorts of developments, we should be opening our minds to many more possibilities, and preparing for more of them.

Hermias
Hermias
1 year ago

It seems like the factors all point toward increased economic value of having a physical body. It would be pretty tragic-comic if the “jocks” ended up being the masters of the “nerds”, in adulthood as in high school.

There is something appropriate to the philosopher in agricultural pursuits – self-sufficiency, silence, careful attention to reality. Plato himself was jacked, so I have heard.

The last few months I have been using ChatGPT and grok. I find them very useful for googling – “find where Plotinus said such and such”, “list some philosophers who discuss X”. I’ve used them both to “generate 10 objections to virtue ethics” for class prep, but the results seemed very formulaic and not intelligent. I’ve never had an interaction that made me say “wow”. Can anybody suggest better ways to interact with it/the best one to use?

Richard Y Chappell
1 year ago
Reply to  Hermias

I think Claude is currently the best at philosophy. You don’t want to rely too much on its “background knowledge” – that will just regurgitate the average vibes found on the internet. You need up copy & paste or upload specific documents into the context window, and ask it to offer careful and in-depth thoughts. Perhaps start by asking for (i) a summary of the key ideas/arguments of the article, (ii) what it thinks the most revealing or insightful objections to the article might be, and (iii) how the author might best respond to such-and-such thoughts that you have.

fwiw, I’ve definitely been wowed by this experience!

Ben
Ben
1 year ago

I’ve been wowed by this experience too, but even so I still haven’t really found a way to make LLMs useful in my own research. Have you?

Richard Y Chappell
1 year ago
Reply to  Ben

Admittedly not! (Though I’m hopeful that they may help me to better anticipate referee objections in future…)

Kenny Easwaran
1 year ago
Reply to  Ben

Weirdly, two years ago, right around the time that Bing added ChatGPT, and tried to seduce Kevin Roose away from his wife using emojis, I asked it a question about the epistemological significance of sigma-algebras, and it gave me a link to a paper in an economics journal that had some ideas in it that were exactly what I needed to get un-stuck in some thinking. That’s mostly luck, and in the time since then the only research-related benefits I’ve had from AI systems are that it’s easier to make illustrations of thought experiments when preparing for a talk, and that it has helped confirm which sections of paper might be able to be shortened to meet word limits.

But it’s been extremely helpful for making new questions for homework assignments. If I give Claude a list of a few questions I’ve already included, or even just the idea for a type of question, it can brainstorm a bunch of examples faster than me. Even though only a third or so of those questions are usable (sometimes with a bit of editing), it’s still much faster than my brainstorming, and I can usually get a few better ones than any I would have come up with – and it’s especially useful if some students want extra practice questions. And it’s particularly nice that the logic questions come pre-formatted in LaTeX, so I don’t have to spend half an hour remembering all the syntax for coding a truth table.

Richard Kim
Richard Kim
1 year ago

Thank you for this clear, balanced, and informative take. I am banking on this particular (and delightfully worded) comment holding true:

Most students are not going to want to sit at home on a computer with an AI-based education system overseen remotely by a few lucky philosophers in Bangladesh. They will want to come onto a campus and sit in a stately looking room with other students where an eccentric person who rambles on about justice makes them feel like they are in Dead Poets Society. At least, that’s my guess.

I doubt though that students watch the Dead Poets Society any more. Maybe we should start including that movie in our classes.

Louis F. Cooper
1 year ago
Reply to  Richard Kim

I saw Dead Poets Society in a theater when it came out years ago but didn’t remember it well. As it happens, I re-watched it fairly recently on computer (YouTube rental).

The character played by Robin Williams is a very good teacher who doesn’t “ramble on” about anything. So the OP is mischaracterizing the film. (Also, the setting is a prep school, not a college or university, but whatever.)

The point is that the OP could have made this particular point about in-person vs. remote attendance without mischaracterizing a movie that the author of the OP likely remembers only vaguely.

Caleb
Caleb
1 year ago

This piece makes a lot of good points especially about the non-ac job market being potentially worse than the academic job market right now. However, I think the AI hype is just that, hype, which companies are cynically using to justify reckless and irresponsible profit-maximizing behavior. There’s no reason to believe that LLMs will continue to improve indefinitely; in fact, improvements in LLM performance are already appearing to plateau. Moreover, there are still many tasks that LLMs are worse at than the average human child and it seems doubtful that throwing more data at the problem will lead to meaningful improvements. I think a long AI winter is much more likely in the next decades than AGI.

Noah Gordon
Noah Gordon
1 year ago
Reply to  Caleb

I used to believe this exact same narrative: “it’s all industry hype driven by AI business people to sell useless products”. I no longer do. There is real evidence out there; I referenced some of it. I think it’s worth noting for anyone keeping score here that the AI skepticism in the comments here has largely been given without evidence, or very weak anecdotal evidence (notable exception in the recent comment by Rob Hughes, which I will respond to shortly).

Philosophers, and especially early career vulnerable grad students who may go on the non-ac market, will be done a great disservice if they uncritically accept this narrative.

Caleb
Caleb
1 year ago
Reply to  Noah Gordon

No offense, but I think most of the evidence offered in your post is also fairly weak and anecdotal, which is to be expected: it’s hard to have conclusive evidence about the future.

However, I think it’s clear that the current abilities of AI are massively overblown. I work in a cogsci department where many people are doing projects involving LLMs. I’ve sat through so many talks in the last 2-3 years about things young children are much better at than our best AI models, everything from intuitive physics to counting. Maybe we’ll see massive improvements in AI in the next few years. But, given that Silicon Valley (and corporate America more generally) has a financial incentive to mislead the public, I think we should be very, very skeptical.

Noah Gordon
Noah Gordon
1 year ago
Reply to  Caleb

Do you have anything to say about any of the following pieces of evidence I offered:

  • The surveys on AI progress expectations from AI researchers
  • The performance of AI systems on various benchmarks
  • The performance of AI systems on various real-world tasks, including medical diagnosis, coding, and persuasion

Meanwhile, what is the evidence you’re offering on the other side? It’s hard to tell from your description. Here are a few vague things to keep in mind when assessing this vague evidence. Current AI systems are trained primarily on text and fine-tuned for doing text-based tasks like writing English and Python. If competence with intuitive physics requires getting data like sensorimotor inputs, it makes perfect sense that AI systems will remain poor at this until they are trained on more data that can approximate the sensorimotor data humans get from walking around and throwing things and so on. Moreover, AI intelligence might be more jagged than human intelligence. It wouldn’t follow that it won’t be incredibly useful in many fields and for many tasks.

Please keep in mind that you can do real harm by misleading people by asserting strong opinions publicly without doing your homework.

Student Philosopher
Student Philosopher
1 year ago
Reply to  Noah Gordon

A piece of counter-evidence from each of your own sources:

The AI survey: “Forecasting is difficult in general, and subject-matter experts have been observed to perform poorly [Tetlock, 2005, Savage et al., 2021]. Our participants’ expertise is in AI, and they do not, to our knowledge, have any unusual skill at forecasting in general“This is not our first AI hype cycle.These benchmarks were developed by the companies doing the research. The comparison of humans to AI on these benchmarks is not really apples to apples.AI Coding”Results indicate that the real-world freelance work in our benchmark remains challenging for frontier language models. The best performing model, Claude 3.5 Sonnet, earns $208,050 on the SWE-Lancer Diamond set and resolves 26.2% of IC SWE issues; however, the majority of its solutions are incorrect, and reliability is needed for trustworthy deployment.“Facts, happening right now, in industry:

A large number of layoffs resulting in employees expected to work overtime compared to baseline.Massive offshoring to lower cost regions around the world.No evidence that LLMs are doing anyone’s job from end-to-end.Evidence that employees are being forced to work with AI output that is slowing them down, rather than speeding them up (e.g. unqualified contractors using AI and not validating it’s output for correctness).Feedback from industry that the costs of implementing AI solutions are far outweighing the benefits (OpenAI can’t even profitably sell it’s current product line).Weakening employee protections, political unrest and economic downturn creating a negative impact on the job market (e.g. AI is not taking the jobs of federal workers, they’re just being laid off for political reasons).You clearly did your homework and I relate to much of what you have written. But we have very good reason to push back against the idea that AI advancement (rather than expectations of AI advancement) is the cause of the current job market. That is what I perceive most people are reacting to here.

Noah Gordon
Noah Gordon
1 year ago

It was definitely a rhetorical mistake to use so much space on this contentious issue. I have a bad habit of, when a point is relevant in a piece of writing, being unable to reign myself in from going into it fully. Predictably, the comments have been dominated by the hot-button, sexy issue of AI skepticism rather than boring topics of offshoring and the gig economy.

If I had to rank the factors I discussed in terms of actual current impact on the job market I would say, purely on vibes:

  1. Gigification and contractification
  2. Offshoring
  3. AI (with expectations having a stronger effect than reality)

I say in the piece itself that the expectations have a stronger near term effect than the reality:

“the effect of AI on the job market in the near-term depends more on employer perception of AI progress than the reality of it”

Alas, the piece is too long. Anyway, here are some thoughts on the specific points:

  • “Evidence that employees are being forced to work with AI output that is slowing them down, rather than speeding them up” — what exactly is the evidence of this?
  • On forecasting: yes, forecasting is in general not that reliable. My forecast as well! On the other hand, we need to forecast to plan and live our lives, and the experts are a relevant source of at least some evidence. Most of us take seriously when climate scientists make predictions about the future climate, and when geopolitical experts comment on what might happen to Taiwan if we enact restrictive export controls on China and so on. No one has a crystal ball, but the experts, one would hope, have a slightly less cloudy one.
  • On the SWE Lancer paper: Yes, Claude Sonnet 3.5 was able to only 26.2% of the IC SWE Diamond dataset. And it was able to solve only 36.1% of the full dataset (SWE Lancer Diamond), which includes coding management tasks rather than just pure coding tasks. This means that the majority of its solutions were incorrect. But those figures are still pretty damn good, I’m sure it’s better than I would do unassisted, worse than a FAANG engineer would do (maybe 90%? just guessing), possibly on par with a bad engineer. On this estimate, almost a quarter of real-life programming tasks on Upwork are already, right now, automatable with Claude 3.5! That number should go up with Claude 3.7. If this is a good estimate, we would be crazy to think that current AI systems won’t have large effects on the job market for programmers.
Tedd Siegel
Tedd Siegel
1 year ago
Reply to  Caleb

Scorchingly realistic, while unrelentingly cheerful. I suspect it is AI generated.

Kenny Easwaran
1 year ago
Reply to  Caleb

It’s also true that for many professors I know, and many highly paid professionals, there are tasks that they are worse at than the average human child.

That’s probably a bit glib, but a system doesn’t have to be good at every single task that children are good at in order to be able to radically re-shape the job market. All that matters is that there are *some* tasks that the system *is* better at or faster at than most human professionals who do that task.

Even if all progress on AI systems completely freezes, so that we are stuck with Claude 3.7 and GPT 4.5 and o3 forever, they will likely continue to reshape job markets (and many other aspects of life) for many years, as all of us learn what they are effective for and what they aren’t.

Simon Goldstein
Simon Goldstein
1 year ago
Reply to  Caleb

‘improvements in LLM performance are already appearing to plateau’

It is important to distinguish improvements from *pretraining* from all improvements. It is possible that simply scaling up compute / data will no longer produce dramatic improvements. Even here, though, we have to wait and see what happens when training compute is 10x-ed over the next year, as new data centers come on line.

But separate from this point, there have been startling results in the last 3 months from new ‘post-training’ paradigms, in particular reinforcement learning in verifiable domains like coding and math. New reasoning models like o3 are not plateauing in *any way whatsoever*. o3 is the 175th best competitive computer programmer in the world. Math benchmarks are falling left and right. Moreover, the new ‘reasoning model’ paradigm has just begun, and will quickly extend from easily verifiable domains like math to AI-verifiable domains, which are far more general.

Ben
Ben
1 year ago

I’ve found it hard to square o3’s impressive performance on “benchmarks” with my personal experience. Have you tried using it for coding or math? Yes, it’s capable of doing cool stuff, but I’ve also found that its code is often annoyingly buggy. It often makes mistakes in its mathematical proofs too… and it gets hand-wavy and vague in crucial steps of the reasoning. It also tends to massively over-complicate things.

Sad
Sad
1 year ago

Another day, another reason to get depressed.

Rob Hughes
Rob Hughes
1 year ago

Gordon’s piece makes some interesting and thoughtful observations about the non-academic job market. I have to push back on the parts of the piece that are specifically about “AI.”

1) Gordon writes that optimistic predictions about the capabilities of LLMs are coming from “sober-minded people” who have “have only weak incentives to exaggerate.”

Anyone who has invested in diversified US stock funds has an incentive to support the narrative that “AI” is a good investment. As of March 13, 2025, 30% of the money in an SP500 index fund or ETF is invested in the Magnificent Seven (Alphabet, Apple, Amazon, Meta, Microsoft, NVIDIA, and Tesla). 27% of the money invested in a US total market fund or ETF (such as ITOT) is in the Magnificent Seven. All the companies in the Magnificent Seven have invested heavily in “AI.” Many other companies have invested in “AI” as well. If LLMs turn out to be over-hyped, and the bubble bursts, that spells bad news for a lot of people’s portfolios.

2) As evidence of LLMs’ growing capabilities, Gordon cites news from December 2024 about the performance of OpenAI’s o3 model on the Frontier Math benchmark. Gordon acknowledges that “there are some outstanding concerns about this particular result.” He links to a post on LessWrong about those concerns. I think it is worth highlighting a key point in that post.

EpochAI, the organization that organized the creation of the Frontier Math benchmark, was “completely funded by OpenAI and shared with them exclusive access to most of the hardest problems with solutions.” EpochAI disclosed this fact on December 20, 2024, the same day OpenAI announced the o3 model and its performance on the benchmark.

3) The optimistic narrative about the future prospects of LLMs is up against a recent finding in computational complexity theory by Iris van Rooij et al., published in Computational Brain & Behavior. Key result (fron the abstract): “As we formally prove herein, creating systems with human(-like or -level) cognition is intrinsically computationally intractable. This means that any factual AI systems created in the short-run are at best decoys.”

https://link.springer.com/article/10.1007/s42113-024-00217-5

The mathematical finding predicts what we observe. LLMs produce impressive results when prompted in a way that is designed to highlight their capabilities. When tested rigorously, they “hallucinate.” If the mathematical proof is sound, we can expect LLMs to keep hallucinating for the foreseeable future. New models may have different hallucinations, and widely-discussed hallucinations may be patched, but the problem of hallucination will not go away.

I have linked to this finding on DailyNous comment threads before. I have yet to hear a serious rejoinder. I know enough complexity theory to follow the proof by van Rooij et al., but not the 2022 proof by Shuichi Hirahara on which it depends.

Noah Gordon
Noah Gordon
1 year ago
Reply to  Rob Hughes

Hi Rob, thank you for engaging substantively on these issues. Here’s my response:

(1) It is true that anyone invested in the US stock market has some financial incentive to hope that AI investments pan out, since if they don’t, that would hurt the performance of those large tech stocks. This might be a good reason to distrust the statements of leaders of these companies like Dario Amodei, who has a large enough platform to meaningfully influence the market by making these claims.

But these are not the statements I was referring to in that part of my post. I was referring to statements like this one by Hieu Pham, fawning over AI performance on the “Humanity’s Last Exam” benchmark. Pham is an AI scientist at xAI with a modest twitter following. That tweet got less than 200 likes. Pham almost surely does not make these claims in an attempt to influence stock prices to preserve his investments. I see statements like these every day in my twitter feed due to following this group of researchers. The survey I cited in the post was an aggregate of opinions of these AI researchers: the survey audience was sourced from people who have authored papers in leading AI journals. These are largely not business leaders with large audiences who can meaningfully affect the market with their opinions.

(2) Yes, OpenAI had access to a large portion of the benchmark questions. It doesn’t follow that they used these to train their model. There is a debate about whether they did, which I left to readers to follow-up on in the citation.

Admittedly, I think this piece of evidence is not that strong. I mostly used it because I have a personal connection to it (described in the piece), which I thought would be more rhetorically effective for the average reader. I could have used any number of other benchmarks, such as QPGA, Humanity’s Last Exam, or AIME. On the whole, the progress on these benchmarks is overwhelmingly strong.

(3) I’m very skeptical that these kinds of a priori results can tell us much about the empirical path of AI progress. Here’s my broad overview of the strategy of the paper, from what I can tell:

3.1 Try to give a mathematical formalization of the task of an AI system (“AI-BY-LEARNING”), which is cashed out as finding an algorithm A that, when run for different possible situations as input, outputs behaviours that are human-like by approximating some distribution. This A would be like AGI.
3.2 Try to prove that this formalization is not “computationally tractable”, which is formally cashed out as meaning that you can find another algorithm which “can be run on any situation s in polynomial time O(n^c) where n is some measure of the input size (|s|) and c is a constant” and which identifies the algorithm A.
3.3 Make heavy weather about the actual empirical significance of the result in 3.2.

Two main comments:

  • No one is claiming that we’ve found an algorithm that perfectly identifies AGI. That’s not even a good description of what the field of AI is aiming for. There are empirical efforts aimed at identifying algorithm A (i.e. AGI). Some of these are getting closer. There are various reasons to think they will keep getting closer. I don’t see how this formal result can tell us anything about how this task will go.
  • You could make the same argument by formalizing the process of evolution in a similar way. Say that the task of evolution is to find an algorithm A that when run for different possible situations as input, outputs behaviours that are human-like by approximating some distribution. Uh oh, now we can prove that finding A is computationally intractable. Guess this means evolution will never produce humans, right? Obviously not.
Rob Hughes
Rob Hughes
1 year ago
Reply to  Noah Gordon

Thanks, Noah, for your thoughtful and detailed response. A few thoughts in reply:

Re (1): I think financial incentives are affecting people commenting on “AI” both consciously and unconsciously. People with a financial interest in a question may be inclined to interpret the evidence in a way that would be favorable to their interests, even if they do not expect their statements on the matter to influence the market.

I hope you will excuse me for choosing not to engage with anything posted on the X platform.

Re (2): I cannot prove that OpenAI used the data that was available to them. OpenAI is making extraordinary claims. The burden of proof is on them. Absent positive evidence that they did not use the data available to them to help train o3, I think their claims about o3’s performance should be completely discounted.

Re (3): Developers of LLMs certainly are using algorithms to train LLMs. I take your point here to be that we don’t need the algorithm that trains an LLM to be guaranteed to run in polynomial time given any input. They could succeed with an algorithm that does not reliably run in polynomial time if the input data is short enough or if there are patterns in the input data that enable the algorithm to complete faster than it is mathematically guaranteed to complete.

As a metaphor, consider the traveling salesman problem. This is the problem of finding the shortest route for a salesman to take through a list of cities, visiting each city exactly once and then returning to the first city. This problem is “NP-hard.” Though the problem has been extensively studied since it was first posed almost a century ago, we have not found an algorithm that can reliably compute solutions in polynomial time, and there is reason to believe we will never find such an algorithm. So, as the number of cities and possible routes between cities grows, the maximum amount of time a computer would need to calculate the shortest route grows exponentially. But if all the cities are located on the coast of an island, and the coastline is a convex curve, and there is a coastal road, it is easy to calculate the shortest route: drive along the coast. All the routes through the interior of the island can be ignored. We don’t need a polynomial time algorithm or a lot of computing power to figure this out.

Developers of LLMs who are hoping to achieve artificial general intelligence (AGI) are hoping that that human behavior will turn out to have patterns in it, akin to all the cities a salesman needs to visit being on the coast of a convex island. What reason do we have to expect that human behavior or human cognition contains such simplifying patterns?

Re evolution: You seem to be assuming here that it is computationally feasible to model evolution. I would think that accurately modeling the evolution of cognition in animals would require simulating the lives of animals in natural environments. I can’t prove that this is computationally infeasible, but it sounds infeasible to me.

Noah Gordon
Noah Gordon
1 year ago
Reply to  Rob Hughes

Glad we could have this productive exchange. Some further thoughts:

(1) That’s a fair point that AI researchers may be influenced subconsciously by their financial incentives. Of course, it’s also true that they can be influenced unconsciously (and consciously) by various other non-truth-seeking motives. On the whole, I think their financial incentives for making these statements aren’t that strong (but that doesn’t mean they don’t subconsciously falsely believe that they have stronger incentives than they actually do). I think speculating about these kinds of motives questions is apt to be unreliable. On the whole, I take the apparent (possibly growing) consensus of AI researchers towards shorter AGI timelines to be decent evidence. We should be informed by the experts, although the experts are human and have various conflicting motives.

(2) I don’t think that we should completely discount the evidence from FrontierMath. I apportion some non-zero credence to the claim that OpenAI honored their verbal agreement and didn’t train on the assessment materials. Even if you think they have the burden of proof here, surely you shouldn’t give 100% credence to the claim that they cooked the books by training the model on the solutions. Nonetheless, I’m happy to admit that the evidence from this particular assessment is objectively not that strong.

(3) With a good night’s sleep and some time to reflect, things do seem to be more complicated and interesting on this point that I initially thought. Admittedly, my main familiarity with complexity theory is memorizing the run-time of some algorithms for CS 101, so I’m a bit out of my depth here. That said, further thoughts:

  • I think there are good general methodological reasons to be skeptical of the empirical implications of the result. We know that natural processes involving flesh and blood organisms found the relevant function. The argument here seems to imply that it is a priori knowable that we cannot recreate that function by running programs on metal computers. It’s possible that this is a priori knowable, but it would be very surprising to me.
  • AI scientists are using algorithms to find AGI. But that’s not all they are doing. Now, maybe there’s a level of description on which it’s all algorithms, in some sense. I’m not sure to what extent the correspondence with concrete reality is good enough for these mathematical results to tell us something empirical, if you understand my meaning.
  • The results in the paper provide the bounds for the worst-case, if I understand correctly, no? I think the empirical implications of the result would be stronger if it were an average-case or best-case result.
  • Yes, I suppose you could say that evolution is doing something non-computationally tractable. I don’t know how to assess the plausibility of that. I wonder if there’s any interesting connection with work on evolutionary algorithms.
  • You ask “What reason do we have to expect that human behavior or human cognition contains such simplifying patterns?”. I would think that the current success of machine learning approaches is good evidence. It also seems first-personally plausible to me. Also, cognitive science in general seems to depend on the truth of this.

This discussion does raise a very interesting general question:

Suppose you have a mathematical result that the computational task of identifying a function F is in the worst case intractable (worse than polynomial). How should that inform your expectations about the empirical prospects of finding F?

I’d be grateful if you know of anything written on this question.

Rob Hughes
Rob Hughes
1 year ago
Reply to  Noah Gordon

Thanks for this reply. Your last question is a good question. I don’t have a good answer. I would think that if the worst case is demonstrably not tractable, we need positive evidence to justify either the expectation or the hope that the actual case will be computationally tractable.

The second paragraph of Wikipedia’s page on evolutionary algorithms says, “In most real applications of EAs, computational complexity is a prohibiting factor… [T]his computational complexity is due to fitness function evaluation.” Imagine what would be required to calculate the evolutionary fitness of an animal with a given genotype. We would need to run a simulation of the animal’s development and of its interaction with its environment.

I know that machine learning has many real uses, and I am prepared to believe that LLMs might have use cases that will have long-term impacts on job markets. The search for AGI based on LLMs is an attempt to build a Tower of Babel. Maybe the people working on the tower can’t see the problems. Maybe from a long distance, the growing structure looks impressive. From the right vantage point, one can see that the tower has arches perpendicular to slanted ground, and that upper floors are being constructed on top of an incomplete foundation. At some point, the structure is going to come crashing down. The open questions are when it will fall and who will get hurt.

Noah Gordon
Noah Gordon
1 year ago
Reply to  Rob Hughes

I think we’re largely on the same page now. (Of course, we retain background disagreements; c’est la vie.)

One last substantive thing I will put on the record is that the Rooji et al paper seems to assume shockingly little about the AGI algorithm to obtain their results. When this class is considered fully generally, does it include various simpler things that are more plausibly algorithmically reproducible? I’m not sure yet, but I’m suspicious.

For what it’s worth, I hope you’re right about the limitations of current ML approaches to AGI. I think society needs more time, and I’m not sure these things will be benevolent to us or that we know how to control them if they are not. Personally, I won’t be happy to see the writing and analytical thinking skills I’ve worked my whole life on get automated and commoditized. Maybe I’m just a pessimist at heart, but I really do begrudgingly think that these things are near at hand.

Thanks for the discussion.

Charles Pigden
Charles Pigden
1 year ago
Reply to  Rob Hughes

Rob your metaphor is rather unfortunate. It may be that the Travelling Salesman Problem is as NP-hard as you say it is, but there are route-optimising programmes on the market that enable fleets of trucks to deliver products to stores from an original depot in a reasonably efficient manner (eg with minimal doubling back) and that can even take into account such things as speed limits and driving conditions en route. (A route might be shorter than another in terms of kilometres travelled but longer in terms of time and fuel spent, something that the business in question needs to know about.) They can even give advice on the optimal loading of delivery palettes. How do I know all this? Because my slightly younger brother Tim (we are both over 65) owns and operates a company that specialises in route-optimising software. (Those interested can check him out. He has a few academic papers on the subject.)

So if the idea that an AGI is going be practically out of reach for much the same reasons that an algorithm solving the Travelling Salesman Problem is practically out of reach, then the metaphor suggests that even this is true, near-enough not-quite-general AIs solving a wide range of practical problems, may well be possible. And that is quite enough to pose a threat to the job prospects of philosophers, as also many others.

Charles Pigden
Charles Pigden
1 year ago
Reply to  Rob Hughes

Thanks to Rob Hughes for introducing us to van Rooij, Guest, Adolfi, F. et al’s . Reclaiming AI as a Theoretical Tool for Cognitive Science.  It is a very interesting paper and I can’t pretend to have anything more than a superficial understanding of its contents. But I do have a question (though let me stress an appropriately humble one). The authors state that: 

“The thesis of computationalism implies that it is possible in principle to understand human cognition as a form of computation. However, this does not imply that it is possible in practice to computationally (re)make cognition. In this paper, we have shown that (re)making human-like or human- level minds is computationally intractable (even under highly idealised conditions). Despite the current hype surrounding ‘impending’ AGI, this practical infeasibility actually fits very well with what we observe (for example, running out of quality training data and the non-human-like performance of AI systems when tested rigorously).”

 Isn’t there a tension here? The authors agree that human cognition is, or may be, a form of computation, which presumably means that they can be modelled by a sufficiently sophisticated computer. But they show, or seem to show, that it is impossible to duplicate human-like intellectual capacities without running into problems of computational intractability. But human beings manage to do these things despite limited computational resources. So how do we manage to do it? If human cognition is a form computation, then it is possible for a computer-like entity, namely the human brain, to do the things that human beings do. In other words it is possible for a computer-like entity to do the things that the authors ‘prove’ to be impossible, since these entities do these things on a regular basis. Contradiction! The contradiction can be resolved if we take the authors to have proved that it is impossible to duplicate human-like intellectual capacities without running into problems of computational intractability using a certain range of strategies. (The authors imaginary Dr Ingenia trains the AI by exposing it to massive data sets). This cannot be the way we do things, since if it were our cognitive capacities would be impossible: we could not do the things we actually do. But this suggests that it might be possible to create an AI that duplicated human capacities modelled on the ‘strategies’ that we actually employ. There are ways that our brains work and if our brains are akin to computational devices then it ought to be be possible to create a device which works in much the same ways and which can therefore duplicate our cognitive capacities. 

If this line of thought is correct (a very big ‘if’ since there may very well be something I have missed – this ‘off the top of my head’ stuff), then the authors have not shown that Artificial General Intelligence is impossible They have only shown that AGIs based on current models and approaches are impossible. It is not that replacing people with AGI-based robots is not a threat – it just that this threat is punted into a more distant future. 

 Moreover, the threat of technological unemployment does not depend on the prospect of high-performing AGIs. Much dumber limited-domain AIs can put people out of work if they can put in an adequate performance at the a limited range of tasks that a particular employee is required to do. So I am not sure that Rooij, I., Guest, O., Adolfi’s brilliant paper is quite as comforting as Rob Hughes takes it to be. 

Noah Gordon
Noah Gordon
1 year ago
Reply to  Charles Pigden

Hi Charles, I think the issue is that there are three distinct claims of intractability at play here. Let ‘H’ be the function that is supposed to approximate human behavior:

(1) H is computationally intractable.
(2) Finding H is computationally intractable.
(3) An algorithm to identify H is computationally intractable.

Advancing (1) would be a (highly implausible) strongly anti-computationalist view that the authors explicitly deny. This would be the view that you rightly deride as claiming that human beings somehow do the computationally impossible.

The problem, from my reading, is that the paper proves (3), but tries to market that result as being equivalent to (2). In fact, (2) is somewhat nonsensical. As the authors define it, ‘computational tractability’ is a property of algorithms. Saying that “finding H is intractable” makes about as much sense as saying “me picking up my cup is computationally intractable”.

Anyway, even if you find (2) sensible, the fact is that we don’t need a computationally tractable algorithm to reproduce H. The real-world equivalent of a computationally tractable algorithm to reproduce H would be a defined series of steps that identifies H no matter what your starting data and situation is. Even if you don’t have this, you can work empirically to get closer and closer to H. This is what flesh and blood AI researchers are doing; they are not trying to come up with an algorithm that will find AGI across all mathematically possible scenarios.

What’s more, it is an empirical fact that evolution and other natural processes reproduced H. By the authors’ argument, this should should be infeasible because the task is computationally intractable, which is absurd.

I find it disconcerting that this paper already has 57 citations on Google Scholar. Perhaps I will author a response.

Noah Gordon
Noah Gordon
1 year ago
Reply to  Noah Gordon

I’ve slightly misstated things here. Computational tractability is a property of formalized problems, not algorithms. The author’s main result says the following:

MR: There is no algorithm that identifies H in polynomial time.

My main points about the insignificance of this result stand:

  • You don’t need an algorithm to find H, i.e. a series of defined steps that produces H for all mathematically possible starting points, in order to actually reproduce H. This isn’t how AI scientists are trying to reproduce H.
  • Evolution and natural processes reproduced H. However you want to construe this in relation to the mathematical result, it shows that MR is no in-principle obstacle to reproducing H.
David Wallace
David Wallace
1 year ago
Reply to  Noah Gordon

An analogy: very plausibly it is computationally intractable for an algorithm to play perfect chess (I.e., always win); certainly the only method we know to design such an algorithm is computationally intractable. Yet demonstrably it is not computationally intractable to design algorithms which play extraordinarily good chess – not perfect in the literal sense, but reliably better than any human player.

(Borrowed from Dennett’s discussion in Darwin’s Dangerous Idea.)

Noah Gordon
Noah Gordon
1 year ago
Reply to  David Wallace

This is a helpful analogy, but there is one important difference here. The formal result claimed in the paper is about the (worst-case) complexity of creating an approximate AGI algorithm, where ‘approximate’ means “succeeds with non-negligible probability”, which is given a formal definition in the paper. So in the chess analogy this would be like saying that the algorithm for playing chess even approximately perfectly is intractable (has a worst-case complexity higher than polynomial).

Kenny Easwaran
1 year ago
Reply to  Rob Hughes

> LLMs produce impressive results when prompted in a way that is designed to highlight their capabilities. When tested rigorously, they “hallucinate.”

What is the significance that you see about this fact? To me it suggests that LLMs won’t be useful for doing many sorts of tasks completely unsupervised for a long time. But, for instance, a system that “hallucinates” proofs for novel mathematical results 90% of the time, but finds real proofs 10% of the time, would be incredibly valuable to mathematicians. Even more so if the “hallucinated” proofs help the mathematician get a sense of what all the dead-ends are in proof methods that might take hours or days for a human mathematician to explore.

If I had a colleague who knocked on my door every day and told me what he thought was a proof of a new result every day, and 95% of their proofs were wrong, they’d be incredibly annoying. And yet with their attempted proofs and my verification of 5% of them, we would probably be more productive than most pairs of mathematicians.

I don’t mean to claim that any system is doing anything as useful as this yet. But my point is just that the mere fact that these systems often “hallucinate” doesn’t mean that they can’t lead to transformative improvement of our research. (One might say that the entire purpose of academia is a sort of controlled hallucination – almost all published research is incorrect in one way or another, but the point is to be interestingly incorrect in ways that help occasional correct things get figured out.)

Rob Hughes
Rob Hughes
1 year ago
Reply to  Kenny Easwaran

I agree with you that for certain types of creative work, an LLM that produced something both good and genuinely novel 5% of the time would be a socially valuable achievement. What do you think are the odds of this happening in the next 20 years? Based on the evidence available to me, I don’t think this is at all likely.

Many of the tasks (some) people would like LLMs to do need to be done reliably well. Given the unavoidable tendency of LLMs to hallucinate, I have two concerns. One is that all the hype will lead many cost-cutting managers to think LLMs are more trustworthy than they are, and to allow LLMs to operate unsupervised when they really should not be.

My other worry is about how LLM adoption might affect education and job training in the long term. Suppose LLMs are developed to a point at which they can do a wide range of tasks as well as an entry-level employee. Suppose an organization can get equally good performance on these tasks from an LLM supervised by an experienced employee or from an entry-level employee supervised by an experienced employee. Many organizations might choose to use LLMs rather than hiring entry-level people, either because LLMs are cheaper or because they don’t talk back. But in the long run, where are these organizations going to get experienced employees to supervise the LLMs?

We are in danger of eating our seed corn.

David Wallace
David Wallace
1 year ago
Reply to  Kenny Easwaran

A way I think of it is that LLMs might work a bit like an unreliable NP-oracle. If an LLM solution is tractably checkable (e.g. a proof of a math problem), it’s usable. If not (e.g. financial accounting) it isn’t, or not safely.

Simon Goldstein
Simon Goldstein
1 year ago
Reply to  Kenny Easwaran

“To me it suggests that LLMs won’t be useful for doing many sorts of tasks completely unsupervised for a long time.”

I think this is wrong. First, Hallucination rates are dropping. Second, AIs can be supervised for hallucination by other AIs. The key is that verification and production are different tasks. Even if the model makes an error when *producing* text, this doesn’t mean it makes an error when *evaluating* the text for accuracy. This is why the model usually ‘takes it back’ when you point out to it that it has hallucinated: it has switched from production to verification

Daniel Weltman
1 year ago

I am not an AI expert, so forgive me if this question is ridiculous, but I thought AI models “take it back” when you tell them they got it wrong whether or not they in fact got it wrong. You can get them to take back anything with enough badgering (and with many of them it takes barely any badgering).

Noah Gordon
Noah Gordon
1 year ago

For posterity’s sake (and to clear out a bunch of open tabs), here’s a link dump of some sources that didn’t end up making it into the piece:

Some data on difficulty for Harvard MBAs on the job market the last few years

World Economic Forum’s Future of Jobs Report 2025

Opinion analysis on the effect of AGI on wages

And a good piece just came out today on NYT for AGI optimism (keep in mind that it’s for a global mass audience).

60
0
Click here to commentx
()
x