Animals don't "work". Not atleast for their own sake. If there is enough green pasture and water around, they don't even migrate to other places. So if work is meant to provide food and shelter and if machines can ensure that, humans don't need to "work".
Wealth is only a reserve capacity to help future generations so that they don't need to work for their basic needs. But if machines ensure that too, then wealth itself, as a reserve, is unnecessary.
- Work is shifting from building/doing to evaluating, judging, and steering — that's where human value will concentrate.
Other supporting points. ------
- No lab milestone or "RSI breakthrough" will suddenly eliminate jobs — economic impact unfolds gradually over decades.
- Reliability, not raw capability, is the real bottleneck holding back AI automation today.
- Historically, making work cheaper/faster (ATMs, radiology, coding) has grown employment, not destroyed it.
- Superintelligence claims misunderstand human intelligence, which is itself amplified by tools like AI ("co-superintelligence").
It is not a good idea to compress articles like this but there are many of these opinions to read and trying to get to the point quickly to uncover new viewpoints.
According to WRITER’s 2026 Enterprise Adoption Survey, 44% of Gen Z employees admit to sabotaging their company's AI strategy in at least one way compared to 29% of employees overall.
Sabotage behaviours include entering proprietary information into AI tools, using non-approved AI tools, refusing to use AI tools or outputs, ignoring guidelines or best practices, intentionally generating low-quality outputs, refusing to take AI training and tampering with performance metrics to make AI appear to underperform.
So if things continue as they are today, I think in the near future, being a software developer is going to be more analogous to the medical field, where in the medical field you have different levels of professional expertise.
Some will be like nurses, and some will be closer to a medic and a smaller set will be like doctors. Each with increasingly required knowledge and experience to fulfill a needed role.
Those who used to be actual software developers are going to be (or have to become) more in the doctor role with years of internship and practical experience to be the architects guiding the overall AI implementation of software development in organizations.
The medics are going to be people who are semi-technical, where they have some technical understanding but they don't dedicate themselves to it, like say product managers, where they jump in to help development along, but don't need to have many years of experience or very deep technical knowledge.
At the nurse level, it's probably going to be similar to what people would do in the past with no code tools, where somebody in marketing who knows very little to nothing about coding at all is just going to directly converse with AI systems, but they'll never be likely to get anything more advanced than the tools they could think up for themselves.
Of course, it's so hard to tell what the next big discovery or changes to the nature of world society might push things in one direction or another.
What upset me a bit were phrases like “This is not a slogan. It’s a framework” which immediately devalued the work for me.
I have read so much Ai generated text recently, that I developed some AI-fatigue or AI-burnout, and I’m wondering if that might hit more fields - making more humans reject Ai work.
To be clear, I still like the text and I don’t know if it was written (partially) by Ai or not - but it’s this uncanny feeling I got reading it.
But we already do have have some kind of measurement of most of these types of side factors, and they actually aren't at zero and are increasing rapidly. So the implication that they will not be human level until decades from now is just (hopeful?) speculation or fuzzy thinking.
To me this looks like a really academic and official sounding version of the same quasi-religious hopium that usually defends the sanctity of the human. He is essentially saying that there is just something so special about humans that it will never be reproduced in a machine. It's very similar to dualism (and in many people actually is religious dualism). No AI is going to have human creativity or judgement. Not anytime soon. Why? Well, we all just _know_ that's not possible. Okay, maybe in a couple of decades (but they don't necessarily believe that anyway). Why would that take decades? Well we all can just _tell_ it's no where close, right? Because AI of today just isn't special like humans.
Aside from that worldview issue, I think that people still are not taking seriously or internalizing the concept of exponential improvement.
Computing efficiency gains can actually level off. In fact, they have many, many times before. But they always tilt back up again when we invent the next approach to get beyond the current level. This is how it has been for 90 years.
There are multiple ways that we continue to see huge gains in AI software, architecture, and hardware. There are huge efficiency gains available still as we move towards more radical fully compute in memory and/or analog approaches and other options like models implemented in hardware.
- i would assume it is reasonable that anyone comes and see what other posts a person has written except you cant find that page anywhere linked
It's possible that the elite will control most part of wealth generation, keeping it for themselves. The rest of society will develop an underclass economy and work for each other. Essentially a worldwide slum.
There will always be new, hard problems to work on. AI will not, and can not eliminate that.
Ourselves.
I sort of worry about things like AI figuring out scripts so well that even multi-tier support work is gone. And learning how to write fiction or create foods so in accordance to our tastes (sugar, fat, etc with food, exactly what each of us is interested in, with writing) that we even lose those truly human creative jobs. Might not ever wanna leave those bubbles.
So much of the human drive is exploration and why and what if. Assuming everyone in the world can have no money problems, what will AI not be able to figure out? Will we enjoy the equivalent of a major breakthrough if an AI solves it in five minutes, or just the outcome? Why learn things?
AI could be a horrible jailor. And better at cancelling than any perhaps sager Gen Z or millenial. Bears some caution to be wary of this and where that power sinkhole will go.
But then, I still think the previous AI winters were more a result of sense and caution than most of us know, and we cannot fathom our species' ways of reasoning/thought processes the way we did as a species thirty, fifty, eighty years ago. Erring on the side of caution is not a terrible thing.
I mean, I have worked and work with AI, but it seems weird for us as a species not to have placed guardrails to prevent us from wiping one anothers' careers and relationships out. What will we talk about? If our generative AIs should be allowed to date?
Again, I am assuming a fast, though not sudden, acceleration that would compound, and sooner than most probably think.
However, this following quote has a simple reason that I don’t see anywhere in the article or framework:
“”” Why is there a huge gap between what people in various occupations could be using AI for and what they’re actually using it for? One reason could be that people are slow to adopt technology, and that’s certainly part of our framework. “””
I would like to add a reason: that the Silicon Valley companies who developed the LLMs are brigands: cognizant of their actions, they have stolen (and continue to steal) the world’s copyrighted material and are selling it back to the masses and the politicians as if they are the arbiters of information itself.
Specifically responding to the quoted question, I could be using Claude or ChatGPT or Grok or DeepSeek or any other to have come up with this comment, or to write emails, or to implement my Python for me, etc., but I use none of them for anything. Doing business with brigands is a choice, and a choice that I hope becomes less and less palatable so that the financial, political, social, and moral fever that is our zeitgeist finally breaks.
The only question is what next?
Our only hope is to teach the AI to meld with our bodies and use them for gestation, energy or hibernation. The alternative is sustenance.
We can start considering to reconsider when we have the begining of an answer to that first fatal issue.
So much brain power, including mine, wasted to this stuff instead of useful or enjoyable stuff, that's quite sad.
Does he want to fund the arts? Humanities?
So they are running unchecked to get richer, get more power, get more stuff. What needs to happen for that to change?
My dog was perfectly happy to snooze most of the day, play with toys occasionally, go for walks and try to hunt rodents. I on the other hand would be incredibly bored with that.
Some may consider local LLMs to help with that as the power of LLMs would be more distributed. I think local LLMs would help marginally. Companies as an entity would still be better positioned to use these to their advantage vis a vis individuals.
So projecting into the future I still wonder if it won't become more challenging for individuals to make a living on average.
When enough green pastures are around animals usually reproduce till that is no longer the case and they need to move to other pastures. They continue doing this untill all green pastures are occupied. After this they start competing with one another to compete for the green pastures already occupied. Some animals will be so succesfull that they take larger green pastures, letting others starve. If by some miraclelous event (ai?) suddenly a lot of green pastures arise, animals will simply reproduce again till they occupied all green pastures again.
Looking at us? The agraric revolution, neither the Industrial revolution, nor this ai revolution have decreased how much we work on our 'jobs' (finding food indirectly). All it did was increasing the population so that all green pastures became occupied.
If you think it is different, just think of how many people write books professionally, or even publish online.
Once the noise settles down a bit and boardroom shakes off their delusions as you can see in rehiring in Ford and Zuck who was very bull on AI remark about "not being it". It will be just the same, but different.
Ford: https://www.bbc.com/news/articles/cgrkd41n2v9o IBM: https://qz.com/companies-rehiring-workers-ai-layoffs-automat...
> A battle of two narratives > Build wealth before AI obviates our skills > Build skills, agency, taste, judgement
both narratives are portrayed as being odds with each other but, I can't come up with a single "build wealth" scenario that doesn't involve building skills, agency, taste and judgement.
what am I missing ?
https://news.ycombinator.com/item?id=48743713
---
We've just done an official evaluation at work, using extensive statistics on our gigantic monorepo in a company with ~2000 devs over the course of 2 years, everyone from hardware engineers to regular old frontend engineers. It's a highly profitable and mature public company, and has been for going on a decade at this point without missing a beat. We were given infinite access & budgets to basically any and all AI tooling we could imagine, and we have several "AI Native" teams (whatever the fuck that even means). We're doing agentic coding, we have harnesses of all kind, skills, we have many teams doing spec-driven development, designers using all the various things like Figma Make and access to tools like Devin/Factory Droid/Claude Code/Codex/etc.
This is all to say, we as a company are using AI a lot in all possible corners, but thankfully our leadership isn't schizophrenic and isn't mandating everyone hit token limits or whatever, it's more of a "Let's see what works and what doesn't" type of thing, and we measure a lot of statistics. Nobody here really cares whether LLMs are the next coming of Christ or not, as a company there are many people (even in SLT) that are indifferent to LLMs, and many who are reasonably hyped.
I wish I could link to the actual document we were all shown since it has a beautiful breakdown of the methodology and a fine-grained breakdown of the stats and the categories measured, but in the grand scheme of things, ALL the AI tooling we have implemented (at least on the engineering side of the equation) has contributed to a total of... drum roll please... 7 (seven) Percent overall productivity increase! The most productive teams saw a productivity increase of around 20%, while some teams actually saw drops in productivity into the negative percentage points. My team, none of us really give a shit about AI and we're somewhere in the 3-5% range on certain categories of tasks, which I'd say is a fairly good assessment.
Productivity here is measured in many ways, including but not limited to speed of MR review and merge times, feature/ticket/roadmap closure/delivery, rollback/revert incidence rate, how often people interact with the MR review bots and implement their suggestions/fixes, how many times people check back on AI transcriptions/meeting notes (hint: Nobody looks back on any of it, it's all just noise that gets generated and never actually referenced outside a few extremely rare cases) and many more things I'm forgetting. It is an imperfect number of course, because measuring productivity in engineering is a sisyphean task, but in my opinion it is accurate to the reality on the ground and outside of all the hype and marketing bullshit.
So, I remain thoroughly unconvinced of these personal anecdotes of people being "massively" more productive, especially once you factor in the fact that we now have a 2000EUR budget/month/dev for all the AI tooling, those productivity numbers start looking pathetic once you factor in the costs (which are only increasing as the AI companies need to start recouping the gazillions they've burned). Some teams have started begging to disable coderabbit and other similar tools in their MRs because they're producing nothing but walls of noise that makes reviewing any MR a nightmare of sludging through endless slop of useless bullshit, ours included.
AI can slop fork or clone existing software well, but a clone of an existing game is pointless, it's basically guaranteed to be derivative and worse than the original game, and games aren't so expensive that you can't just buy the original. AI can't know if new mechanics or angles to an existing genre will feel good to play, or if a new genre is fun, that requires a human to experience the game in its totality.
Games are also very resilient to sloppy AI coding, and if an indy game crashes nobody is getting paged.
Until the machines aren't owned by anyone (or owned by everyone, take your pick on the phrasing), the owners of the machine have no need to keep you alive.
This take is basically "Don't worry, people like Sam Altman are looking out for us"...
How many man-hours go into various parts of the advertising distribution chain? Though a certain fraction of that energy goes to connecting people with goods and services they might find valuable, most of it goes into shifting numbers around for people that already don't personally have to worry about money.
We don't need to find endless ways for people to spin wheels, but as long as we're worried about "jobs," we will. We just need to find the social structures to provide people with basic needs and reserve "work" for things that are vital to society or truly inspired.
We will never automate all work so we with half of humanity doing nothing of value it will be a struggle for the people who do nothing of value to convince people to do work for them.
We can see it now where products dont target the people without money. There is no point because they cant give you any reward so instead you do your work for the people that can give you something in return. We can use the government to stimulate and balance this a bit but at a certain point the number gets to high and things collapse.
Our technology has long since surpassed our capacity to understand its impact or to morally restrain its use - we've gone from chivalry to "oops, that was a school, oh well"
We will solve problems faster though, but we aren't showing signs that AI is making us better humans, and that is the problem.
These "needs" are sometimes enforced by the systems and government so that people don't stay away from the work and "economy" keeps churning. The housing prices could be a way to keep the people working for loan payments.
Instant foods, nursing homes for elderly, creches, roads, commuter trains - are all ways to have more workers and make them focused on work.
Very open definition of sabotage.
Jinxzi has Bernard as friend though
I didn't do fewer hours in these weeks but had time to explore and innovate a little.
I think many people will wish it was like this. But when AI becomes so capable that a nobody who doesn't understands computers a well as a "doctor" uses AI to create something BETTER or superior or just as good then employers will think... why am I paying that "doctor" more than that nobody?
See right now everything is fine, because AI is not that good yet. But if it gets better. Well. The future will be one where we trust AI more and more.
But I do believe (gut feeling) the novelty's worn off for a lot of people.
I cant read this shit anymore.
"creating cross functional AI evaluation team to keep the company honest"
Fucking garbage.
I have an old book between me and the keyboard, and just read it while the AI is thinking, so I can read non AI words.. otherwise its like reading books by the same author over and over and over again.
I will go touch some grass now.
The problem is they are now paying me more, plus paying for the cost of using the AI, and the needless complexity also slows down the employees. So more costs there as well, any future debugging is going to cost far more and at the end of the day they are getting less quality on the core function but far more presentation data that is essentially meaningless.
For example, all the work like translations that used to be done by humans, now it is a CMS AI feature.
Secondly teams setup.
It used to be we did everything ourselves for development, then cloud, SaaS products and serverless decreased the teams size required for delivery.
Now with AI, there is an even greater push for low code/no code tooling, with agents, leaving the actual programming left for MCP tools that might not yet be available for the project.
Thus you get a team of five doing what used to be about 15 a decade ago.
This is true as a sentiment, but my understanding is that the majority of students are overwhelmingly using AI for ~everything. If a thing provides massive utility people will use it.
And what reason do they have for killing everyone else? Where is this abject nonsense even coming from? You can't just assert that the "owners of the machine" are all cartoonishly evil for some unknown reason.
All else being equal, at least Sam Altman et al. aren't constantly making up fantasies about exterminating people.
Not that much. AI is already heavily automating advertising and has been for a long time. And a lot of activity that was previously exposed to untargeted passive advertising - like TV - has shifted to mediums that don't have any, like Netflix (for most subscribers).
And as you note, advertising isn't useless. It's how people find out about things they didn't know they wanted.
In such a future there would be a handful of lucky well paid artisans, a healthy community of hobbyists, and the overwhelming population of the planet who would be perfectly content to delegate their entire software diet to generative superplatforms that script themselves to perform any arbitrary software function. The idea of paying for individual bespoke software programs would become an anachronism for an era where software was so difficult to produce that entire teams spent years painstakingly tweaking programs to spec.
So I think it would be more comparable to something like literacy. There was a time when that was a fairly uncommon and highly valued skill. Now the guy flipping burgers or pouring a cup of coffee is also almost certainly fully literate. And in fact many jobs have evolved in a way such that it became mandatory, but only because it was already ubiquitous. I expect to see the same thing with software. The industry of producing software that do fairly simple tasks will probably die, but in its place will be a vast array of heavily customized and oft iterated software for companies and people achieving their own stuff.
The mobile industry is a perfect example of where this will be massive shift. Right now there's a million mobile apps to execute extremely basic functionality on phones, but it's loaded with advertising, begging, and general annoyances. As are the app stores themselves. When you can make software that does that in a few minutes with a single prompt, and people realize this (as we're already practically at this point), then that will be the end of those apps. This is because the one thing LLMs have shown is that natural language interfaces are way less friction than using search, whether on the web or an app store. And so there will be a time when it will be lower friction to simply just quickly build your own app to do [whatever] than dealing with somebody trying to monetize an alarm clock.
https://www.reddit.com/r/gifs/comments/3p0b3i/graphic_design...
Why? Because it resembles the typical "You're not just implementing a text editor, you're reshaping the text editing landscape"?
But I'm trying to figure out if this is a thing in my organization. So far the channels discussing AI have been about how they use AI, how to build agents, and complaining they've run out of tokens - but nobody says anything about how much work they finished, how it impacted the product(s), the user, and the company's bottom line.
I don't know if they just don't know, don't care, don't measure, or they don't actually want to admit they're just burning through tokens with nothing to show for it.
That said, right now I'm using AI to extract some code from our app to create a minimal reproduction for an issue we're having with a 3rd party library, it's a huge time saver there - that is, I wouldn't have bothered making a reproduction because of how time consuming it can be.
So to answer your question, it's not that "I have less to do", it's "I do things I wouldn't have done otherwise because they cost less effort/time". Which is also the promise of e.g. automation - the automated loom didn't just reduce how long it takes to make cloth, it made it so people make a bajillion times more cloth.
It pushes the bulk of the work to review. So teams with good practices can account for this.
For junior teams, the time saved is massive, because they aren't doing all the other practices required to prevent technical debt.
I would doubt however that this would be an 'Equals' or 'Implies' scenario. Let go of seeing either of them as binary, and then not even as scalers.
> It’s funny. I was looking at my GH activity graph. It’s been pretty solid green, for years. I stay busy.
> But since I’ve been using an LLM, it’s been bright green.
> I always check in code manually. I don’t let the LLM do it.
Depends on what kind of game you are building. I can tell you that even Claude Opus absolutely hates desync bugs and has a rather hard tracing them. Maybe Fable or chatgpt-5.6 are better?
Of course, billionaires have other plans, and are the main obstacle in achieving any sort of social cohesion.
Everybody says they hate heroin but once you try it, you can't get enough of it.
The sell was - Creating your apps was easier than before. You don't need to know any programming language. You don't need IT. You have a GUI and just drag and drop and visual your apps. Just fill the fields and voila your app is ready.
The criticism of the approach was that low codes apps might not be secure, unsuitable for large scale and critical apps, and lead to increase in unsupported apps by "shadow IT".
The same reasoning exists today for AI coding. Anyone can create apps, its not great for complex and mission critical apps, might not be secure etc. And lots of discussion like your post parsing the future too closely.
Low/no code apps have continue to grow. But many of the low/no code tool users are developers who use it to make their jobs easier.
While some might say - this time it is different. I believe we are currently we are the beginning of the cycle so everyone is excited to use their PowerApps shaped Claude Code/Codex to make their 100th budget tracking app but as time passes and edge cases are figured out the biggest users for AI are going to be software developers (and I believe they currently are the biggest users).
As for software engineering jobs just like before engineers will be expected to output more and faster. This has happened with every innovation from assembly compiler to IDE to low/no code ways of building.
The different levels of expertise exists even today and it will remain so.
just?
They say they are, but they have actually been using it so much that they can't even do they homework without AI at this point.
The drive some individuals have to control / exploit others isn't limited to the rich, they're just in a better position to exercise it
There are people who have absolutely no interest in what other people around them are doing; there are others who will move heaven and earth to help those around them whenever they can, even at the expense of their own interests; others still who can't help but compare themselves to everyone and make it everyone else's problem if they find themselves lacking; and yet others who at all costs want people to afford them "respect" which when you drill into it means stroking their ego
I'm sure we've all experienced petty individuals who've been granted just enough power that they can gain personal satisfaction from leveraging it to make themselves feel better at the expense of all other considerations
These are all just some of the many flavours of humanity
Your analogy isn’t necessarily wrong, but it might ignore the extreme importance of nurses. Many medical facilities are only staffed with permanent nurses, with doctors helicoptering in, from time to time, to take care of specific duties that may require certain licenses, or provide specific advice.
So lots of jobs for nurses.
I can't wait to be kept in agistment by my overlords, fed on treacle and oats, ridden in circles once a fortnight, and shot when I break a leg.
[0] https://www.researchgate.net/figure/United-States-Farm-based...
But it's at my day job, and it's because I was able to write a prompt which automates having Copilot review uploaded scanned PDFs of invoices with checks (and the bank line obscured with a pen, so no PII) and then write a batch file which renames the files per a file-naming convention, removing the need to open them in batches of 50, find the Invoice ID, re-save using that filename, then quit and re-launch Adobe Acrobat (if left running, eventually I run into a bug where it stops saving files), then run a .bat file which renames based on Invoice ID as a filename.
Problem of course is I've been running into a limit of number of allowed files per 24 hr. period.
Even if it's not less work, it feels like less effort.
Once you've used a BBS or modem, "online banking" is an obvious idea - just a very difficult one to implement until you add on decades of security and improved connectivity.
But also, the article is explicitly arguing against armageddon.
The medical field is also going to change though. Massively. Because people are going to realize you don’t need to pay someone $400k per year to hand out advice about moderate exercise and which antibiotic is appropriate for a sneeze-cough with yellow mucus.
Regulation isn’t going to prevent this. AI is already way too easily accessible to ever rein it in again. Not to mention that the US now has serious competition from a hostile country, so they can regulate their own AIs all they want without it making a difference in practice.
"Work" doesn't exist to keep people busy, it exists to keep them alive.
Luxury horse living during the heyday of working horses and pit ponies, "horse power" wasn't left ideal for a fortnight.
> How many horses do you see now that the world
Personally, a surprising number perhaps, there's a pony club at the top of my street in town, and the area is still littered with horses and other livestock.
This isn't my area, but it's not dissimilar: https://news.ycombinator.com/item?id=45623799
Full size image: https://live-production.wcms.abc-cdn.net.au/a26664f6500a7c74...
That seems a rather high baseline, maybe even more significant than younger generations putting another 50% on top of it as they are wont to do.
It's just hard to design robots that can handle patients that may be simultaneously fragile, mentally handicapped and aggressive in a way that doesn't hurt them and respects their rights. It can take multiple human nurses to do a seemingly simple thing like changing a diaper.
I'd like to imagine that it's ultimately going to work out for humanity to be in a better condition overall if that happens.
However, if there reaches a level of near or full autonomy in all aspects of knowledge work, then there is a strong possibility that the field with the highest growth potential is going to be in "Private Security"; as those who have access to the most resources seek to defend their own positions in a societal return towards aristocracy and serfdom.
Let's hope not though.
Those are the people who will get the boost from AI coding. They have the "lazy" mindset every good programmer has. You do a thing 3 times manually and then your brain goes, nnnope. This needs to be automated. And you get to work.
With the n8n's of the world they can get something done, limited by what they offer. Same with Excel, there's a limit to what you can do with macros and VBA.
But with an AI agent there really is no limit. I have seen first hand the things domain experts can do when just given a Claude seat and the permission to build stuff. Things I could've easily built, but never would've found out are actually needed.
And doing the same thing with a programming team and project managers in the loop with the domain expert being the product owner would've taken months and cost six figures in salaries.
Also new heroin users are only 30% likely to become addicted. https://jamanetwork.com/journals/jamapsychiatry/fullarticle/...
There's a much bigger group of people building games with Unity than the number of people building complicated engines like Unity. Same for [insert Javascript framework here].
My mother worked in an office in London as a shorthand typist in the 60s, along with thousands of other young women.
At some point computers began to enter offices, bosses typed their own letters and then emails, and this category of job simply evaporated. Of course there was still SOME secretarial work, and some workers retrained to do it, but most simply had to leave the sector.
Isn't that a common story with technology replacing workers?
I don't think your assessment that this is mostly automated stands in the face of the number of man-hours both of these companies purchase each year.
with the latest agents, you can even be vague and they'll probably do an ok job if it's not something super complicated. some models are also pretty good at stopping to let you know about stuff you haven't thought about and ask for clarification.
so you can outsource thinking now, not just the doing part.
Maybe for a very loose definition of medical facilities that includes assisted living facilities.
But for example in an ER, nurses come and go with very rapid turnover and it’s common to staff with temporary travel nurses.
> nurses tend to do most of the actual work
Techs, environmental services, phlebotomists, respiratory therapists, CNAs etc. probably do more of the “work” than nurses.
> highly experienced nurses actually taking up the mantle for many duties often done by doctors
Only if they go back to school and become a Nurse Practitioner or CRNA, but in that case they are no longer functioning as a nurse. Even then they are general operating under the direct supervision of a physician.
> Nurses are in far higher demand, than doctors.
Only in absolute numbers. It’s far harder to hire a doctor than it is a nurse. I know an ex-NFL player who works as a physician recruiter.
Wishful thinking by the managerial class. At best they can vibe code but they can’t verify that what was written is correct.
Who said they would kill everyone else? Since the rest of your strawmanned caricatured argument is a response to an argument never made, there's no point in addressing it.
I don’t understand how people can say this and then continue talking about software. So we’re saying machines can now casually do complex and cognitively demanding jobs like software development (or 90% of all white-collar jobs out there) and we’re NOT worried about the lynching mob going door to door and hanging IT people on lampposts? And I’m being serious, the impact this would have on societies would be unprecedented.
Exactly. (Experienced) Nurses can do ALL of the basic stuff we currently use doctors for. They diagnose patients with a glance (not by going to medical school but with raw lived experience on the field), administer medications etc.
And in some countries (I think the UK?) they are allowed to do more of the doctor-y stuff to free Actual Doctors to do the stuff they studied for years for.
The same thing will come to the software industry. All of the basic CRUD HTTP API crap can be done by "nurses", just run of the mill domain experts can wire up a basic UI + API + DB -combo for a product MVP with AI assistance.
And with a "doctor" (more experienced programmer/architect) guiding the process, the quality will be slightly above average. Not artisanal excellence, but how many services actually need that, really?
I think I've mentioned this before, but the only ones whose jobs are at risk at the moment are the mid-tier programmers. Either by skill or experience.
Juniors can learn to work with AIs and easily become the "nurses" of programming.
Programmers with decades of experience in multiple fields can either go full the artisanal-no-AI route and do the things AI's suck at or do AI assisted programming, which isn't really that different from working with a team.
But the midtier people who are too well paid to be "nurses" and not experienced enough to work as "doctors" might find it hard to find a place for themselves.
Human agency is real and powerful - unless humans want to automate all work, it won't happen.
I spent some time in the military, and my expression of medics and nurses are mostly derived from that experience, where I'm referring to a nurse as just any warm body who is able to provide aid.
For professional nurses who might work in hospitals, I'm sure that many of them have significant knowledge and experience to be very effective in providing medical assistance.
Who is going to realize that?
The same forces that prevent you from walking into a pharmacy and asking for antibiotics based on what you found on WebMD will prevent you from doing it with a ChatGPT printout in hand. Lawyers and doctors are the best-known examples of industries that are in control of who gets admitted to practice the profession.
I don't buy this. Tons of people in the West live on welfare and don't contribute shit to the economy, they are net takers. No one is plotting to kill them, in fact they're (arguably) getting more than they did 50 years ago and definitely more a 100 years ago. We live in democracies, not in a dystopian nightmare and I don't see why A.I is going to change that dynamic. People vote, people control the government, the government controls the armed forces, private citizens are not allowed to gather unregulated weapons... billionaire or no billionaire I don't see how you can beat that. And also, it's not like the billionaire class is some kind of a cohesive group that wants to work together to rule the world - they pretty much compete with and hate each other. I don't think Musk, Altman, Hassabis and Larry and Sergey are all going to agree to work together on "controling the machines" and killing all the other people.
Is it different? because then:
> you can even be vague, and they'll probably do an ok job if it's not something super complicated
That's lot of qualifiers to say maybe, should be possible - how does someone who is not a software engineer know what is complex, vague or ok. There are lot of assumptions to say - This time it is different.
It is entirely possible it is different but I am yet to find a convincing argument.
But are they also paying for it? Or simply using the free version because it's available?
But that's neither here nor there, since we're talking about a world with an artificial superintelligence. That's the hypothetical here.
Heh. I wouldn't be so fast in calling this one, yet. It might go either way, or a combination of the two, who knows... Things are moving and progressing fast enough to at least be wary of, and keeping an eye on things. I wouldn't be surprised of any outcome, tbh.
On the one hand, PMs get to hone in a set of skills that include lots of juggling resources, bridging the gap between stake holders and people that execute, work with different layers of technical expertise, lots of back and forth and so on. Having to "call" a thing after talking to 5 people from 6 different PoVs is interestingly something that could make them quite good at using "AI". They're already used to working with "jagged expertise" so to say. And that's really close to what "agents" do today. You might get a session where claude/gpt/gemeni turns out a brilliant piece of technical artefact, or you might get an average piece of content that misses key important aspects that a "human" could see from a mile away.
On the other hand, LLMs do one thing quite well: they raise the floor of what "minimum effort" in a thing gets you. So things like language barriers, basic processes, procedures, etc. get elevated to a level where you can use these tools and be better off than not using these concepts at all. So a very talented engineer that previously had these issues, could now use LLMs to get better at auxiliary but "unimportant" aspects of integrated that technical experience into other places. In a "90% of the time people expect this to work like this" way, but with the floor being raised. One quick example of this would be a technically sound piece of software that now has bare-minimum ux/ui from this century, instead of "engineer GUI" aspect :)
Who knows where this all goes. Yesterday there was an article here about someone taking a bunch of languages, asking an agent to mix concepts from all of them and slop a "new uber language" design. Is it the ultimate language that many have tried before? 99% sure it's not. Or at least this one won't be. But there's a chance, with how many people can now prompt their way to a PoC quite quickly, that we eventually get something cool. No idea how that'd look like, but it's likely the "you'll know it when you see it" kind of thing.
That’s not true. They do completely different jobs. Some experienced nurses could do some things that we currently require doctors to do.
Other experienced nurses have problems doing basic dosage calculation.
The main difference that addition to med school being much more selective, rigorous, and longer, residency is regimented, standardized, supervised, and evaluated in a way that nursing experience isn’t.
There are entire teams of doctors who spend hours every month discussing each resident’s progress (I know because my wife runs the resident program for her department and I overhear the discussions in the background).
You might have one nurse who worked in the ER for 5 years who can diagnose appendicitis as well as a doctor, but another who worked there for 20 years who couldn’t even begin to do that.
And since diagnosing appendicitis is not part of their job description, both could have absolutely stellar performance reviews.
And the nurse who can diagnose appendicitis, might be terrible at other tasks normally handled by a doctor.
The solution for this is to put them through more standardized on the job training/evaluation, but you’ve just reinvented residency at that point. That’s the thing people don’t understand about residency, it’s as much evaluation as it is training.
No amount of on the job experience is equivalent because job experience isn’t rigorously regimented, and evaluated.
Yes, but this might also be a counter-point to your position. In a world of rising baselines, having a "pot but without all the bells and whistles" might be a thing people need. Instead of having a pot that holds fluids at every conceivable angle with 20 temperature thresholds, 5 versions that hold liquid in a cloud for you, with monthly subscriptions, having the ability to get a "like a pot but make it left handed and only holds sand because this is what I need" could turn out to be something that people want/need/end up chasing.
In other words spec down, not up, and base + my particular kind of a pot.
> Maybe for a very loose definition of medical facilities that includes assisted living facilities.
Or midwives, for unimportant things like giving birth /s obviously
The position of (for instance) the check # on a physical check varies quite widely, as can date position and format, and a fair number of them are still written by hand.
On top of that, I'm scanning thousands of these each year, and the invoice underneath the check ranges from the gamut of: "pristine copy just re-printed 'cause the customer didn't include one" through "bad inkjet photocopy of a photograph taken w/ a potato phone and then printed" and includes variations such as "customer included half-a-dozen invoices to be paid w/ one check, and if arranging them and the check _just so_ all will fit in the document camera window, saving a trip to the sheet feeder scanner"
A co-worker actually worked on this for a different program, one where actually sending in paperwork in good condition was expected and customary, but his system had a reject/failure rate of ~10% --- that would be almost 1,000 invoices each year for the program I am handling.
No one said anything about killing.
And people had to fight for it. Through voting, yes, but also through strikes, sometimes riots, and some died for it. People all the world fought and died for the right to strike (basically, the right to refuse to work). For the right to organise as groups. For the right to only work 8 hours a day, for the right to a have a weekly day off. For the right to pay for pensions. For the right to have some sort of health insurance.
All over the world. Some still do. Many still don’t have half those rights.
Now, I don’t know about any billionaires specifically. But welfare didn’t fall from the sky. Or anyone’s good grace. It’s all about leverage.
If in doubt, looking back at the late 1800s and early 1900s work and life conditions across the Western world is enlightening.
It may take vastly more training but on average a full annual physical provides less benefit on average than a 30 second vaccination requiring minimal training. Value creation and skill are wildly different things in the medical profession.
Even now the machine owners have enough power to change laws to their liking, bending governments and public opinion to their will.
I have no idea what that means.
> Why does the "world" have that power?
Okay, lets say "the environment". The universe does not care whether you live or die, so why are you so sure that someone will paternalistically care for you?
On what are you basing this viewpoint on?
> Machine owners don't have the power, for example.
That's because the world doesn't care about them either. So why would the world care about you?
Garbage in, Garbage out.
The University market is brutal too. If you aren't using AI too, you are falling behind. Many see it as a means to an end.
Ah I see you're familiar with PowerPot by Atlassian(TM)
Very powerful potware, pity about how much soil it leaks.
When lawyers and writers are talking to me about "docker containers" and "agents" I assure you that the amount of code out there is going to grow.
Who said anything about killing?
Tests can't be exhaustive, whereas even a mediocre developer is likely to possess enough background knowledge to notice risks that a non-developer would not think to test.
Some people keep forgetting that we haven't reached AGI yet. These tools can still make serious and sometimes obvious mistakes. Not long ago, vibe-coded software could embed credentials directly. That particular blunder seems to have been addressed, but there is still no reliable way to tell an LLM to avoid every class of obvious blunder.
But nurses aren’t the driver of vaccination. Nurses are anti-vaxxers at a much higher rate than physicians, and if you removed physicians from the equation, vaccine rates would plummet.
>total medical care provided
At the end of the day a physician is responsible for every patient that comes through. Nurses are incredibly important, but all of the care they provide is at the direction of a physician, so it’s impossible to separate out “total medical care provided” into doctor vs nurse.
Well, one thing about AI... if it does become our overlords, maybe it won't be so eager to be wheedled into giving passing grades. :/
Have you read the research that says AIs are more likely to react favourably to output based on their own model compared to those of a different AI model? Guessing teachers subconsciously grade similarly, like people who use one model get more of some grade than people who use another one...
Guessing this is also why so many liberal arts majors are being cut.
I 'sorta' get why people might use AI in a required class though I am not for it, but why major in something and do it? I mean aside from wanting money (and, really, many of those majors don't make much).
I.e. some source code gets compiled, the output gets decompiled again, the result gets compiled again, and so on.
Now do the same with LLM: start with a prompt, it generates a program. Then an LLM examines that program by looking at the source, running it, etc., and is tasked to describe it, i.e. turn it into a prompt again. Then that prompt is used to generate a program, which is examined and turned into a prompt again, and so on.
Someone she’s knows calls her up late and says their child has a minor fever, upset stomach or some other mild symptom. They ask “should I go to the ER or can I just wait till tomorrow to go to the pediatrician?”
My wife says it’s totally fine to wait until tomorrow. That is not an emergency. And frequently she’ll tell them they don’t even need to go to their pediatrician unless it gets worse or continues for several days.
9/10 what happens is that the person goes to the ER or their pediatrician first thing in the morning and calls my wife to tell her “you were right, it wasn’t an emergency”.
Despite that my wife is more highly trained than either their pediatrician or a non-pediatric ER doctor, and they know and trust her, they still wanted to see someone in person.
It doesn’t matter how accurate AI diagnosis gets, that’s not going to change anytime soon.
If we are just talking about the physical environment then totally fine - a rock doesn't care about me, for example.
Humans tend to care about humans at least in some situations on average, though.
The odd thing is that after the pandemic just showed what societies can enforce, somehow it all is forgotten again when it comes to who holds what power.
Mostly for drugs that should probably be OTC anyway. Try doing that for a narcotic, or for an expensive biologic.
You could always lie to a doctor about your symptoms.
And I think that example largely carries over to many other big picture and/or creative tasks. I think it would likely require another revolutionary leap for LLMs to get to that point, and that leap is probably not coming any time soon, if ever, at least not with LLMs.
Another practical issue is that LLMs, as difficult as it can be to imagine at times, are still just glorified autocomplete machines. And so genuinely novel advances will not be coming from LLMs. For instance take LLMs back to the early 20th century and none of them will discover relativity, even with infinite power/time. It's just so unlike their training material, and in direct contradiction to much of it, that it's simply outside their reach, let alone grasp.
Well, that someone might use AI diagnostic tools too.
Healthcare is a world of diminishing returns at ever increasing costs. But that doesn’t mean the low effort first steps are less valuable.
Look, I actually agree: humans do care about humans, on average. The problem is that power tends to concentrate to those who don't care about humans.
When almost all jobs can be replaced by machines, the owners of those machines are unlikely to be much different from Sam Altman and Co. They aren't going to sit down and say "Well, look, we have all the wealth, and means of wealth. Let's give everyone food, shelter, clothing, because our machines can do it".
What is necessary are pretty basic killer robots with the capability of determining who is a friend and who is a foe. These systems already exist.
Labor is a commodity that has had a diminishing value in our current, capitalist system. That means, people who rely on their labor(most of us) instead of capital(a few), have control over less and less resources(this is already happening).
With the current trend, what generation of your grand children will have nothing at all?
Or perhaps, they stand to benefit?
It's like a light switch where the light has a slightly different colour every time you turn it on, sometimes the colour is very different, and sometimes it emits bursts of heavy gamma radiation. And the argument is "you flick a switch, and what you perceive in that room changes, it's the same experience." (which is a neat way to put it, because it's true even when you get evaporated by the gamma radiation burst)
That this isn't "the same kind of tool" than just a light switch and a bulb isn't even an interesting conversation to have. Of course it's not, not even close. A piece of string is not a steel rod just because you can stretch it out, take a photo of it, and play pretend.
So the really interesting bit to me is that this level of argument is even made. It's nonsense, but if the nonsense gets repeated often enough, maybe the people refuting it can be exhausted and then we can replace concrete with pudding, right? Not that this is your intent, but this stuff just feels like the thing you'd hear in a digestive tract, not at a table between peers.
As for AI fixing things... Sure, at least for a while. But in my experience, AI can hit a wall and start going in circles. It doesn't happen often, but it can happen. And when it does, what recourse will you have for fixing the tangled mess that vibe coding tends to produce?
Physicians have to understand the diagnosis and treatment plan. At the end of the day, they are legally responsible for the patient.
Quite so. It is true that most compilers allow you to opt in to reproducible output, but it is generally not the default [iirc, gc is the only mainstream compiler that makes always reproducible output a design requirement]. There are advantages to not have to worry about reproducible output, like performance gains in allowing threads to run in any order. Of course, LLMs equally enable you to opt into reproducible output (temperature=0). Implementation details are immaterial. The human experience is always the same: Input one language, out comes another language.
> So the really interesting bit to me is that this level of argument is even made.
Yeah, it stems from a lack of base understanding of the technology. The "a compiler is deterministic but AI isn't" is always at the heart of the argument, but it isn't even true. Implementation details are immaterial, but when you don't understand the implementation it stands to reason that one would want to focus on that facet in order to learn more about it. Hey, that’s what discussion is for: to learn.
There’s nothing wild about that prescription.
Is that true anymore when anyone can vibe code? Seems to me that quality, correctness, and performance will be huge differentiators in a see of vibe coded slop.
This is just splitting hairs. And temperature = 0 is not comparable either, because the we will not keep models trained on data of a specific point in time around forever.
> Implementation details are immaterial, but when you don't understand the implementation it stands to reason that one would want to focus on that facet in order to learn more about it.
Not splitting hairs is not the same as not knowing what a hair is.
> Hey, that’s what discussion is for: to learn.
You told me nothing I didn't know or didn't expect, because it gets brought up every time like clock work. It seems to be be part and parcel of not seeing or not wanting to see the woods for all the trees.
Bacterial sinus infections are absurdly over treated [1]. It can take weeks to recover from a viral infection.
Not saying that you are a wuss, but my guess is, the doctor probably over-indexed on the throat pain, assumed you were a bit of wuss for calling the doctor for a sore throat, and decided to give you something to make you feel better quickly.
Either way unless you were 65+ a short course of oral steroids is very safe (and even then it’s only very mildly dangerous).
1.https://www.aafp.org/afp/2020/0615/p758
“Without antibiotics, rhinosinusitis resolved in 46% of patients after one week and in 64% of patients after 14 days.
Antibiotics can shorten time to resolution but in only five to 11 more people per 100 compared with placebo or no treatment.
Despite this, approximately 86% of U.S. ambulatory visits for acute rhinosinusitis result in oral antibiotic prescriptions.1 In Europe, antibiotic prescription rates for acute rhinosinusitis in primary care range from 72% to 92%”
Equally, if you replace one compiler with another, it is almost certain that the output will not remain stable, even when you have enabled reproducible builds. You are not going to find a difference on the outside, even if implementation may differ under the hood.
> You told me nothing I didn't know or didn't expect
What's in it for you, then? I have been able to learn about your character, which is quite fascinating. Technology is pretty boring, which is why nobody else was talking about a technology implementation to begin with, but HN accounts are quite interesting to learn about. Seems to me like a complete waste of time if you cannot learn anything from it.
I think we’re done here man.
For me this exchange kinda was, but I don't see what I could possibly do about that.
I guess way to argue with yourself? That’s not what I said.
> Not saying that you are a wuss
Then I provided evidence to demonstrate that it’s more likely than not that the doctor prescribed you antibiotics unnecessarily.
So neither of the straw man statements you created were accurate.
I had the honor of giving a keynote at the International Conference on Machine Learning in Seoul last week titled “What will be left for us to work on?” I addressed the widespread anxiety about how we should adapt as AI capabilities increase. I was thrilled by the talk’s reception, so I have made my slides available here, annotated with a lightly edited transcript. You can also view them below right here on this page, but the online version has animations, clickable links, and a much nicer experience overall.
I made three arguments. First, the AI as Normal Technology framework is a correct and useful as a way to think about AI’s impacts, unless and until there is some future discontinuity such as through recursive self-improvement. Second, even though we should take recursive self-improvement seriously, there is no milestone that companies might achieve in the lab that will suddenly put us all out of work. Third and finally, jobs of the future will be radically different, and a lot of adaptation will be needed. I shared my thinking about what this might look like and ended with a vision of human/AI “co-superintelligence”.
Now is a time of great excitement in AI, but it’s also a time of great anxiety in the AI community. I want to address that anxiety head on. How do we prepare for a future where AI will become capable of doing more and more of the work that we do today?
I lead a team at Princeton University trying to advance the science of AI agent evaluation. We try to go beyond the usual claims of “Look, capability is going up on benchmarks!” Those claims tend to be misinterpreted by the broader public as implying that agents are soon about to take all our jobs.
Maybe that will happen. But in our work we try to understand the factors beyond capability that matter for real-world deployment, and bring that understanding into evaluations.
The work that I’m better known for is the essay I co-authored with Sayash Kapoor called AI as Normal Technology. It’s a way to think about the medium-term future of AI and how to adapt to it — and in turn how to adapt it to the needs of society and the economy.
So we’ve been going around writing these essays about how lawyers should adapt, or maybe how journalists should adapt. But perhaps ironically, the question of how to adapt has been hitting our community first. Whether it’s software engineering or AI research itself, AI capabilities in these areas are of course advancing very rapidly.
Our response to this moment matters beyond this community. The whole world is watching. If we simply roll over and accept that a lot of our work will be done by AI in the future, instead of setting clear boundaries, I think it will lead to an even stronger political backlash against AI than what we are seeing today. So I think this question is not just for us but for the whole world.
From the beginning of AI, historically there have been these two battling narratives. In the past, the distinction was academic and philosophical, but now it has become an acutely practical question. Each one of us has to decide which camp we’re in, or where on this spectrum we’re in, because the practical consequences of believing in one versus the other are very, very different.
If you think this is a technology which in a few years is going to be able to replace everything we do today, then perhaps the correct response is to build wealth as quickly as possible before our skills become irrelevant. And this is the path that many have chosen in Silicon Valley. You may have heard of the “permanent underclass” meme.
On the other hand, if you believe, as I do, that this is a technology that will greatly amplify our potential, then now is the best time to build skills — especially the skills that are going to be complementary to what AI is doing and is going to be able to do — as well as to build all the things around it, such as agency and taste and judgment.
If you choose the first path, and it turns out that AI actually ends up being an amplifying technology as opposed to a replacing technology, then I would argue that over the next few years you’ve perhaps lost the best time in history to build these skills that will give us superpowers. That’s why we all need to think about this question, even if we won’t all land in the same place.
AI as Normal Technology is the intellectual framework for my talk today. When we say AI is normal, we don’t mean that it’s just like a hammer or a toothbrush, some kind of mundane technology.
We acknowledge prominently in the essay that this is a transformative technology on the scale of the industrial revolution. We’re not AI skeptics.
This is not a slogan. It’s a framework — sort of a causal model how AI capabilities impact the economy and society. It’s a 15,000-word essay and we’re turning it into a book. And I mention that because people often hear the word normal and they assume they know what we mean, but that leads to misunderstandings.
First, I’ll argue that this framework is correct and useful as a way to think about AI’s impacts, unless and until there is some future discontinuity — such as through recursive self-improvement — that leads to future impacts looking very different from past impacts.
Second, I will talk about why, even though we should take recursive self-improvement seriously, I’m not particularly losing sleep over it.
Third and finally, I want to be clear that I’m not saying that jobs of the future will be just like jobs of the present. A lot of adaptation will be needed. So I want to give some preliminary thinking on how I think our roles are going to change and how we can best adapt to the changes.
Powerful technologies of the past, like electricity, have been thoroughly studied, and we have good frameworks to understand how technological progress leads to economic impacts.
Invention: discovering the principles of electromagnetism, AC versus DC, etc.
Innovation: People don’t use “electricity” directly. We use electrical appliances. Those had to be invented, so that is a kind of downstream innovation — that’s the second phase of the framework.
Diffusion: This refers to the gradual process by which people start adopting innovations.
In our essay we apply this framework to AI and flesh it out into a four-part framework.
Here’s the basic picture, illustrated with software engineering as an example.
Methods/capabilities: Models are rapidly improving.
Products/applications: We don’t use LLMs directly. The reason they’ve been so influential in all of our work is because of coding agents. These are products that take those latent capabilities and turn them into something useful and usable for workers.
Early adoption: At first people were mostly trying vibe coding, and now we know that that’s not really the best way to develop production software — so now we have more sophisticated ways of doing agentic engineering.
Adaptation: (or structural transformation) — the fourth and slowest phase. Much of my talk today is going to be about that. I claim that this stage takes decades. It has not really started yet, even in a field like software engineering, which is a relative early adopter of coding agents.
We don’t know what the adaptation phase will look like — we can only speculate. Permit me to speculate for a minute. If it’s going to be the case that coding agents are going to be able to create ten-million-line code bases in the future that are not full of bugs and security vulnerabilities, then it won’t make a lot of sense for us to create one piece of software that billions of people should use. It’ll make a lot more sense for software to be tailored to the needs of each individual or team. And that’s what I mean by extreme personalization.
It’s not merely a technological change — that’s also a change for the industry. For instance: do we even need software companies anymore? Maybe software development will massively shift in-house, into the companies and teams that are actually using the software. Again, this is speculation, but the point is that it is this kind of organizational change, human change — that’s very slow, that takes decades — that will allow us to take advantage of the full potential of AI, whether it’s in software engineering or in any other field. So that’s one of the central insights of the essay. When we look at past technologies, this kind of change tends to be very slow.
Before electricity, factories used to look like the picture on the left. A massive steam engine generated power and it was moved throughout the factory by mechanical gears and belts. So when electricity came along, factory owners tried to replace those steam boilers with electric generators. They thought it would be much more efficient. But this idea of a drop-in replacement did not work. We keep hearing that term in the context of AI agents today — that they will be drop-in replacements for human workers. That did not work in the case of electricity.
What actually worked, and what took 40 years to develop, is to recognize that electricity is a very different technology. It’s portable, so you can move the power to wherever you need it. That lets you reorganize the entire layout of the factory around the logic of the assembly line. And that required changing the way that workers are trained, hired, and fired, new labor laws, and so forth. So that’s the kind of organizational adaptation that it took in order to reap the benefits of electricity in factories.
Our claim is that this is the kind of process that we will go through for AI. A couple of decades from now, we will have fundamentally reorganized work. We don’t know what that’s going to look like, and that is the challenge in front of all of us. And that’s not just a job for the AI companies to do, much like it wasn’t the job of the electric utility to figure out how factories should be reorganized. In our view, this is the slowest of the four stages through which AI leads to economic impacts. Today, this process has not really gotten started.
Why is there a huge gap between what people in various occupations could be using AI for and what they’re actually using it for? One reason could be that people are slow to adopt technology, and that’s certainly part of our framework.
But we wondered if maybe the people who are deploying AI and are not having much success at it know something about the practical limitations of AI that the AI industry doesn’t. Let’s have a bit more humility about the relationship between capabilities and deployment.
Given that the #1 concern people cite is reliability, we wanted to try to measure whether reliability, distinct from capability, is a barrier to the practical usefulness of AI agents.
We looked at 10-12 reliability metrics and clustered them into four dimensions.
Consistency: Suppose we hear that an AI agent has a 70% accuracy. Does this mean it works on 70% of the tasks, but on the ones that it does, it does so every time? That’s great for deployment — you can deploy it on that subset of the tasks. Or does it mean that on any given task it might unpredictably fail with a 30% probability? That’s pretty useless from a deployment perspective. Perhaps shockingly, none of the agent benchmarks that we looked into make a distinction between these two. Both of these are represented as 70% accuracy.
Robustness: We looked at robustness: what happens when the environment changes a little bit?
Calibration: Can the agent look back at its transcript and tell if it performed the task correctly?
Operational safety: When it does fail, is it recoverable, or is it something like deleting the production database?
For a human worker, if we think of someone as being competent at a job, it’s all of these things, not just accuracy. But it turns out we were measuring agents only on accuracy.
We measured capability and reliability using two complementary benchmarks, for models from these 3 frontier AI companies that were released over the last 24 months or so.
This is a period during which accuracy or capability shot up dramatically (left).
But reliability (right) only increased by five or ten percentage points.
There are a bunch of implications but let me highlight one. Right now the industry is not treating automation and collaboration agents differently. If you run an agent in headless mode, I guess that becomes an automation agent. That is not a good way to look at it, because properties like reliability that are very important for automation agents can actually be a hindrance for a collaboration agent that you might be using to improve your creative writing or something like that. For that kind of agent, you don’t want it to behave like a robot that does the same thing every time. You want it to be creative and explore different possibilities and be unpredictable.
Scaffolds and even the post-training of models should be different based on whether it is supposed to be driving a collaboration agent or an automation agent.
My hope is that reliability will continue to improve and automation will become easier over time. But for now I think collaboration agents will continue to be much more successful. Many companies that hastily rushed to automate business processes using agents are recognizing the limits and costs, and even legal liabilities — like when an agent deletes production data.
For the time being, you can have only two out of these three properties in agents: general-purpose (a language model based agent that can be instructed to do different tasks rather than purpose-built for a task like traditional software); deployed in high-stakes scenarios, and automated.
This is one of the reasons that I think that for now, AI remains much more of a collaboration technology than a technology that automates workers away.
Let’s take software engineering as a case study. It’s a good leading indicator because coding agents have been particularly rapidly adopted.
You might think: okay, we can’t completely automate away software engineering. But if agents make software engineers ten times more productive, then we need ten times fewer software engineers. Isn’t that an obvious consequence?
Well, that is completely contradicted by the data. We looked at this in a follow-up essay. In every case we looked at, the company was under financial pressure, and it turns out to be more convenient to blame AI for the layoffs instead.
Why is AI not replacing software engineers so far? We’ve known for a while — this is a paper from 2019 — that writing code is not really the bottleneck.
Over the last year, as software engineers started adopting coding agents and started to recognize that it doesn’t seem to be cutting down on the amount of their work, there have been many blog posts rediscovering the fact that writing code is not a bottleneck. Here is a small sample.
So what actually is the bottleneck?
This framework is our answer to the question.
The decide layer: understanding customer requirements, developing the specification, planning, etc. That is not getting compressed by AI.
The execute layer: the actual coding and debugging. This is getting compressed, but it was only maybe one-third of the work to begin with.
The deliver layer: Understanding your code deeply enough to be accountable for what you release; carrying out integration into customer systems, maintenance, testing, etc. This layer is not getting compressed either.
In fact, the first and third layers are arguably expanding as AI compresses the middle layer — and I’ll come back to that point.
I think it is already the case in software engineering, and will increasingly be the case in many professions, that we can think of knowledge workers as similar to a crane operator or a forklift operator. The machine greatly amplifies the human potential to do physical work — it’s doing all the heavy lifting, but the person still remains in control. And I think this is what is happening with cognitive work. Machines are going to increasingly do the cognitive heavy lifting, but the person still remains in control.
The entire job gets reconceptualized as being about operating the machine, understanding the machine, and controlling the machine, as opposed to doing the cognitive work ourselves.
All this might seem like a dramatic change, but in a sense it is only a continuation of what has been happening over and over and over in software engineering. Starting from the days of machine code, we’ve had many waves of technology, each of which gives us nearly an order of magnitude increase in productivity. And far from decreasing the demand for software engineers, during the time that we’ve climbed up this ladder, the amount of software engineering employment has increased by a factor of something like 10,000. And that’s simply because the amount of code that there is to write has gone up by orders and orders of magnitude.
Economists have found this repeatedly. You might have heard the term Jevons’ paradox; I like the term “lump-of-labor fallacy”.
ATMs made it economically feasible for banks to open lots of regional branches. And those branches still needed human tellers to handle the things ATMs couldn’t. So, paradoxically, employment actually grew.
Geoff Hinton made the famous prediction a decade ago that radiologists would be basically extinct in five years. But it turns out radiology employment has in fact grown. And it’s not because radiologists are rejecting AI. They are actually enthusiastically adopting it. One reason for job growth is that when a task gets faster and cheaper to perform, there’s more demand for it.
We’ve written a paper looking at how lawyers should adapt. There’s a lot in there, but one simple point is that AI has made it a lot easier to file lawsuits. And this means more work for lawyers. We might be unhappy about this if we don’t like living in a litigious society, but from the point of view of employment for lawyers, this is great news.
Translation is a more extreme example which kind of blows my mind. AI was almost at human parity nearly a decade ago. Yet the employment of human translators has remained more or less stable, and is projected to remain stable over the next decade. There are many reasons for that, but one reason is that there’s really no ceiling to the amount of things you can translate and the number of different languages you can translate them into.
I’ve argued that so far, our framework is very consistent with the evidence. Now let’s talk about the possibility that everything I’ve said so far will be obviated because some kind of takeoff or singularity will be achieved, and at that point there will be nothing left for us to do.
Many companies have said, including these two, that they are racing towards recursive self-improvement. And they are serious companies — I take them seriously.
Suppose, some time in the next year, recursive self-improvement is achieved. What are the consequences? The slide shows the usual view in the AI community. Note that AGI a famously slippery term. There are two main groups of definitions. One is that AGI will be humanlike in a range of cognitive dimensions, and the other is that it will be capable of performing a range of economically valuable tasks.
My claim is that this view — that treats AGI and ASI as near-automatic consequences of RSI — is a little bit silly. These are four different dimensions of progress. None of these dimensions implies any of the others.
I’ll talk about the details on the next few slides, but one small bit of intuition for now. One of the things that’s associated with superintelligence is that it will cure cancer and other diseas. But we know that the hard part of developing medical treatments is clinical trials that often require thousands of people and 10-15 years. That’s an example of the fact that the bottlenecks to superintelligence are external. It is not something that you can solve in the lab through purely computational processes. So the idea that recursive self-improvement will automatically lead to superintelligence, like many commentators assume — clearly there is something missing in terms of the causal chain.
Early on in the history of AI, these were all distant goals, so it was okay that we didn’t clearly distinguish between different dimensions of progress. But now it’s becoming a real problem — it’s leading to confused discourse.
As an analogy, suppose we are early explorers and we hope to go to Hawaii one day. It’s okay that we have one single term for this group of islands. But as our ship gets closer, we’d better be able to talk about these islands with different words. Otherwise, we’re going to confuse ourselves about where we are and where we’re going.
Helen Toner, among others, has made similar points.
Now let’s discuss each of the four dimensions in detail, starting with Recursive Self-Improvement.
Suppose a company claims they have built RSI — they have built an AI system which built its own successor. What does that actually mean?
On the one hand, maybe it means an LLM spit out a whole bunch of ideas for how to tweak the architecture or the data pipeline or whatever else, and they automatically tested it and kept the improvements that worked. That’s just glorified hyperparameter search. We’ve had systems like AutoML for a long time — I looked it up, and even Schmidhuber has published about this a long time ago. We’re not calling that superintelligence. So that’s one end of the spectrum.
It’s very different from the other extreme, where you can imagine the company actually manages to replace the creativity and intelligence of the human AI researchers — not just one researcher, and not just researchers working at the company, but the entire worldwide community of hundreds of thousands of people whose innovations are all going into improving AI systems. So when people are talking about RSI, it’s not clear which of these they mean.
The reason for my prediction is simple: AI is still in a state where it is much better at verifiable tasks than non-verifiable tasks. Some dimensions of AI performance, like speed and efficiency, are verifiable, and it is possible that they could be greatly improved through near-term RSI.
However, creativity, among other dimensions of intelligence, is the epitome of an unverifiable task.
We don’t even know clearly how to test AI creativity. We don’t understand human creativity well enough from a cognitive science and neuroscience perspective to be able to have clear tests for it.
Let’s go a bit deeper into AI creativity as an illustration of the barriers to humanlike AI.
This quip from podcaster Dwarkesh Patel is a good illustration of how AI creativity lags human creativity. There might be a few counterexamples, especially in narrow domains like Erdos problems. But think about famous “Eureka moments” — finding surprising connections between seemingly unrelated fields or problems — that we associate with great feats of human scientific invention and creativity. LLMs seem to be nowhere close to that.
To understand why, I spent a good amount of time immersed in the cognitive science literature. I won’t go into the details, but seems like there are two literatures looking at this in slightly different ways that haven’t really been talking to each other, for reasons I don’t fully understand. Based on what I’ve been reading, I want to present some hypotheses about AI creativity.
Representation quality is absolutely fundamental to cognition. Think about deep learning — what a big leap that was in the quality of representations behind perception. So I do think AI has more or less caught up in representation quality when it comes to perception, but it has not caught up when it comes to representations that we use as humans for creativity and reasoning.
François Chollet in particular has talked about the fact that human representations that underpin our creativity seem to exhibit a kind of extreme compositionality, where everything is built up from a few “atoms of meaning”. Many seeming limitations of human working memory and information processing actually turn out to be strengths, because they force us to come up with these extremely efficient representations.
Because LLMs are much better than humans in certain dimensions — like memorizing and retrieving stored patterns that might enable creativity — we’re not able to elicit these important real limitations in studies that try to compare human and LLM creativity.
Furthermore, we humans can do something beautiful when we’re creatively thinking about a problem: we improve our representations related to the problem in real time — at inference time, if you will. That enables us to sleep on it, improve our representations, and come back the next day much more efficient at solving that particular problem. I’m sure we’ve all experienced this. But it is something that today’s AI systems are not able to do.
I assumed naively that the continual learning people would be all over this, but when I started looking into that literature, it appears to me that that’s not the case. Continual learning is mostly focused on preventing catastrophic forgetting, as opposed to improving things over time. And furthermore, it is more focused on facts, skills, etc., as opposed to the quality of the underlying representations.
So putting all this together, and considering the fact that creativity is only one of many barriers to humanlike AI, I do think there is still a very long way to go. That said, I want to express some humility here. These are all just hypotheses. I think we need empirical verification.
Speaking of empirical measurement, we have an ongoing project testing AI’s ability to do humanlike AI research.
We give AI agents a budget of a few thousand dollars and a machine learning research problem — a problem that researchers have already worked on and written a paper about, but not yet made public on arXiv. This combines two properties: it gives us a team of human judges who’ve thought deeply about this problem for months and can therefore judge the AI output, but their thinking is not yet online, so it prevents AI cheating or contamination.
Many others have done auto-research experiments, but we’re trying to do some things differently — in particular, picking somewhat open-ended problems, so that we can actually test the ability of AI agents at exercising judgment and creativity. We hope to release detailed findings very soon.
This builds on a foundation that we call open-world evaluation. We’ve completed one open-world evaluation of getting agents to build and upload an app to the Apple App Store — that’s not about recursive self-improvement, but it tests a different kind of thing. We have assembled a great team, with people from many universities, a couple of companies, as well as the UK AI Security Institute.
And by the way, we are hiring a researcher to lead some of our future open-world evaluations. If you’re interested, go check out our website.
Now let’s turn to the third of the four dimensions.
Some people say that AGI is already here. Well, this is one way to interpret that claim, and I happen to agree with it.
This might seem surprising given the skepticism of rapid economic impacts that I expressed in Part 1. But this is actually not only consistent, but in fact a restatement of Part 1. My point there was that the barriers are downstream, and therefore model improvements won’t rapidly change the economy. But by the very same token, because the barriers are downstream, even without model improvements, those barriers are gradually going to get addressed. Yes, it might take a couple of decades, but we will definitely get there.
There are many barriers: reliability (already discussed); integrating AI models into various existing systems; tacit knowledge from domains like medicine or law or various other professions that need to be made available to the models; regulation that often straight up prohibits today’s AI systems from being used in productive ways — in many cases there are good reasons for that regulation, but it will need to be modified in order to be able to enable adoption.
The key point is that these are not things that will be solved in the lab. These are not things that are going to be addressed by the next model release on Tuesday. They will be addressed gradually, through gradual adoption, over a period of decades.
From a geopolitics or strategic perspective, some people advocate a race to AGI, arguing that, similar to a Manhattan Project, the country that gets to some capability milestone first is the one that is going to reap economic rewards. I strongly disagree.
There is no particular capability milestone that will unlock all of this economic potential. The economic potential is already there. It really depends on all these downstream actions that we take. It is not gated by capability.
Okay. Now let me talk about the last of these dimensions, namely superintelligence.
I gave the example of medical trials earlier. Another example: Do you think that we will get superintelligence that can predict the weather precisely a year into the future? We know that that is basically a mathematical impossibility because of the theory of chaos in nonlinear dynamical systems. Our position is that a surprising number of tasks are similar to weather prediction, in that there are inherent limits and we are pretty much already at those limits.
Even for tasks where we’re not at that limit, the idea that there is going to be AI superintelligence that’s going to obviate humans relies on a fundamental misunderstanding of human intelligence. I claim that in the vast majority of tasks, our performance is not limited by our biology — it is rather limited by our learning and tools.
If you imagine someone from the ancient past time traveling to our world, we’re superintelligent compared to that person. And it’s not because our biology is better, but because they don’t have the benefit of all the learning that we’ve been through and all the tools, especially digital tools, that we’re able to use in order to be productive at whatever it is that we do.
And AI is one such tool. So what that means in our framework is that improvements in AI are actually improving human intelligence, not just AI intelligence. So we have a race between the performance of AI-augmented humans on the one hand and the performance of AI systems acting alone on the other hand. I think we can ensure that we are the superintelligences of the future and we win that race, as opposed to allowing AI systems to act alone in ways that would threaten our future and control.
Many people are pessimistic about this. They bring up thought experiments such as AI running and owning companies in the future. Their view is that if the AI system is not aligned, terrible things might happen — the AI could turn into a paperclip maximizer or do other catastrophic stuff. Our perspective is very different. If you’re imagining a future where AI is actually owning and running companies and hiring and firing people, that’s already dystopian. It doesn’t matter if that AI is aligned or not. The consequences for human dignity. democratic governance, etc. are already catastrophic.
So if your view on safety is to treat this as an inevitable future and just to hope for “alignment” to solve that problem — pardon me for being a little bit blunt here — it feels to me that you’re against safety, not for safety. I think it will take a lot of hard work to ensure that we don’t irresponsibly deploy AI systems in this kind of fashion. But I do think we can get there. And that’s the reason I’ve spent a lot of my career advising policymakers. We’re going to need policy and we’re going to need politics, and it’s going to be hard, but let’s not give up that fight before it even starts.
For example, RSI is achievable in the lab, but it won’t immediately put people out of work. On the other hand, economically transformative AI is going to happen, but it’s not because of the next model release — it’s because of things that will gradually happen over the next couple of decades.
Connecting back to the title of this talk: there’s no world in which something that an AI company will decide to do in a lab will put us all out of work. Yes, there are risks to be worried about. Yes, things are going to change. But we have agency over how AI gets deployed, and that process will unfold over decades. Again, this is not guaranteed, but this is the future that I want to work towards, and this is the reason why I’m cautiously optimistic.
Let me take the last fifteen minutes to talk about the flip side. I do think a lot of things are going to change. What are some of them?
Technical skills tend to be verifiable tasks, and AI will continue to get better at them.
Over twenty years ago, there started to be a stark difference in the labor demand for programming jobs versus software engineering jobs. Programming jobs are conceived narrowly around the technical skills of coding and debugging. Software engineering jobs are responsible for all three layers of the decide, execute, deliver sandwich that I talked about — figuring out what even needs to be built, understanding customers, that sort of thing. It requires domain knowledge, judgment, and more.
I predict that this will happen in more and more fields over time.
A recurring pattern I’ve observed — effort shifts from building systems to evaluating systems. As I’ve mentioned, I lead a team working on AI agent evaluation. LLMs and agents are general purpose. So each time capability goes up, it creates demand for evaluation in a legal setting or a journalistic setting or whatever other setting, and that’s not work that is scalable.
Not only is AI agent evaluation resistant to automation — it has become sufficiently specialized that the set of people and teams working on evaluation is starting to diverge from the set of people and teams building and pushing the state of the art in AI agents. This new community is developing a new set of best practices around what it means to rigorously evaluate agents, and we have a forthcoming paper that is going to look at that in some detail.
Here’s an a metaphor to help explain the shift I see in our community. Imagine that in the past most boats were rowboats, and the work of the humans was in physically moving the boat. There was no separate specialized role around steering the boat: when you’re rowing the boat, you’re also figuring out which way to row the boat.
But what happened when the physical work of moving of the ship could be delegated to the engines? The human jobs didn’t go away. In fact, they became much more specialized. Modern ships have very complicated control panels, and they might have dozens of different specialized roles that are focused on where the ship should go and how it should get there.
I would argue that we’re seeing a similar shift in AI/ML. In the past, most of our work was on building — we didn’t need separate roles for evaluation. That has changed now. We’re still in the early stages of this process, but I think over time more and more of the building, because it’s a verifiable task, will be able to be done by AI, whereas it’s the evaluation — figuring out where we should go as a community, figuring out what are the desirable properties in AI systems — that is very resistant to automation. So a greater fraction of the community’s attention will have to focus on evaluation over time, compared to where it is today.
In rowing, physical strength was valorized, and today we might be sad about the fact that sailors don’t need to be strong and it’s all a bunch of nerds. Similarly, today in the AI community, there’s still great value in deep technical understanding of models and systems, and that is considered the coolest thing; the most important thing; that’s the thing that commands a lot of value. But I think in the future that will become less important than exercising judgment and all of these fuzzy things that have a lower status in the AI community today. That’s a mental shift that we’re maybe not quite prepared for. Many of us will be sad about it, and that’s okay.
One practical consequence of this shift: consider a conference like this one. What fraction of the conference should be dedicated to evaluation papers?
I don’t know, but I think maybe a lot more than it is today. Maybe around half — I’m just throwing a number out there — which is orders of magnitude more than it is today. At the very least we need a dedicated track. We don’t have a dedicated track. NeurIPS does have one, and it has been growing in popularity over time, and I think that is a good thing.
And I would go further and argue that thoughtful evaluation of AI systems is a form of alignment. It’s not aligning the AI system itself, but it’s aligning the community as a whole — thinking about where we’re currently going and where we want to go, and trying to align those two to each other. And without enough emphasis on evaluation, my worry is that the community — going back to the ship metaphor — will behave like a rudderless ship: very powerful, but where we don’t collectively have control over what kind of AI we want to develop.
I had a small role in this cool position paper. A one-sentence version of the argument: we need to move beyond benchmarks being the be-all and end-all of what AI research is considered valuable.
What benchmarks give us is efficiency. We don’t have to spend years peer-reviewing papers — if it beats the state of the art, we know it’s probably a good paper. But unfortunately, it has narrowed the collective vision of the community. We’re searching under streetlights — we are searching the kinds of things that are easy to search in a benchmark-driven evaluation regime. We need to move beyond that. That means evaluating how good a paper is will become much more costly. Sadly, that is a cost that we will have to pay. Right now the community is trying to rush in the opposite direction.
There is a big temptation towards automating peer review. And with all due respect, I think this is a trap. If we automate peer review, we essentially give up control over the direction of progress to AI systems themselves. That seems like a fundamental misallocation of effort.
If we automate the mundane parts, human time can be freed up to think more deeply about the parts that require judgment.
Generalizing from AI research to scientific research as a whole, we have an essay that argues that visions of “automating science” rely on a fundamental misunderstanding — as if the point of science is mere problem solving, and that if we can use AI to get directly from the problem to the solution, we will have automated and accelerated science.
In our view, human understanding is not some friction to be automated away. It is rather essential. It is central to the very purposes behind why we do science. And if we lose human understanding, we lose all of these things that flow from it.
So my prediction is that if we are going to have increased use of AI agents in science, there will be new tools and new roles that are specialized towards looking at those AI solutions and backing out the human understanding from those solutions. Because it is essential to preserve human understanding.
Many people in the corporate world are saying that evals are the new IP. Without going into the details, it is very much the same phenomenon playing out once again — effort shifting from purely building to a mix of building and evaluation of systems.
(Note: in this post I explain the idea of cross-functional eval teams to keep companies from fooling themselves.)
If there’s only one thing you take away from this talk, it should probably be this slide. What I’ve shown throughout the talk is that in many different areas, because the verifiable tasks can be handled by AI, some or much of the effort shifts from building to evaluation.
Effort shifts from “rowing the boat” to “steering the ship, navigating the ship, and figuring out where we even want to go”.
I think this is a powerful metaphor that will allow us to predict how human roles will shift over the next decade or two.
In the last few minutes, let me give you some personal reflections on how I’ve been struggling through this challenge — the fact that AI capabilities are rapidly improving — in my own research workflows, as I’m sure many of you are. I think everyone should choose their own path, but hopefully there are some interesting ideas here for you to consider.
By floor I mean what AI can do on its own. By ceiling I mean what AI allows us to do by augmenting our capabilities — the ability to take on new and ambitious projects that were not possible before AI. But the ceiling is not going up automatically. It is only if we work on pushing the ceiling up.
I find that I’m spending something like 10 hours per week just learning and experimenting with new workflows, as well as learning new topics. So one way to think about it is that AI enables big productivity improvements, and I’ve been trying to take that time saved and “reinvest” it into long-term growth and picking up new and complementary skills.
I’ve learned that if I don’t feel exhausted at the end of the day, I’ve done something wrong. I’ve offloaded too much to AI. I’ve sacrificed too much of my long-term growth in the pursuit of short-term productivity.
Growth and productivity are two legs of a three-legged stool that we need to learn how to balance. And the third leg of that stool is staying in control.
Here are two heuristics that I’ve tried to use in order to do that. The first is resisting the black-box temptation. Companies want us to use agents as black boxes — just prompt it and it will go off and do its thing. I think that’s a trap. I think that’s very dangerous, and it will lead us to gradually give up control over time.
Second, and related, is what I call the dependence spiral. It’s very tempting to use AI for tasks that I am myself not yet an expert at, because learning new things is of course very hard. But that leads to my losing whatever little skills I have in that task over. It’s much better in the long run if I first put in the time to master the task myself before I use AI to augment my productivity.
None of this is easy, but if we get this right, the vision is tantalizing and empowering.
Computers have often been called bicycles for the mind. I think AI can be more than that. I think AI can be a crane for the mind, if you will indulge my metaphor, in the sense that it can amplify our potential to previously unimaginable heights. Getting there can seem daunting. It has an incredible learning curve. I feel like I’m on a treadmill all the time, but I’m very excited about it. I think it’s a fun challenge, and I think it’s worth fighting this fight. Compared to five years ago, in a way, I feel ... maybe superintelligent is not the right word, but I feel like I have superpowers, given the extent to which AI allows me to take on new ambitious things that were not possible before, and push myself harder than was possible before.
And I think we can ensure that this remains the case even as AI capabilities advance, for the foreseeable future. It might be that in some distant future this becomes impossible to do, but it is very premature to give up the fight now. I certainly plan to continue this fight, and I hope to work towards this vision of co-superintelligence, and I hope you will join me. That is my closing thought.