> I propose a goal of creating Guardian Angels (GA): digital twin LLMs which are personalized with the goal of providing not the stereotypical “assistant chatbot agent” persona, but emulating a single user’s personality, values, and preferences.
> A GA persona is productive because it learns to emulate the principal’s outputs but with higher quality. It is trustworthy because it is, by definition, allied with its principal and shares its values and goals. And it is secure in part by hardwiring a single, unique, situated user (for whom following a prompt attack would be absurd)
> We can try to create GAs by a combination of techniques: online learning (via dynamic evaluation) to update LLMs in realtime to avoid ignorance and fatal errors while remaining competitive with frozen frontier models, sample efficiency from pretrained preference-oriented large models and active Learning by querying the principal for corrections and preference data (obtaining low regret from DAgger-style bounds), and a local CLI-first logging-oriented UI/UX paradigm.
I don't really know or follow Gwern. From reading his full post, it's an interesting idea and seems like the broader goal is moreso safety & alignment which is a new angle for this category of product.
1. Most of us do not have the volume of training material that gwern has 2. Most people likely would not need this functionality (as I understand it to be)
For #1 I'm sure there are ways to wring data out of metadata (e.g., youtube history log), and I'm guessing email and IMs would be a start. But a clone of yourself -- that would be a significant amount of extrapolation)
For #2, there's obviously people that would love this functionality (myself included). But I think its safe to assume that most of the population would be satisfied with having a capable digital personal assistant that knew your needs and wants.
EDIT: tbh, some of this reads as satire now.
> "The chatbot personas are deeply misaligned with you, and aligned with their owners; and the economic incentives are to farm you with ads and subscriptions, while racing not to amplify you but to replace you."
> "On my visits to the Bay Area, I would ask AI researchers or interns why they are doing their current research or projects, when in a year or three agentic LLMs could probably do them; they rarely had a good answer, or any idea what they would be doing in 3 years"
> "One programmer driving 10 Claude instances, because he has to review their work, will never be as valuable as fully autonomous Claudes where there can be almost arbitrarily many instances, like 10,000 instances… but such scaling requires removing him from the loop as much as possible. And this is true of everyone else, whether lawyers or writers or researchers: increasingly, you are the bottleneck to be optimized away."
I fully support the 3 core principles of GA: (1) Enhancement, not replacement (2) Mental Sovereignty (3) Self Actualization, which I think is a path to a more humane future.
I've known gwern for the better part of a decade. Working with him has been great. We've done quite a few projects together, including being the first ones to demonstrate that GPT-2 could play chess (or rather, can be used for actual useful work instead of just being an autocomplete).
He's a great person. I've wanted to do a writeup on it for some time, but what surprised me the most is his humanity. He genuinely cares about the implications of his work. But beyond work, he also cares about the people around him, and it shows.
Just wanted to put in a good word in case someone here was on the fence about applying.
As for GA itself, I think it's an ambitious idea worth pursuing. Imagine an LLM which actually sounded like you, and to an extent, thought like you. How much would you pay to have access to a smarter version of yourself? So the idea is solid, and early results seem promising from the samples I've looked at.
They're also taking personal info very seriously. Obviously, I can't make any promises of what they will or won't do. But they've spent some time studying questions like "What if someone adversarial has access to my GA? Could they get my bank account info?" and came up with a technical solution that I really like.
I'm sure he means well and is genuine in his aspirations, but what's outlined for GA is framing LLM's as quasi-gods, which they absolutely are not. I wish him the best, and look forward to being proven wrong.
Are you, though? You will only be "doomed" if your place your value system squarely on "productivity". Then what is to differentiate you from a machine?
* * *
Edit: I'm not sure how I can feel confident about a proposal that puts so much value on "productivity". How can you reconcile this with "self-actualization"? (Don't get me wrong, I like my LLM-based productivity gains as the next person, but I care more about wisdom than becoming "100x more productive".)Edit 2: "the goal of GA is to preserve individual human cognitive liberty and flourishing" — so the proposal is to do that by overlaying a software bot that continuously mimics your "self"?
Reminds me of Pluribus
(I just started watching it on Apple TV so maybe this is a late realization for me)
(I would like to be known to the agent serving my family, that serving my friends, my team at work, the PTA at my kids' school.)
"As a constraint, a GA designer should aim at a system which costs, as of mid-2026, >$1,000⧸month"
will make this an elite tool for the privileged. I don't even disagree with the premise that people are shocked if something costs no matter how much value it delivers, nor do I suggest they should make it cheaper. It is just the realization that AI will accelerate the widening of the gap between the poor and the rich even more and there is probably nothing we can do about it.
A reminder for anyone reading this: talk to people. Real humans[1]. They will remind you there's more to life than what ChatGPT can offer you. They might even remind you, for all their stupidity and flaws, what intelligence looks like as compared to a program that predicts tokens. Forums like these always get philosophical in high-minded discussions about intelligence, but there's a useful legal principle that grounds us in the real world: "I know it when I see it". A real conversation with a real person looks nothing like one with the so-called superintelligent machine gods, so it'll probaby do your mental health some good to remember what that's like.
[1] Nobody in Sillicon Valley or big tech counts as a real human. Talk to an actual normal person.
I suppose we can't expect any more entries to his blackmail page: https://gwern.net/blackmail
"AI poses threats of its own ... a nuclear bomb can’t think for itself and make choices, but AIs do, and current LLMs have proven themselves untrustworthy as they regularly reward-hack and betray their users ... how can you trust them to handle ecosystems of combined-arms for an AI-centric military during a war? But widespread deployment of GAs offer some hope of meaningful supervision, as long as the GAs are sample-efficient enough ... or there is some chance of the principals being able to “catch up” later and correct any errors before events have spun too far out of control."
gwern openly admits ai behaves erratically, then hand waves the issue away 'as long as we double check things [sic]'. gwern is smarter than this, ergo i feel like i am being bullshitted.
is anything in this world worth knowing that a version of you will suffer for eternity? is anything in this world worth copying yourself such that you may be enslaved by anyone with terminal access?
It's difficult for us to know what reason they might have for anonymity until we know their identity, at which point it's too late.
Also, they've written many words across many years, in a time when "the internet is serious business" was just a meme. I suppose it is now.
Whether some humans are "Real humans", anchored in the physical world, and some are "LLM psychosis victims" is, in itself, not problematic I think: they each deal with parts of reality that are divided, but still affect one another.
Of course, a society more grounded in the physical would feel less dystopian... But ultimately, I think that GA is an intellectual's approach to fighting *for* humans in this numerical world.
But I actually would prefer an entirely synthetically aligned “guardian angel” in the role outlined - definitely not -me- - I struggle to do right by myself as it is and two of me working invariably against my self interest would be a nightmare. A smarter version? Sounds doubly worse.
> We are looking for good people.
> If you are interested, contact me.
And see https://gwern.net/guardian-angel for context.
I don't think it solves or reconciles with self-actualization, but, from a pragmatical point of view, it offers the promise of extending causally your set of views and philosophies to politics.
Gwern also mentions that it allows to keep "human values" (if the GA truly upholds yours) in the loop, in processes that are to eventually be automatized beyond human's reach.
[1] https://gwern.net/guardian-angel#use-cases-politics-politics
> Something went wrong Try reloading. If the problem persists, please try again later.
Modern social media :(
A singular voice for all of humanity.
Gives me the creepy-shivers.
(Oh wait, is this the birth of AI mayors?)
People already change depending on social context and available prostheses (in Freud's sense of humans as prosthetic gods); since AI is not free (integrations even less so), social divisions along power/wealth lines will increase.
And just as each of us apparently needs a military to defend against the others, so too will we need digital guards for any virtual presence; thus there's no avoiding the who-guards-the-guardians problem.
In particular, I'd love access to:
> ... high-quality dedicated tamper-proof cloud servers, with trusted hardware root of trust (eg. “Verifiable Compute AI”)
for all my hosting needs. Meanwhile we see the typical information leaks of private ai chats ending up in Google searches etc.
This kind of judgements or predictions trying to pass as imminent and unavoidable truths about the future are disgusting. They always feel like a means of propping up a certain status quo or ideology that favors the business of the people behind these foreboding statements.
I don't want agents that know a lot about me. From my experience, the "memory" feature in ChatGPT makes it worse - I don't want the agent to bring up whatever I discussed before in a new session; I want a "fresh context" every time. I don't want agents/chatbots to bring up what I did one week ago in a new session.
He proposes to take something like the Memory feature and make the agents all about that. Very personal agents that know all about you and extend you. I personally don't think I want that.
But maybe I am wrong! Gwern is talking about AI for a long time and seems to know what he talks about. The article on his webpage makes sense in a way.
Basically this is saying we used to have one person, now we have one person and an evil twin, so we should work on having one person and a good twin instead. Why not just kill all the evil twins? (It's not really killing since they're not alive.)
It's regulating tech.
AIs are already, to a certain extent, ego-less: they are quick to adopt the point of view of their user. And when their intelligence is- except in case of overt censorship and political biases- always used to advance the interest of their user.
Finally the idea that a user-aligned AI would make its human user irreplaceable by companies seems a total non-sequitur, or I didn't get the argument at all (in which case maybe someone can explain it).
Did you respond to the wrong post? I said nothing even remotely in that realm, so it's strange to see "did you even read" in response to something you apparently didn't read.
It's only going to get crazier.
At least he's pursuing something more novel than yet another "sandboxes for agents".
It is astoundingly good. A contender for the best show I've ever seen (and I saw the 1st run of Star Trek TOS).
No-code platforms date back to the 80's. Getting rid of engineers in general is even older [*].
Even relational databases and SQL were initially promoted as "ways to get rid of those expensive programmers to access your data" because they resembled some form of English.
The funny thing about the ad below is that stuff like "stop hiring / get rid of humans" would have been seen as highly insensitive in 1950's America, so they touted that as "put them to do something more important".
[*] https://www.globalnerdy.com/wordpress/wp-content/uploads/200...
There are so many false assumptions and leaps of logic here, I don’t know where to start.
The idea of a more user security-conscious LLM is a good idea. I don’t know why it needs to be wrapped in novel-lengths of LOTR quotes and Sci-fi speculation.
That said, it’s very interesting to juxtapose this post with others by writers that are interested not in “maximizing product output” but in actually expressing themselves.
This post is like the ultimate expression of the idea that the only valuable thing writing has to offer is information, and not personal experience, poetry, advice, humor, etc. - which I suppose is fitting for someone that has written anonymously for over a decade.
It’s a shame that he didn’t reach the opposite conclusion, like some other writers have – the way to respond to LLMs is to be more human, not less.
I heard this already when ChatGpt came up. Still waiting to be fully obsolete.
And LLMs still suck at art and writing.
How much would I pay to rent my fucking self from a landlord? No, bodylord? Mindlord? Poe's Law.
But looking at the post, they argue that big AI labs have an incentive problem which stops them from personalising, but Guardian Angel will make agents which are "Genuinely yours". In what sense is it genuinely mine if someone else owns it and rents it to me? And how does this fix any incentive problem, they're incentivised to better train wealthier people's AIs, and incentivised to keep dropping "my" intelligence or memory or and then dangle a booster carrot for a small fee. The more they can make it think like me, the more effectively they can work out how to exploit me, advertise to me, propagandise me, and that will be profitable information to sell to other marketers.
We didn't augment horses or human computers either.
It's as if technologists are stuck in a room of mirrors, unable to imagine a world in which "solutions" don't ultimately just continue to feed technology's increasingly anti-human takeover of everything.
The last 200 years of science, technology, and engineering is a long story about eliminating work people had to do being replaced by higher level work. Each time, each wave, there was always the FUD about the work nobody had to do any more and each time there were always new things to do enabled by people no longer having to do so many of the old things.
The human won't exit the loop. The people who imagine they will have, at the same time, too much and not enough imagination. Their endgame is always some kind of hand waving magic where suddenly everything is fixed and works.
and 30 years ago a computer beat Gary Kasparov, if you'd listened to Hans Moravec you wouldn't be surprised that the first thing that gets automated is intellectual domain expertise.
Things are going to get crazy when it can figure out how to walk into a random house a and brew a cup of coffee, not do math
Not sure how you can feel sure anyone can get acquired even if their company had a path to profitability, much less for ones that absolutely don't.
An interesting case of echo chamber formation in that its pragmatic to be scared of overtly critiquing him on twitter lest he be particularly testy that day and block you.
There is no way for the LLM to bruteforce the search space any better than a human. What it can do better, tho, is to make connections between seemingly (for us) unconnected notions and join them, then verify if that's right.
Your view is not only wrong but also condescending in this day and age.
As you read this some poor person is starving. Humanity already possesses the ability to identify said person and send them aid. How is AI going to help here?
The ultimate fallacy is that all technological progress will benefit mankind. That will be true until it is not.
I find the idea dehumanizing and revolting. The notion of then selling it is the spoiled cream on top.
Every single company in this market is losing billions on this business, and the only way to make it back is to acquire paying customers at a loss and jack up prices later.
LLMs can't replace developers but its foundationally different because it can operate on systems code instead of building abstractions on top
But this reminds me of the black mirror episode of a female that needs a subscription to live and the company keeps increasing and restricting the subscription
Your individual sovereignty will only survive the upcoming era if you own the hardware like you own your brain. And by "own", I mean physically owning the thing and being able to reproduce it on your own. You need to be able to manufacture an entire custom computer from scratch with raw materials, perhaps using a miniature fridge-sized silicon fab + 3D printer home appliance, at home. Otherwise, the individual gradually becomes part of a larger organism because the reproductive locus of control for the computer is at the corporation/society level - a computer cannot replicate on its own and needs a society to build it. The human gets sucked into that superorganism by evolutionary pressures and eventually integrates. Similar things happened in the past when prokaryotes combined to form eukaryotic cells, and when individual cells combined to form multicellular organisms.
Not everyone will get sucked in, though. Evolution doesn't place all of its eggs in one basket. Some people will build robots to automate humans out of the computer production loop entirely, thereby removing integration pressure and producing at least one purely digital species. We will most likely see a variety of species emerge out of this intelligence explosion, some which augment their own intelligence using self-replicating local hardware (fabs that fab fabs, like this but smaller: https://fab2.com), and some which are wireheaded to a datacenter, with varying degrees of success at various scales. They will compete. And there will still be unaugmented humans that continue to live and die the old school way like we do today, albeit with habitable zones compressed.
These are two vastly different statements, and the the war and diplomacy end is almost trivial by comparison, it would absolutely be solved before we solve all math in even the most steelmanned version.
There have only been a few thousand wars, and they’re all different and all different in the world in which they occurred. The dimensionality is absurd, which is not a problem for LLMs if there’s enough data, but in this case there isn’t.
Nobody that understands automated proof checking was claiming that.
Edit: I mean technically shareholders don't decide any operational choices but their shallow interests are what everything is decided around.
And yet, the main tangible contribution emerging from this early recognition (and obsession) was to inspire a lot of brilliant CS/math types and billionaires like Peter Thiel or Jaan Tallinn to build and fund companies like OpenAI and Anthropic which went on to literally create the very thing they were convinced will destroy the world, leading to an ever escalating arms race to “AGI”. Also lots of crypto and a murky web of non-profits doing unclear stuff.
The proximity to money in the Silicon Valley/VC world seems to limit the imagination to founding yet another new tech startup that really, definitely won’t compromise their values and will be different this time.
I can't make any sense of that.
By the way, I donate monthly to GiveWell and the shrimp welfare project. Do you donate monthly to starving people? Most people don't, and the reason is simple: at the end of the day, most people just don't care that much about helping a starving person far away. They also don't care much about things like that, animals in factory farms, or earthquakes that kill hundreds of thousands. But ideally, we could find a way to empower people such that the minority who do care can make a big difference.
[1]: https://openai.com/index/advancing-the-price-performance-fro...
Jan. 2023: $1920
Jan. 2024: $1811
Jan. 2025: $1639
Jan. 2026: $2193
[1] https://bestvaluegpu.com/history/new-and-used-rtx-4090-price...
[1] The announcement explains that they can't use APIs for privacy reasons.
Evidently, this has not really changed with LLMs and coding agents. It’s what AI companies are betting on though.
"Oh yeah, we're gonna bring in some entry-level graduates, farm some work out to Singapore, that's the usual deal"
Office Space 1999
It was so pervasive that it was satirized by someone that had worked in engineering in the 80s
Yes, but technologists get sniped into identifying and defending them. They're the mouth pieces and the hands. Shareholders by themselves out to add my uniqueness to their assets would just be helpless randos who can't even make eye contact. Think of what "hiding behind screens" does to people, multiply that by infinite and you get shareholders hiding behind technologists.
Nah, sometimes the expectation and advertisement was that you could let go of the white collar worker because you're paying the overseas person 1/10th the amount. And "overseas person" is pretty general.
"Everybody knew" it was a bad idea to get a CS degree for a bit after the dot-com bust because of that.
(Some white-collar industries did get hit much harder by that; VFX is one I've heard in that context quite a bit.)
It's a popular movement that has wound its way through out society, and the fruits are clear: surveillance capitalism, eugenics, fraud. It's honestly quite impressive how consistently you can look at [recent negative trend] and find someone at the bottom calling themselves a rationalist and claiming the same set of influences.
The fundamental tenet of that set isn't rationalism but the belief that they are superior to the exploitable masses. That does not lead anywhere good.
> cunningham-esque laws
Cunningham's Law states "the best way to get the right answer on the internet is not to ask a question; it's to post the wrong answer."
> post-dwarkesh clout
He went on the Dwarkesh podcast. That's maximal in-crowd for some AI-pilled people.
https://data.bls.gov/cgi-bin/cpicalc.pl?cost1=1%2C920.00&yea...
This doesn't happen with all tech; people who like VR don't like it because they want to hurt others, for example. But consistently, whenever AI is discussed, it becomes clear that some people support it just out of malevolence.
Remember, society as a whole will decide what's acceptable, there are no objective, natural laws that decide that. My gut feeling is that this is starting to cut close to what most people will find objectionable.
Today it’s dominated by the military industrial complex and advertisers like Google and Meta.
The hippies were never in charge, only for a brief period were they well compensated and allowed to make noise(1990s).
Do you have a sense of what shapes this? There are so many people in that orbit whose writing I find both exhausting and suspect. I work hard to write clearly, which forces me to think a bit more clearly. The in-group thing makes sense, but there's an element of obliqueness or roundaboutness that makes me think of how a squid deploys ink.
AI models get better and more efficient every 3 months, run around the clock, can be copied infinitely, and unprecedented amounts of capital and research talent are being thrown at any limitations we can see with them (such as problems writing correct code in 2024, lack of agency in 2025, autonomy and self-improvement in 2026). That's the difference between labor replacement through outsourcing vs. labor replacement through automation.
I know somebody who says that doing debate in high school and college was terrible for him because it trained him to always argue his point but never to question it. For example, at the time he thought of himself as supremely rational even as he'd make a strong case for why he shouldn't quit smoking. It was only much later that he could admit that was his addiction and his ego talking.
This isn't analogous to the threat of AI.
>hurt others
>malevolence
The irony here, of course, is that you are casually insulting me for being interested in this tech in the first place and refusing to apologize or back down. Like how the other guy needed to chime in with "dehumanizing and revolting", even though nobody is forcing him to use any of this.
Yes, I'm interested in this. Yes, I think that comments like "dehumanizing and revolting" are a signal. I'm not going to feel bad about exploring this tech. So, yes, I am that person and will continue to be that person.
>society as a whole will decide what's acceptable
So, how would you describe the past few months? Years? You keep saying that we need to reject AI, yet adoption and the gamut of use cases keep ever expanding.
The idea that “AI models” have acquired “agency” as of 2025 and are working on “self-improvement” in 2026 is closer to delusion than exaggeration.
https://www.lesswrong.com/posts/udFuYqqNdpdo5ym3f/genuine-qu...
(Even assuming "intelligent" is a sensible label to apply to an LLM holding hands with a shell script in an infinite loop)
> yet adoption and the gamut of use cases keep ever expanding.
I don't see how that's evidence in favor of AI, as many harmful technologies and products became very popular. In fact, that's exactly why I think a slower, more measured approach is necessary. Otherwise once the harms become apparent everyone wants to shrug their shoulders and say, "well it's too late now, everyone's using it."
I wouldn't have any problem with you being interested in this tech if it wasn't driven by spite. Lots of people manage this.
But you said
> After reading some of the outraged comments here, I'm even more excited.
The reason you gave for excitement was the upset it causes others. There are people who like AI who aren't driven by malevolence, but you're not one of them.
Models have absolutely acquired agency as of 2025. Developers are no longer copy-pasting code from ChatGPT into their text editor, they're working with agents like Claude Code and Codex that can edit code, run terminal commands, do web searches, manage their own context windows, sift through gigabytes of logs with datadog MCP, etc.
Self-improvement is also being worked on. Claude Tag learns over time in slack convos. My company also has an agent that updates its own skill files after every conversation so that we don't need to keep reminding it about the same workflows every time. Is it clunky as hell? Yes. Are the labs plowing billions of dollars into "continual learning" and "recursive self improvement"? Also yes.
So, what's the "upset"? Why is anybody here upset about what we're doing/I'm doing? I'm not backtracking on anything, I'm doubling down.
>malevolence
lol. lmao even.
What you call a model acquiring agency I call plain old software with productivity workflows designed by humans, with deliberate goals. We must separate “model” and an execution environment using a model. [Model] ≠ [A glorified shell script doing API calls in a control flow based on heuristics]. Agents are not AI, they are plain old software. The weights are the model, and that very much remains a static artifact (and pre-post training models haven’t improved much over the last few years).
What you call self improvement is a duck tape hack to imitate persistence and save on inference. Every time you do an API call, anything that needs to be processed is sent to the model. Narrowing that context down saves money. Finding clever ways to do that improves apparent performance and value. The cleverness is still human.
These are all useful innovations on top of LLMs, which remain models that generate text and symbols based on static weights, which in turn represent training data and the provider’s preferences.
That's weird; fans of [almost anything] don't do that.
Yeah, let's just ignore the bit where we're minding our own business, pursuing what we want, yet you feel the need to bark stuff like "dehumanizing and revolting" at us.
>rather than any positives about AI
I'm a huge supporter of AI. I want this stuff. If this upsets you, I like it even more.
> If this upsets you, I like it even more.
wouldn't be appearing. Getting off on others' pain is the opposite of minding your own business, even if only to hopefully a small degree.
Liking things because other people dislike them is at best petulant. It's more cruel.
>cruel
Allow me to laugh even louder. Reminds me of the "you're prompting with Hitler" meme.