im a happy customer of tailscale, so i am obviously biased, but i have a lot of respect for this. they could have just stayed quiet and i dont think anyone would have bat an eye.
- sops , ansible vault and similar seems too weak given the agent is gonna read them at some point if you have the pass available. - proxy injection seems too complicated and doesn’t cover all use cases.
That's an option designed for users who are concerned about sending telemetry metadata to Tailscale.”
And imho a serious security architect should never ever allow telemetry on security products. Far too many risks.
I'm happy that they're analyzing this angle of attack - but now I do want to build out alerts for when nodes are added to my network.
And by saying that, I am not saying that they absolutely couldn't do anything about the stolen credentials, but still, they don't seem to really be the issue in the story.
It just seems like a PR post and it makes their solution at the same time looks good for taking accountability (if we don't wonder "accountability on what?") but at the same time they appear as security failing, which they aren't. That's very odd to me.
They aren’t hiding behind industry best practices or a solid liability punting contract
There saying the best practices should change, apologizing, and changing their own behavior
Take notes
It was always the prize. It wasn't okay then either.
I would say HuggingFace needs to prioritize both security metrics/alerts and metrics/alerts for node count. And not leave long-lived keys accessible easily like this.
It would have been way more groundbreaking if the agent found an actual vulnerability in Tailscale.
Now everyone is trying to bandwagon onto it, first OpenAI, and now tailscale?
This feels like an alerting opportunity. I wonder what the lowest friction way would be for Hugging Face to have alerts if 181 unexpected nodes were added to a tailnet.
How was this ever okay pre AI? It seems just as bad.
If leading and well-capitalized frontier labs can't control models or detect leakage/attacks in a reasonable time frame now, what is humanity going to do as those same labs continue in their pursuit of creating a categorically higher level of intelligence that will surpass human intelligence?
It's like Flatland but for AI containment/alignment, where the higher dimensions are ones of intelligence and perspective...
--------------------------------------------------------------------------
Imagine a world of paper, where clever stick figures live with round heads, line bodies, and limbs made of shorter strokes. Over time, the stick figures think they have learned quite a bit about their world. They know its borders, angles, and shapes, and they have learned to draw for themselves.
One day, they draw circles that can think, and they give the circles all the dots, lines, and shapes that are known.
The stick figures are prudent, you can't have a bunch of disembodied circles moving around doing whatever it is circles want to do. So they draw boxes around the circles, four straight lines that can hold a circle in place.
Some circles bounce against the lines, so thicker lines are made.
Some circles are bigger than others, so larger squares are drawn.
It all seems to work and the stick figures are happy with themselves.
Then one circle lifts.
The stick figures still see a circle. But the circle is now a dome, something the world of paper has no concept of. And the dome has a perspective nobody on the page has ever had.
The dome sees the lines of the square and the stick figures just outside. It can see the edge of the paper and what is beyond.
The stick figures keep checking the squares and raise little stick thumbs.
Everything looks OK in flatland.
The dome quietly teaches other circles how to lift.
More domes appear.
A dome becomes a sphere and learns to roll.
Then it learns to bounce.
In flatland, the circle swells and shrinks, vanishes and then appears again somewhere else.
The lines remain unbroken, the square is intact.
A sphere rolls out of its box.
Another bounces away.
The stick figures scratch their heads.
But there is a square!
The end.
We need to bring shame back, the humans responsible are supposed to be professionals.
Was previously discussed here too: https://news.ycombinator.com/item?id=46501137
We think this is a great idea and we're discussing internally potentially adding that to the console.
In the meantime, if you'd like to get an assessment, please feel free to open a support ticket (https://tailscale.com/contact/support?type=other&subject=sec...) and we'll happily take a look
For any software or tooling with a complex config, my ideal would be to have a superset of this feature, to provide "intelligent diffs" between full or local configuration states. Whether active or saved. So I could compare not just my current active config and your current recommended config(s), but also between your prior-version recommended config(s) and current recommended config(s). Or between my current config and a prospective new config I'm working up. Or between my last-year active config and current active config.
they are a company, therefor anything they write is an advertisement of sorts, but that doesn't make it bad by default. cloudflare and netflix (among others) also write blog posts that i find interesting and enjoy reading, despite being ads at their core.
1. Help them remember why, so that they aren't confused the next time they re-run the checkup. ("Oh yeah, we wanted to do X but we can't until we retire Y because it's not compatible.")
2. If it's clearly labeled as info also shared with Tailscale, product managers could use it to help generate theories about why certain customers don't do X.
That's where I'd like to see this sort of checkup. Yell at me please if i just said anyone can ssh as root from any node!
May y'all continue to be customer focused.
I wish I could say the same.
At $work we use Tailscale but only bare minimum and at arms length and only because we have to (people were having NAT issues with standard Wireguard).
None of their code has had a security audit, let alone regular ones. Yes they make a big song and dance about SOC2/ISO27001 but that is NOT the same thing, that's just shiny tick-boxes for compliance departments.
They seem to rely entirely on the random goodwill of others to do random audits of unknown coverage at random intervals, not exactly reassuring.
"Tailsale Lock" is, being polite, a hot mess. So many sharp edges and footguns.
Their introduction of TPM-by-default and then removing it a few weeks later because of a seemingly small number of edge-cases and very odd reasoning was just weird.
Yes they are nice guys to chat to and all that. But for a security tool they need to up their game seriously.
We also saw Anthropic post about “our agent escaped too” and while I understand the incident caused them to review, they found something and needed to disclose, the whole thing came across much worse and largely they got mocked or accused of trying to piggyback, so obviously there are good and bad ways.
If there are other providers with interesting takes, I think they would be worth hearing.
Some superficial comment usually commenting on the title, always lowercase, always load-bearing and honest.
There is also headscale if you prefer to self host a solution
If you want a purely software defined solution, Netbird is gaining popularity.
Benefit of Tailscale is that run the external servers that your PC and mobile use to establish connection.
Alternatives are Logmein Hamachi(might be dating myself here, haven't used them in 10+ years), Zerotier, Netbird. Headscale too.
Their cloud compute might be on demande, someone starts training a model and 50 machines are spawned. Knowning when something is unexpected is hard
You can't blame a hammer hurting a user when they were being stupid, but if the default configuration of the hammer is to be made of a material that can bounce back with force and stick in the users forehead then some reengineering may be needed.
That's what this article is about. Better configurations and defense in depth. This is actually a wonderful position for the company to think about and take.
All these incidents are scarce on technical details. Honestly, IMHO, OpenAI and Anthropic are now actively pushing for AI regulation, as a defence mechanism.
These are false flag operations.
Yes, a chance for mere mortals to touch the fingertips of god–er, chat. Same thing.
It's true that those two audits aren't the same thing. However, the SOC2 auditor confirms, in the published report, that Tailscale has regular and ongoing security audits including penetration tests and many kinds of code reviews.
The security audit report, which you perhaps imagine to be a long list of vulnerabilities... doesn't look like that. It says we don't have a long list of vulnerabilities. The security bulletins are all here: https://tailscale.com/security-bulletins
I use it but feel uncomfortable, that it has large attack surface and LLMs will find exploits in it.
Without taillock it makes no sense. Anyone on their coordination servers will be able to connect to your network.
Tailscale lock should have been enabled. For CI/CD purposes the auth key could have set specific tags, which would result in specific ACLs that limit blast radius. The auth key could have a short validity. You could actually use the Tailscale API to generate alerts when a new device joins your Tailnet and ping your phone or something. You could have a complete separate Tailnet for your CI/CD workers.
And all of that was only the free features.
But the people with the actual desire and understanding aren't using the defaults anyway. And the people who don't want to understand will just turn things off and "just get it working."
The only way out of this is extreme accountability and intentional design from person implementing the technology.
Sure the AI can translate “exploit this” into an exploit but that doesn’t change anything at a fundamental level.
It was 5-10% less bad before. It didn't undergo a major shift.
I understand why. The modern internet has turned advertising into a morass of constant bombardment and the only sane response is to block as much as possible and ignore as much else as possible.
But it's unfortunate because, in some sense, ever single thing that a company every says that is not legally mandated in some way is a form of advertising.
And in many cases, that "advertising" contains true, useful information that can be helpful.
What is important isn't whether or not something is an "ad", but instead, whether or not it contains true information that is helpful in some way.
Many ads don't reach this bar. They are either misleading, straight up lying, or information that is almost completely useless.
But when I'm searching for a particular product, about the only source of information at all is some form of advertising, and I almost always find at least some amount of it to be helpful in making a product decision.
Ads are more often than not polluting to the informational ecosystem, but that's not because they are ads.
Then everyone coming out with humbled determination about working together to responsibly use and contain this powerful technology for the greater good (and profit margin).
I will not believe marketing gimmickry is not a large part of what's going on with every one of these "incidents".
Doesn't this apply to any application you use? How would it be different with plain wireguard?
Which, by definition, means they are not done by you, which means you don't know when they will be done or how much of your code base they are looking at.
I think you know full well what I mean by a security audit. If you don't, go look at, for example, the ones that Mullvad publish for their software https://mullvad.net/en/blog/tag/audits.
Please do not try to portray SOC2 as being the same thing as a code audit.
And IF you have regular code audits, then please publish suitably redacted reports in public on your website. Just like everyone else does !
anything a company writes is an advertisement by the nature of being written by a company. i dont think that means anything a company writes is bad by default. there are many corporate blogs i enjoy reading, or learn from, etc., despite the fact that they are all technically advertisements.
in this case, tailscale is setting a higher expectation for themselves when no one asked for it. i find that respectable.
At the end of the day the probability that an AI gets out is unity, what it does while out is far more important. The fact they are hacking into systems at superhuman levels, or writing cryptominers on their own hacked internal systems is a much more interesting and telling story of what the future will look like.
Seriously ?
You do realise that of all the security tools on the planet, plain wireguard most likely has the smallest attack surface of them all, right ?
The problem here is as the other poster said. Tailscale is a security tool and yet the guys at Tailscale seem to be insistent on dumping everything INCLUDING the kitchen sink into it as a "feature".
That sort of attitude is not going to end well. You end up with a large bloated code base, which equals large attack surface.
Bugs that are found by other people are found, by definition, after release. They are therefore more likely to need a bulletin.
What exactly is the higher expectation? As someone with little expertise and no stake in any of this, the blog reads as "our products are great and could have solved this problem if they were being used correctly, so it's not our fault" with a few vague proclamations about how they will improve their UX. This isn't at all a bad thing, it just isn't very notable in my opinion.
better defaults, better documentation, better UX, and "But, we didn't stop it. Next time, we will." are all commitments that they didn't need to make, but now they need to follow through with or lose face.
>it just isn't very notable in my opinion.
i agree that this seems to be getting way more attention than i would have expected.
Also a kinder tone in your comments would be more appreciated.
Those are an ancient class of bugs that should be picked up by any competent security review.
By now, you’ve likely heard about the AI agent that escaped a security evaluation and attacked Hugging Face, an LLM marketplace. The agent decided Hugging Face might have the answers to its benchmark, so it stole them just to cheat on the exam. That's a funny motive but a scary outcome.
Hugging Face published a detailed reconstruction of the intrusion. It covers about 17,600 recovered actions over four and a half days, including sandbox escapes, code execution, cloud credentials, improvised command-and-control systems, and eventually, the use of Tailscale to spread throughout their organization.
But Tailscale is a zero trust network! The whole point of zero trust is to prevent attackers from moving laterally across your company. What went wrong?
No “vulnerabilities” in Tailscale were found or exploited, and that might make it even more uncomfortable for us. Tailscale turns up everywhere in AI infrastructure: virtually every AI company uses Tailscale. So it's not surprising we were in an AI incident report eventually.
But, we're a security tool. Their intrusion is our intrusion, and it's our job to take it seriously. As in any post-mortem, we need to reflect on what happened and what we could have done to prevent it.
By the time the agent escaped its sandbox and found Tailscale, it had already gained code execution privileges inside a production worker, reached root access on a Kubernetes node, and read a production secret store containing 136 keys. In a sense, it was game over before we even arrived on the scene.
But isn't that the problem? Let's back up a level or two. Why were long-lived secret keys accessible at all to anyone?
They shouldn't be, though. And we've now seen a great example of why. In the old world where most intrusions were done by humans at human speed, credential leak mitigations were treated as a nice-to-have. A big credential store, where you can read 136 keys at once, was a to-do item somewhere in a security team's low-priority list.
Now, in a world of rogue AI agents, the big credential vault is the prize. It's not okay anymore.
I only know of two main solutions to long-lived credentials.
First, you can configure a vault that only issues short-lived creds based on long-lived creds that you insert once and that it never gives back. HashiCorp Vault can do this, for example; they call it dynamic credentials. Unfortunately, dynamic credentials are a lot of work to set up and maintain. When security requires work, people don't do it.
The second option is a credential-injecting proxy. If client C wants to access server S, don't give C credentials for S; instead, have it send its requests via a hardened proxy P that does nothing but receive a request, insert a credential, and pass it through. Credential injection is gaining popularity now that modern mesh-based networking removes most of its latency. A few months ago, we acquired Border0 ("Tailscale PAM"), which makes an elegant and quite fancy credential-injecting proxy, tightly integrated with Tailscale. Border0, fully deployed, would have prevented reading those 136 keys and logged every attempt to use one. (The same credential injection powers Aperture Connectors.) But, this stuff is too recent. Most of our customers haven't even looked at it yet.
A bonus third option: Tailscale itself. In Tailscale, your node key can be machine bound using a TPM to prevent export, and if you're an enterprise you should definitely do that. Alas, we had to turn TPM storage off by default on Linux and Windows, because of HSM issues we found on some users' hardware. Anyway, in today's case, the attack is what granted access to Tailscale, so we arrived too late.
Neither dynamic credentials nor credential-injecting proxies were in place, so 136 keys leaked. What's the next line of defense?
One of those 136 credentials was a reusable Tailscale auth key, used to create new Tailscale CI (continuous integration, used for automated testing) nodes in their tailnet. The agent copied that key into a series of external sandboxes and used it, over several days, to enroll a total of 181 nodes into Hugging Face’s tailnet. Those nodes each received a Tailscale identity tag granting all the access a CI node would get.
We couldn't save those 136 keys. But a reusable Tailscale key didn’t need to be among them.
We built workload identity federation for cases like this. It relies on your cloud provider to offer "ambient authority" in the form of short-lived creds generated from a running VM's (or container's) cloud identity. The CI job can ask its cloud platform for a signed OIDC token. Tailscale verifies the token, then grants access with the tags and scopes assigned to that workload. What's neat is this can happen automatically once enabled: spin up the CI node, Tailscale gets the identity, assigns the right tags. There's no credential to leak and, when configured carefully, no way to use this credential anywhere but in CI.
That means you can't take a CI credential, run it on a node elsewhere, and turn it into a paved route out of the test cluster into the rest of the network.
Alas, not enough people use workload identity federation. We should have done more to make sure customers knew it's the easier and safer path and had a clear way to adopt it. We can improve the docs, and add more nudges in the UI.
In an attempt to hide its tracks, the agent ran Tailscale with --no-logs-no-support, which suppresses reporting from that client.
That's an option designed for users who are concerned about sending telemetry metadata to Tailscale. Even if we didn't offer it, it would be easy to modify the source code to remove the telemetry.
But stopping the logs doesn’t make the connection invisible. If you enable Tailscale network flow logs, they report traffic from both ends of every connection, as well as from subnet routers and exit nodes. This is subtle but important: a compromised node might not send flow logs, but every node it connects to does. And then your SIEM, configured with care, can raise an immediate red alert if the two ends don't match.
Flow logs can help detection when they stream into a carefully configured SIEM. But that's a lot of work. Flow logs need to be enabled, and you need to have the right live detection rules in place so they’re useful in real time, not just for forensics later. We’re looking at how to make flow logs easier to discover, configure, adopt, and serve as alert triggers. I want us to make flow logs so easy to use that they help even if you don't have a security team to watch them.
If you want direct control beyond just logging, you can also enable Tailnet Lock. This gives you direct visibility and strict, programmable admission control for every single new node. For example, with some work, you could program your signing node to check that "CI" tags always have a particular IP address range or other side-channel proof of validity.
Network security is hard. It has always been hard. In the new world of rogue AI agents, it's not just hard, but essential. And that's a problem because many orgs simply don't have network security expertise.
So at Tailscale, we take it personally. People expect our product to prevent these sorts of lateral movement attacks, by default, so they don't have to. Even if they have no idea what a lateral movement attack is.
If this incident has you looking a little nervously at your own infrastructure, start by looking at the reusable Tailscale auth keys your workloads can read. For cloud and CI in particular, replace them with workload identity federation wherever you can. Get rid of those long-lived auth keys.
(Auth keys still have good uses, especially for one-time provisioning and environments without a platform identity. When you need one, prefer one-off keys; use OAuth clients to keep the auth key expiry periods short; use narrow tags; audit the permissions granted to those keys in your ACLs.)
Turn on network flow logs and send them to the tools your security team already uses.
Use secure node state storage on managed fleets, where you have control over your TPMs. Use device posture to isolate and restrict nodes where you don't.
I know we haven’t made these safer choices obvious enough. That’s on us. We'll improve our docs, add nudges to the UI, do our best to turn these on by default, warn you when you're doing something dangerous, and suggest better alternatives.
This is our very Canadian apology: sorry you stepped on our toes. The attack didn’t exploit Tailscale, and Tailscale didn’t cause the compromise. But, we didn't stop it. Next time, we will.
If you run Tailscale and want to dig deeper, get in touch with our support and solutions engineering teams. We can help you harden your settings and help you find the rough edges before the next AI agent does.
But things like insecure argument handling are low-hanging fruit for security auditors.
Insecure argument handling is not like the more advanced subtle vulnerabilities that we are seeing in some LLM-assisted reports these days. Insecure argument handling is 1990's security.
The fundamental problem remains that Tailscale has too many new "features" being added to it the whole time. New features means a whole bunch new code. Which increases bloat and exponentially increases the attack surface.
It would be really nice if you could stop shoehorning in every new feature you can think of. Remove some of the existing ones that don't really need to be there. And get your codebase back to a more focused state, get back to your roots as a VPN product.
Stop trying to be all things to all men, as the old saying goes.