Out of all the big orgs, Amazon is probably the one that has the least regard for keeping technically interesting projects alive, and the certainly will bulldoze it for some dumb reason when the next re-org comes.
That said, the acquisition by Amazon is interesting in that it seems to overlap with their existing DB offerings? Which makes me wonder what exactly it is that they are getting out of this.
I hope they just let the team carry on and don’t contaminate it with all the other craziness going on.
I know the projects will remain open, and ostensibly still contributed to in the same direction, but surely AWS thinks of this as another piece of a product suite to compete with Databricks.
Nontheless, I think the open embedded query engine approach DuckDB is spearheading is larger than one project, and I remain excited about the broader ecosystem (especially datafusion).
"SedonaDB is an open-source single-node analytical database engine with geospatial as a first-class citizen"
Apache License 2.0 https://github.com/apache/sedona-db/blob/main/LICENSE
MIT License https://github.com/duckdb/duckdb/blob/main/LICENSE
https://duckdb.org/docs/current/core_extensions/spatial/func...
"AWS has committed to supporting the continued development of DuckDB and its wider community for the long term."
Those are just words man.
"We also worried that scaling DuckLabs into a much larger sales, support, and operations organization would pull our attention away from the technical work and open-source community that made DuckDB successful in the first place."
Given AWS's services arm seems like a good play for a team. Congrats!
Quote from the article:
"As the CWI representative on the DuckDB Foundation... When DuckLabs spun out of CWI, we created this foundation, which holds all IP of open-source DuckDB, and will continue to do so." - Peter Boncz
Are databases not a solved problem? Why are there lots of different databases? Why is one faster than the other? What's different between them?
Now, with Snowflake and Databricks earning big bucks in the intelligence era, time to gain market share by having DuckLabs under its ownership?
From my own experience I can say it integrates far better into your Rust app than DuckDB does.
Well over 100 monthly contributors, too.
I've been eyeballing DuckDB and LanceDB as part of AI agent memories. This gives me a vibe that AWS will use DuckDB somehow in their ai agents sometime in the near future after seeing the potential.
Purely acqui-hiring?
I trust the person leading DuckDB, nevertheless.
I use DuckDB extensively for local dev as well as a parquet viewer.
one feature I can't wait from the DuckDb team are real-time materialized views.
Well ...
Either way that news is awful. Europeans selling out to US corporations and then wondering why Europe is not competitive in software engineering anymore. Funniest thing is Draghi keeps on saying that; well, recently we also heard that the USA has full access to all national police databases in the EU now. It seems as if Washington remote-proxy-controls the EU. Quite amazing to see, too. The amount of bribe money flowing must be legendary.
DuckLabs will be an AWS subsidiary, not subsumed into AWS itself.
Why do we need a source distribution for a well regarded MIT licensed project? Because it's not easy to contribute code to DuckDB if you don't work at DuckLabs. The CI used to take 5 hours for a simple bug fix last I looked (may have improved since I flagged it on social media).
There are two that I'm aware of:
* Haybarn: https://rusty.today/blog/duckdb-extension-distribution-gap/
* Pygmy-Goose: https://github.com/Pygmy-Goose/pygmy-goose
Pygmy-Goose is focused on making agentic workflows faster by splitting the repo, making git worktrees cheap and 5 minute cached CI on 3 major platforms.Current Google cache of this page states:
> "The DuckDB project is governed by the non-profit DuckDB Foundation . The Foundation and DuckLabs are not funded by external investors (e.g., venture capital)."
But "(Last updated: Aug 2026)" and I was not able to find this text anymore.
I wonder if they took vc money after all. I guess not, but still, it seem like to be difficult to live as an independent open source company. Getting a big co as a parent/sponsor is probably the next best thing.
I hope the future for DuckDB is still bright
I'm curious what you mean by this. I would have said the exact opposite - AWS tends to keep projects around for a very long time. They haven't acquired very many open-source projects, but the few that they have are all still running as far as I know.
2026-08-26 | 10 min
Today, we’re announcing that DuckLabs will join Amazon Web Services (AWS), which is expected to be effective in early September.
Our team will remain together in Amsterdam, continuing our work on DuckDB, DuckLake, Quack, and the broader community. Joining AWS gives us the resources and reach to bring this technology to many more developers and organizations, and to pursue ideas at a scale that would have been difficult for us to reach alone.
Most importantly, the foundations of the project will remain firmly in place. DuckDB and the other open-source components of the “Duck Stack” will remain free and open source under the MIT license, with the nonprofit DuckDB Foundation continuing its stewardship of the projects.
This is a significant moment for all of us at DuckLabs. It marks the end of one chapter that we are immensely proud of, and the beginning of another that we believe can take DuckDB much further.
We founded DuckLabs a little over five years ago to give the team behind DuckDB a stable, long-term home.
At the time, DuckDB was beginning to gain real momentum. The first commercial contracts to prioritize features were materializing, and venture capital firms were calling. We chose a different path: a bootstrapped company, fully owned by its founders and development team.
That decision shaped DuckLabs in ways we value deeply. It gave us the freedom to build patiently, to put the technology first, and to grow without losing sight of why we started. From a small group gathered around an ambitious open-source project, we grew into a team of more than 30 people in Amsterdam, all while continuing to invest in DuckDB and the community taking shape around it.
During those years, DuckDB traveled much further than we imagined when the project began. Today, we see more than one million downloads every day. Developers around the world use it to explore data, power products, teach new ideas, conduct research, and build systems we could never have anticipated ourselves.
Watching that happen has been one of the great privileges of our professional lives. But the limits of our existing model also became increasingly clear. As founders, we worried that DuckDB’s growth would eventually outpace our ability to support it. That our small company could become a bottleneck for the project, the team, and the people building businesses on top of it.
We also worried that scaling DuckLabs into a much larger sales, support, and operations organization would pull our attention away from the technical work and open-source community that made DuckDB successful in the first place.
Our partnerships work best with highly technical organizations – often companies with substantial database expertise of their own. Reaching a much broader group of users requires us to solve more complete and specialized problems, serve the needs of different industries, invest substantially more in infrastructure, and reach people who may never think to seek out an analytical database directly.
We believe the DuckDB revolution can grow by another couple of orders of magnitude. To give it that opportunity, we realized we needed a different setup.
DuckLabs and AWS have already been working closely together for more than a year. During that time, we came to understand how our teams collaborate, what each brings to the table, and what might become possible by combining DuckLabs’ technical expertise with AWS’s infrastructure, scale, and customer reach.
That experience gave us confidence in taking this next step.
Together, we plan to use DuckDB, DuckLake, and Quack to help power a new generation of data services. Our ambition is to reach people who may eventually depend on the Duck Stack every day, whether they interact with it directly or encounter it quietly inside the products and services they use.
For the DuckLabs team, joining AWS creates the space to concentrate on the technical work we care about most while operating at a scale we could not easily achieve on our own. It allows us to think further ahead, be bolder, tackle harder problems, and bring the ideas behind DuckDB to many more people.
AWS has committed to supporting the continued development of DuckDB and its wider community for the long term. That commitment matters deeply to us.
Quote “DuckDB is an incredible open source project with an amazing community; it is broadly used and very much loved by S3 customers today. After about two years of working closely with Mark, Hannes and the whole team at DuckLabs I’m excited at the opportunity to help the project have an even broader impact. Also, and maybe a little selfishly, I’ve found the DuckLabs team to be one of the most technically deep, humble, and high-velocity teams that I’ve ever had a chance to work with and I’m delighted that we get to do a lot more of that.”
– Andy Warfield (Distinguished Engineer and Vice President, AWS)
We know that an announcement like this brings important – and understandable – questions for users, contributors, customers, and partners.
DuckDB, DuckLake, Quack, and the other open-source components of the Duck Stack will remain free and open source under the MIT license. The nonprofit DuckDB Foundation will continue to steward these projects, and the DuckLabs team will continue to contribute to the project and remain together in Amsterdam.
The open-source projects will continue to serve a broad community of users, contributors, platforms, and vendors. The openness that allowed DuckDB to flourish will remain central to its future.
We care deeply about the trust this community has placed in us. Protecting that trust has been one of the most important considerations throughout this process.
Quote “As the CWI representative on the DuckDB Foundation, I would like to congratulate Hannes, Mark and all DuckLabs employees with this new chapter. When DuckLabs spun out of CWI, we created this foundation, which holds all IP of open-source DuckDB, and will continue to do so. I am delighted that AWS is committed to keep advancing open-source DuckDB, and I anticipate it to even accelerate its innovation. Via the DuckDB Foundation we will also make sure that the voices of its supporters and all members of the open-source DuckDB community at large, will continue to be heard. I am very proud of DuckDB and its creators; this development underlines it is state-of-the-art technology, which realizes many ideas from the database architecture research group in open source, for everybody’s benefit.”
– Peter Boncz (Database Architectures group lead at CWI Amsterdam, DuckDB Foundation board member)
Quote “Over the course of five years, DuckDB’s open nature, its sheer hackability, and the remarkably friendly community of developers has turned the system into one of the premier platforms for database research as well as teaching. For both, it is essential that we can inspect and tinker with DuckDB’s kernel. I am thus excited to learn that DuckDB will remain open source under the umbrella of AWS. An even wider range of opportunities is in reach now. We are glad to be able to be a part of this new era for DuckLabs and DuckDB.”
– Torsten Grust (Professor of Computer Science and Database Systems research group lead at Universität Tübingen, Germany)
Joining AWS will give us greater capacity to invest in both the technology and the community around it. Going forward, the DuckDB Foundation will include a technical advisory board, so that leading community members can provide their input on the project’s technical direction. We also plan to open the extension stack so that extensions signed by other developers and organizations can run in DuckDB.
There is a great deal of work ahead, and many details still to shape. We will share more as these plans develop.
Quote “Amazon putting its weight behind DuckDB is going to add a ton of momentum and strengthen the ecosystem.
This is great news for those of us who believe in DuckDB as the platform on which the future of analytics is being built.”– Jordan Tigani (CEO, MotherDuck)
Quote “Amazon is the ideal home for DuckLabs. DuckDB is the center of the vendor-neutral open data stack, and Amazon has the commitment to openness, and the track record of working with the entire cloud ecosystem, to enable the DuckDB project to continue to thrive in this role.”
– George Fraser (CEO and co-founder, Fivetran)
DuckDB’s success has always belonged to a much larger community than the people working inside DuckLabs.
It belongs to everyone who has used it, contributed code, reported a bug, written an extension, answered a question, published a benchmark, taught a class, built a product, challenged our assumptions, or recommended DuckDB to someone else. Every one of those acts helped the project become what it is today.
We do not take that support – or the trust behind it – for granted.
Joining AWS gives the DuckLabs team an extraordinary opportunity to bring the Duck Stack to a much larger audience while continuing to invest in the open-source technology at its heart. We enter this next chapter with the same curiosity, care, and technical ambition that brought us here, alongside a team that has been through the entire journey together.
To everyone who helped us reach this point: thank you. We are proud of what we have built together, excited by what now lies ahead, and looking forward to building the next chapter with you.
You’re also welcome to check out our press release and the media kit.
Maybe this will be the rare case where founders survive corporate shenanigans and keep on doing their thing. But I'm bummed, as I've heard that optimism many more times than I've seen it happen.
It's both a great time to be gold mining in tech, and also a terrible time to be a bit employee.
I can’t say that’s true for me and my company. It’s not all lies or anything, it’s a fragmented view that emphasizes conflict and the most visible 10% of what the company actually does.
What big corp is not a total mess?
KuzuDB folks worked on something called GRainDB in 2022: https://vldb.org/cidrdb/2022/graindb-a-relational-core-graph...
But circa 2023 decided to write their own. Work continues as LadybugDB. One of my long term goals is to find an integration point with DuckDB's table implementation as the "node table". Conversely at some point DuckDB could implement all the join algorithms and LadybugDB's REL table in their code base.
For now, any talk of Graph on DuckDB is limited to DuckPGQ and the more recent entrant DuckGQL (both of which don't touch the storage layer which is the main reason why LadybugDB exists).
And DuckLabs is the engineering behind DuckDB.
If I'm correct, this seems like a signal that AWS is coming after the segment of customers wanting to host DuckDB (MotherDucks customers).
Correct me if I'm wrong though, please.
"Are cars not a solved problem? Why are there lots of different cars? Why is one faster than the other? What's different between them?"
I think the best way to approach the subject in an easy to grasp way is to ask Gemini or another frontier AI to teach you the basics, they will do a surprisingly good job and they'll be able to react to your questions with INFINITE patience.
"DuckDB has been an in-process database since day one. But people have asked us – very persistently – for a client/server mode, and we have finally caved"
This was already satisfied by numerous projects and products, and feels like a "me too" attempt to capture AI-based workflows. DuckDB always felt like "SQLite for Analytic Data" but I fear these changes and now acquiring the org leading technical direction is where they deviate for good. AWS is so unnecessary for what DuckDB can (and should, IMO) be; MongoDB jumps to mind as a cautionary tale.
What do you mean by solved problem? I don't own DuckDB (sqlite, postgresql, etc). If I think I can create something as good or better than DuckDB, should I give up doing so (and get filthy rich with an acquisition) because someone thinks databases are solved? Solved databases aren't mine.
Every tool is a trade off between effort to create vs power of the solution.
Effort is generally expensive so most things settle on some general purpose local maximum. If you had infinite effort available, you could build bespoke hardware and software from the ground up to solve every problem. It would be faster and more power efficient than any solution available today.
CPUs win out over integrated circuits because the same CPU can be used for ~every software problem, so by using a CPU you benefit from everyone pooling their efforts to improve the general purpose CPU rather than their own specific niche. But when you reach a certain scale/requirements it makes sense to do something more specific. This is one reason why we have standardized GPUs. Still general purpose but more specialized than a CPU. Or think about how Bitcoin mining moved to ASICs, because they need to do one specific thing as fast and as power efficiently as possible.
So for databases, when you get to specific scale and requirements the same kind of specialization starts to make sense. DuckDB or Clickhouse for analytical loads, TigerBeetle for high scale transactional stuff, etc. And that scale is aggregated across ~all software users, i.e. scale of analytical workloads being big enough to support analytical DBs.
Also as time goes on and industries develop the cost to develop specific solutions can go down.
It’s not a solved problem because each iteration of technology doesn’t just fix the mistakes of the past, it’s an evolution to solve the problems of the present.
I was really surprised when I first came across DuckDB at how good at it is for its target use cases. It is a game changer for me for the "local analytics" space, and its ability to scale up to a large degree helps a lot.
It is simply awesome to be able to point DuckDB at a mess of CSV and other files and have an instant database on top of it that I can run regular SQL over, and it is fast and just works.
AWS did not acquire the DuckDB technology itself, which is MIT-licensed open source and governed by the DuckDB Foundation which holds most of the related IP [1].
What they did acquire is DuckLabs, the Amsterdam-based services and development company behind the technology which is owned by and employs the creators and major contributors.
And MotherDuck is a US-based venture-funded commercial company, whose cloud-based data platform is centered around DuckDB but has been significantly expanded recently, including Python pipelines, an agentic context layer, and a visualisation layer.
learn about it here: https://query.farm/haybarn/
There's no single set of requirements and desired properties that people have for databases.
What queries does it accept? How does it persist data? How does it manage replication and partitioning across multiple servers? Are questions with many answers and the right one varies by application.
My guess is that Amazon wants official hosted versions and doesn’t want to go through something like the Redis fiasco with licensing. In that case, they ended up having to support their own development anyway (with Valkey), so they might as well just buy the team.
DuckDB team probably gets a nice package and pay bump, but it’s really unlikely they’re getting hundreds of millions from this.
My attempt at an answer: no, databases are not solved.
Specifically, different databases are better or worse for different use-cases.
For example, Postgres is a great "all around database" - you can use it for a lot of different things. As a "relational" database, it's really good if you have a table full of users, a table full of order, and you want to see all orders made by a user with ID=123. You need to answer questions like that a lot (eg every time someone on a website loads a page) and you need the answer fast (hundreds of miliseconds at most)
However, say your use-case is more like... you've got 100 billion rows of billing data ("joe was charged $123.45 on 2026-03-07 for a shirt, blue, size 11, brand foobar") in one table. You don't care much about joe, but you want to be able to find out how much was billed, total, in 2026-03 for blue shirts (or all year for brand foobar, or all time, for size 11). Postgres would struggle with data of that volume - you'd need a really big expensive database. A "columnar" database like duckdb (or clickhouse) might be able to answer those questions better.
Anyway, different databases are better/worse for:
- Large piles of data that you need to query in seconds
- Huge (petabytes) of data that you need to query in minutes, but can query in parallel
- Many related piles (like a standard relational database)
- Cases where you're mostly getting or retrieving single items (key-value stores)
- Huge piles of data that represent a long stream of events in time (time-series datbases)
- Piles of data that look and act more like files (object stores)
- When you need strict transactions
- When your need is very write-heavy
- When your need is very read-heavy
- and probably many others - I'm not even a huge data guy :)
So it all depends on your use-case. There are still cases that are not served well by any existing database - eg "filtering billions of rows, in milliseconds, by an arbitrary portion of several dozen very-high-cardinality columns" (to use an example that came up recently for me IRL) :)
Indexes, transactions, a first-class storage format are all things that come included with DuckDB, that you won't have with Datafusion.
From the MD "about page":
The idea for MotherDuck came after Jordan Tigani, MotherDuck co-founder and chief duck-herder, saw DuckDB in action and thought, “Wow, this is amazing! someone should really build a serverless version.”
...
Hannes and Mark, who founded DuckLabs to focus on the core technology and build the world’s best analytics database, were looking for partners who would build a commercial cloud offering.
* Cloudflare also seems to have gotten the same idea right with R2 Data Catalog
100 items max per Transaction BatchGet 100 items, 16 MB max low write limits on same key Item size 400 KB max etc.
Heck AWS is the sole reason all these projects needed to go through these license changes to prevent AWS from completely destroying their business models.
Of the few examples I have I my head, I'd even say that the fate of a product that has been acquired by Google is probably even better than those of Google's internally developed products. e.g. Waze is still alive and kicking 13 years after acquisition under its own brand and hasn't been completely swallowed by Google Maps. The Nest brand also stuck around for quite some time.
Buying the team and having them all quit is such a bad outcome for AWS, especially they would not keep any of the IP, which is owned by the DuckDB Foundation. I don't think they would be that short-sighted.
Is profit sharing, like what you're implying, actually happening? If so, in what quantity?
These kind of arrangements generally get the investors their money back plus a modest return. The employees generally don’t get much, relative to the risk they took.
The things I hear from recent AWS departures are very concerning.
I can't think of many other languages/frameworks where one of the worst places to install it is from the primary sponsor/maintainer.
Currently it says “Today, DuckLabs remains independent and fully owned by the original creators. We've deliberately chosen not to take venture capital, so we can focus on sustainable engineering, correctness, and keeping DuckDB open and MIT-licensed for everyone. “
Future will tell if that ends up being true.
Check this curated list of databases from Carnegie Mellon University: https://dbdb.io/
Assuming historical returns, your money doubles roughly every 7 years, so within the rest of your lifetime, that 1 million should turn into at least 8. That's an extremely comfortable upper-middle-class lifestyle on the interest payments alone. If your children don't spend it all, your grandchildren would easily have private jet money by the time their parents retire.
DataFusion is being used as a building block for a growing number of databases and data processing engines in Rust, as it offers the necessary primitives.
nothing prevents to build single database which would cover all such answers. Its engineering, funding and distribution problems: no-one built it yet.
Presumably the issue is Haybarn-specific?
MotherDuck, a semi-competitor of theirs has raised $100mil in funding. DuckLabs has reportedly taken no external funding (so all ownership is with the founders) and was profitable with 30+ employees.
On the sidelines, other companies behind beloved open source products, like Astral, Astro, Bun are getting bought left and right.
With that as a backdrop, I think they should have been able to get quite a good payout.
DuckDB to the rescue.
Also embed the crap outta DuckDB locally in local agents running on VPS, etc
Weirdly I’d put Microsoft above modern-Google, and that’s still a low bar.
If I had 5 million I’m going to retire and never code for money again.
I do want to make small video games though, make some music. Pay for a friend’s kids college.
I wouldn’t waste a single extra hour making more money. Usually the golden handcuffs fall off after a year or two. That’s why Heroku went to crap, all the core people left.
I wish I could get a role to work on OLTP databases. PostgreSQL seems to be a fascinating topic so that's on my plate.
1. You're selling a system to a customer to use within their own AWS account, that you will have no access to, and it needs a transactional datastore (not just an object bucket) of some kind. The fact that it costs nothing by default (particularly valuable when the customer is trying to deploy a proof-of-concept), scales more-or-less perfectly without anybody touching it, requires zero day-to-day maintenance by you or the customer, and all it will ever ask is that you throw money at it, is very, very much a feature. One example I'm familiar with in the wild is Teleport: https://goteleport.com/docs/reference/deployment/backends/#d...
2. You have a huge OLTP workload that fits Dynamo's KV patterns (e.g. Amazon.com shopping carts, which is what it was originally built for). You don't care how much DynamoDB costs (in either dollars or engineering limitations) because any alternative would melt your face off if you even tried.
Most of the pain that comes from Dynamo is people who try to use it as a primary datastore in place of a relational database just to get the serverless pricing model. It's not worth giving up the flexibility on greenfield systems. It does become worth it to give up the flexibility when your system is mature and you don't have genuine flexibility anymore anyway.
The constraints are what let this happen. Unconstraining it might make a better generalist product but part of what you're opting into with DDB is the dumb "put an item in get an item out semantics" and the other side is knowing that it will still work if that volume increases dramatically.
Amazon will drive these people to quit within a couple years over filling out MBRs and threats about how the MBR isn't good enough (MBR = monthly business report, pure bullshit theatre that drives the whole company mad 2/4 weeks every month).
Google would drive them to quit over a longer timeframe, with no threats or harrassment, just because every time they try to ship something cool, someone else has a reason not to do it and that kills you eventually.
For the same price, you can run a much more capable PSQL instance with way better features, but now you're on the hook for it being up 24/7.
On my team we were expected to work 90+ hours a week, and were told we were slackers if we were unwilling to work Sunday afternoons. But I've heard from others that some teams are a lot better.
I wouldn't want my info to be in those DBs.
So give your tools, your life, free bug fixing, priority attention to me because I am getting my paid job done. Why they need money anyway, they can leave on reputation of OSS contributors. Also not to forget I donated 5 dollars last year so now give me full certified audit of your finances of last 5 years.
For example, I want to know how to calculate the charge of incremental export. One blog says it is charged by the amount of change logs (but the official doc doesn't say so), which makes sense. But how do I estimate the amount? My hunch is: Put + Write + (1~100) * TransitWrite + Update + Delete + (1~25) * BatchWrite.
Hah. Never have I ever heard a sentence that described Amazon less.
One blog says it is charged by the amount of change logs (but the official doc doesn't say so), which makes sense. But how do I estimate the amount? My hunch is: Put + Write + (1~100) * TransitWrite + Update + Delete + (1~25) * BatchWrite.
You would also have to train your kids to responsibly use the money without demonstrating it, as you’d just be hoarding it. People don’t have a good track record there, either.
I think the best you could do generationally with a million would be to try to invest moderately and draw down a small percentage (2-3%) to demonstrate fully considered use of the money. This would keep you in a middle class income but give you more ability to donate charitably, vacation together, let one spouse retire earlier, solve a financial crisis for a child, etc.
Letting them in on the thinking would give them a good chance to handle a high six/low seven figure inheritance, depending on how the investing goes.
The DuckDB foundation owns some equity in MotherDuck, which is a data lakehouse platform based on DuckDB, and all three have been moving closely together in making DuckDB better locally as well as in the role of a query engine that really threatens a lot of amazon's role in data lakehouses. DuckDB in the hands of a good team means that you can use AWS almost only for storage, instead of using any of the managed services.
Buying a house in desirable areas is going to be in the ~ $1M range, college is going to be hundreds of thousands, etc.
Common definition is somewhere between 1.5 to 5 million per child.
The CI extension build runs are still public on GitHub, but may have expired, if you bump your git sha ref it will rerun.
Even if you did it, I'd be surprised if a single code base is optimal across the spectrum from resource constrained microcontrollers all the way up to IBM Z Series mainframes and everything in between; along with the full spectrum of persistence from in-memory MRU (or similar) caching to full two phase global consensus as well as optimistic eventual consistency.
There's just too much scope to be the best in class at everything, and even if somehow you did it, the maintenance cost of all the options would be huge.
Best you'll get is as computing continues to increase, the cost of using less than the best becomes more reasonable. There's a lot of database applications where any reasonable database works, and there's a lot of databases that are reasonable in wide application. That may well lead to fewer databases being available, but it's unlikely to converge down to a single database. Just as most engineering domains don't converge down to a single solution for all applications.
I can't speak to how successful it is, but Daniel/Curl have been running with this model[0][1] for years. Probably not too different from 'drh and SQLite[2][3].
[0] https://curl.se/support.html
[1] https://rock-solid.curl.dev/
Couldn't have said it any better myself having worked at AWS for close to 8y. I was there during their peak growth years and remember very well how some of these processes that were once not as maddening, devolved.
Yeah, I've never personally worked at any of those places, but collaborated on a few open source projects with Googlers and Xooglers and the slow grinding wheels of their "processes" that wear people down over time were very apparent. Nevertheless, it has been a breeding ground for many interesting technologies, even if it often suffocates them over the long term.
You can find monthly archives of every reddit comment on Academic torrents. These are huge NDJSON files compressed to .zst, named like 'RC_2026-01.zst'. The size is ~60GB compressed, 350GB+ uncompressed, per month.
Most of the size is taken by the actual comment text. But I was only interested in calculating how many unique commenters subreddits have in a month so I only wanted to extract a few fields from it and discard the rest.
If you use traditional tools like pandas or load the data to a database and then query it, you would quickly run out of RAM or storage space, especially when doing it on a basic laptop like I was. But with DuckDB it's just this:
``` SELECT lower(subreddit) AS subreddit, author FROM read_json(['RC_2026-01.zst', 'RC_2026-02.zst']) ```
DuckDB automatically handles decompression on the fly, figures out the schema, manages RAM so you won't OOM and so on. And even on my laptop that query takes like, 2 minutes? Which is super impressive to me.
You might think that's got a very low ceiling. But, even though it's a bad example in many ways, OpenAI showed that the ceiling is very high. And if you're morally flexible even higher.
Wouldn't be possible nowadays to LLM your way out of this?
Really no idea, just asking.
You really underestimate the value of support contracts.
>Independence: We are fully founder-owned and not externally funded. Our priorities align with the long-term health of the project and its users.
is a marginal niche, but everything else can be handled by one DB. You think we have PG with extensions (including duckdb) already covers most of the ground, making something like that distributed is also feasible task (there are projects), architecting system to make core engine embeddable is also feasible task.
At a more sane expected return of 5% annually, you get $50k a year to live off of to just keep what you have (or rather, watch it slowly erode in value due to inflation).
That is basically a one person income, maybe two adults if you pinch pennies and live in a crappy apartment or a low-end house in a midwestern suburb.
You can get a lot more lucky with a million bucks than you can with 10k if you gamble, but there are no guarantees. Risk tolerance is the most impactful variable. For those in the low risk tolerance group, I think you'd need at least 2 mil these days. And that's with being frugal, as well as probably not having much left over for kids.
Looks like this is the failed build https://github.com/Query-farm-haybarn/haybarn-community-exte... but logs have expired
Honestly, the whole thing felt LLM to start with, tons of people lacking context and throwing words and summaries around.
I don't know how you came to OpenAI as an example given that they famously succeeded while going closed-source with the release of ChatGPT.
When a bigger entity (e.g. AWS) decides to undercut the original creator/vendor (e.g. redis,elastic search), MIT code des not help.
Yes I think motherduck is going to get ... well motherducked bad by AWS
Not that I like either actor in this particular play, but this isn't exactly an example of community oriented goodwill on Amazon's part.
Open search wasn’t precisely born out of the kindness of their heart and their defence of FOSS values… it’s more like a predatory practice to capture market.
There are no technical reasons for having more than one codebase, only economical and coordination-related.
You need more than $1.5m today should be $6m by the time they're 20. Depending on inflation, cost-of-living, tuition, etc. that might be enough.
People talking about generational wealth aren't saying, "what if they live a frugal life in the Topeka suburbs."
I know some lawyers who specialize in these choices. Everyone thinks their choice doesn't smell, but there are the resources to make informed choices.
They used the software in a way that complied with the license. When the license changed to no longer suit them for future releases, they forked from the original license.
Remove the names here, and in context, every single person familiar with open source projects and licensing would think this is a nonissue. They used the software as the license was intended, and exercised the rights given to them by that license. That freedom to fork if they don't like where the project is going, or things change and the project is no longer viable for them with the new license, is one of the key freedoms enshrined in the whole FOSS movement.
If getting a million dollars wouldn't affect how much money you can leave your kids, you already have generational wealth.
Also wealth does not start at the ability to live an extravagant life without a salary... You're wealthy way before that point.
There the question turns to "you may sue me, if it breaks" as reasoning. In reality sueing will rarely work, but having a business contract satisfies the company's board and insurance about using the software over an "AS IS"-license alone.
We can debate “live comfortably” if you want, but no 4.5 people are doing that from the investment proceeds of 1m.
What you are describing is kind of more like social mobility.
Well said.
Where I live, one million dollars would allow me to pay off my house, open healthily sized investment accounts for my kids, pad my investment account, setup a trust and, overall, set my family up for a comfortable life in the future. I don't see how that isn't generational wealth.
Generational wealth is definitely not "affect[ing] how much money you can leave your kids." That's an equivocation - if you're leaving your kids a dollar, another dollar will "affect how much money you can leave your kids."
edit: you need a million dollars to securely retire at all, and that's if your parents, kids, or you don't get sick. If they do, a million is not only not "generational wealth" but it may not even last you three years.
Technically any asset transferred to the next generation is “generational wealth”.
But in the context of a startup sale, when you say generational wealth almost nobody assumes you mean the ability to pass a few thousand down.
The common understanding is that you have enough money that future generations do not need to worry about money, assuming they maintain an average or slightly above average lifestyle and use the money responsibly.
Basically the bar is higher than "something you can pass down". It is enough that the next generation does not need to worry about making money either.
Speaking for yourself I guess