These weaknesses are why we[1] added FIFO endpoints, Polling Endpoints, and what we call "Svix Stream" as ways to do ordered state synchronization (each with its own tradeoffs). This lets people consume the events in the way that best fits their use-case. We are working on more things to make the state sync even easier. I'd love to hear about more challenges people are facing with webhooks, as we want to make these things better.
OP: I'd love to hear more about your thoughts there, and will send you an email in a moment.
P.S, if you're unfamiliar, please check out Standard Webhooks[2]. It's a spec we created to help with signature verification that has been adopted by OpenAI, Anthropic, Google, and many others. We are chipping at one webhook challenge at a time. :)
1: I'm the founder of Svix (mentioned in the post), we do webhooks infrastructure as a service.
But there is no way (for a merchant) to get the latest 'true' state as held by Adyen. So you better hope your data is exactly in sync with the notifications you got from the webhook (which it never exactly is, because there are so so many points of failures, and unlike what this author says, the docs aren't thát well presented to hold the same model as the PSP does. It is often close enough though, but you are constantly gardening your implementation, because the model also changes on their end with little information in the changelogs).
The "latest state" data exists though! If you open the customer portal it is presented to you without problem.
The article presents a good framing and is well-written, but doesn’t really propose anything new.
(I'm not serious about 1954 in particular, I am about hoping somebody here knows the CS literature better than me)
One complication in this approach involves access-windows: What if my system is only supposed to be seeing stuff that happened during two separate weeks in the year, because those are the spans when it was subscribed or authorized?
So the data-host would need to maintain a concept of "connection history" for other services, and also use that to filter/modify its real event stream, inserting artificial "initial state" roll-ups of events that happened in dark periods.
Tried to talk about this on X until the CEO of WorkOS wanted to bring it in private, then proceeded not to help at all. https://x.com/grinich/status/1913035839866835297?s=20
* Database sync of log events. * "Real" time updates.
Webhooks can already mostly handle the "real-time" event portion. Of course there are problems. as the article expounds on, but a lot of those won't magically get solved with other solutions either. Distributed real-time communication is hard. Webhooks are good enough for this purpose.
For the database of log events, personally I'd rather just have a SQLite DB I can yank whenever. Don't give me CSV or JSON or whatever I have to parse and manage, just give me a SQLite DB ready to go. I'd love you for it. I'll just take a whole fresh copy with everything thanks. Maybe you limit it to to the last X events, say 90 days or 365 days or whatever, depending on sizing of events, but just send it all every time I fetch and I'm happy enough. If I need to generate a delta to keep some other DB in sync, well that's my problem. Just give every row a stable identifier.
Problems listed are signatures, dedup, buffering, bootstrap, cron. Everything other than signatures and bootstrap, can be solved by having a counter in every webhook payload. It will increment each time. When you receive a webhook and the counter does not match, the consumer can fetch the missing data from the events API.
I agree with the author that providers simply saying "at least once delivery" is insufficient. they should have solutions that does not require an architecture diagram.
Bootstrap is better served with a bulk events API so you don't make one call per request. It can have an after/cursor pagination. Solutions that work for our internal Kafka might not be suited to work across services, over the internet.
The key, and only thing that matters, is that the cursor rides in your database, and is therefore transactionally consistent with the event. That's the whole magic trick.
We've done event streams like this at the bank I work at for years.
With SCROLL, consumers are responsible for choosing when to ask a provider for updates. Without a mechanism for knowing when data has changed, consumers will be forced to be pessimistic and poll providers for new data on some cadence.
I see two issues with the proposal: (1) SCROLL will lead to an increase in unnecessary network traffic for both the consumer and provider, and (2) because a consumer cannot know when data has changed, the lag between a consumer's local model and the provider's data model will be larger when with Webhooks.
How does the provider know what event the cursor you provided refers to?
Sounds like external state that needs to be managed ("cursor" -> timestamp)
On create a user or invoice for example sometimes it will return an error, yet it actually created the entity. This means you have to check manually after creating everything to know if its created properly.
Then you have the issue that sometimes quickbooks takes a while to update, and locks the company file while it does some background magic. This means you cannot immediately do the existence check, and also sometimes the check errors or times out which essentially means you need to keep checking forever until you can properly reconcile your db against theirs. But with hundreds/thousands of transactions per minute this state is never reached. You perpetually live in a state of trying to catch up but never managing it.
When I brought it up with Quickbooks dev support their response was literally "Its your job to make sure things are created properly in our system".
How did we get to this place where we started putting up with systems that cannot ever be trusted?
Both drafts request a subscription with a GET plus a header:
Scroll Request:
GET /scroll/feed/customers
Prefer: stream
Braid Request:
GET /customers
Subscribe:
In both systems, the GET leaves its response open to stream events. SCROLL responds with application/x-ndjson. Braid subscriptions are a 209 Multiresponse, with content-type application/http-history. This lets them support more than just JSON. You can send updates to the state of CSV, or PNGs, XML, HTML, plain text, or any media type.The author noted that it's hard to get adoption. Well, the reason that Webhooks are so common is that they are bog-standard HTTP. For this to get adopted, we need to put it into bog-standard HTTP. So we need to go to the IETF, and and extend HTTP in a general way to support state synchronization. It should just work for any existing HTTP media type (not just JSON), and any resource/URL (not just special /scroll/* URLs), and any way of marking timestamps (not just the ordered strings proposed in SCROLL).
Then we can bake this stuff into HTTP, and thus into all our bog-standard libraries, utilities, and code, and you won't have to reimplement the same sync-logic-over-webhooks again, and again, and again.
Reach out if you're interested!
That aside, what I don't understand (especially having worked on the side of the webhook sender, which is itself really tricky to get correct/performant/cheap) is why more companies which broadcast webhooks don't, say, provide direct access to Kafka topics, S3 buckets with ordered data objects landing, SQS queues, or any of the alternatives to those things.
"But it's irresponsible to expose an internal-use-only datastore directly to clients" goes one objection. But plenty of log-store systems have the notion of sharing a subpart of the log with a less-than-trusted external peer, so while exposing Kafka directly might be asking for the same kind of trouble as exposing your customer's SQL database for authenticated connection over the open internet (e.g. "we said you could issue reads, not that you could open/close TCP connections a million times a second! You just took out our message broker!"), exposing, say, an S3 bucket or Kinesis stream is much less risky because those systems have put some thought towards semi-trusted sharing.
"But everyone is used to getting HTTP webhooks and doesn't have the expertise to connect to something else"--that'd be true if, say, reading from a websocket or Postgres NOTIFY stream or Kafka topic or S3-change-notification stream were advanced techniques, but libraries around those things are so good nowadays that even the most web-tech-only low-skill developer can probably integrate with them with minimal hassle. Maybe it's just that a lot of shops literally only know how to run their code in a webserver, and have never deployed any other kind of application service/cronjob/queue worker? That seems unlikely to me, but I might be surprised.
"If we do something weird our competition will beat us on ease-of-use" goes another objection. But is it that hard given the libraries available? And can't you hedge back on the ease-of-use sell with "our data is fresher and more provably ordered and correct"?
I'm glad that SCROLL exists as a possible solution here. I'm just puzzled why more people haven't been using existing technologies to achieve this property.
Do most webhook senders literally not have a log store? Are they just firing webhooks in the middle of business event handlers and giving up synchronously if they can't be delievered?
Because if that's not the case (and I don't think it's the case), then it seems like the SCROLL API is ... basically just the Kafka consumer API. Or Kinesis. Or SQS. And so on.
Webhooks aren't at-least-once, nor at-most-once, nor are they guaranteed in-order. Some people build systems to make them more reliable, but if you really care about the data you need to think of a webhook delivery as best-effort, a bit like UDP.
That's before you get into all the extra complexities around these systems being owned by different people. For example, either or both system might have to roll back their database. Or either side might have a long-term bug in how they process webhooks, and now you have months of broken data.
My view is that the only reasonable thing is to start with the process that gets things back into sync if everything is broken. That almost certainly involves a poll or query of at least the upstream side, and maybe both sides.
I find that if you put a decent bit of engineering effort into that "disaster recovery" synchronisation, it can often act as the main or only synchronisation process for quite a lot of systems.
Stepping up from that, it's often useful to introduce webhooks as notifications only; that is, to provide a signal that some or all of the data is stale. You have to do a bit of consolidation, but this approach is usually enough to get completely reasonable latency for the kind of applications the author is describing.
Only if that wasn't enough for speed/scale reasons would I reach for a truly "push-driven" fast path. But you always have to be able to disaster recovery assuming the stream is wildly out of sync.
Some bits of the author's idea seem reasonable: certainly, I would love for there to be a standard protocol to request new data since some cursor or since some timestamp, ideally with some webhook notifications to give hints on when to poll.
The problem I have with the author's idea is that it is very strongly event-based, but the desired outcome isn't event-based. The desired outcome is almost always "the state over here looks like the state over there". Relying too strongly events ends up at the same kind of problem another level down: the "disaster recovery" script ends up wanting to compare the states anyway to figure out whether the events are broken.
Going fully event-sourced can work (although, I think, less often than advertised), but it really relies on everyone collectively agreeing on the same event stream being the source of truth. Once you start doing work across multiple organisations then that coordination is relatively rare.
What really surprises me is the variation in maturity on this topic. There seem to be people at all experience levels who are both doing this well and doing it badly. I have worked with people with decades of experience whose whole design just collapses if you ask "but what if X?" for some really banal values of X like "we have an outage for more than five minutes" or "we have to restore the DB to yesterday" or "someone, one time, accidentally merges a bug into master".
As an aside, I do find the obvious LLM-ness of the blog post and the proposal a bit disheartening. These are problems that require diligence and precision of thought. LLMs may be able to achieve those things, but that level of quality just isn't expressible in "Claudish".
Documenting unexpected or intermittent behavior: The easiest bug fixes of all.
> 1. Trigger a side effect: send the receipt, start the build, ping the channel. > 2. Keep a copy of the provider’s data correct:
On high level
1. System can either be PUSH or PULL, webhooks are essentially push and towards the end the OP is exploring the possibility with PULL. The caveat is that OP already iterated the PUSH mechanisms thrice and is aware of all the hardships and is somehow hoping that PULL would solve them. Unfortunately the grass is same on other side too.
a) The availability of the server can always be questionable in PULL mechanisms and its a lot of load on servers to support this kind of data at scale in bulk to multiple customers. You are essentially getting into database table scans. Its becomes a lot more costly with NOSQL databases.
b) The customer would end up making way too many calls to server even if data is not available or there would be additional latency when data was updated and when it was queried. This is one of the reasons why servers prefer to push instead of pull if they can find a listener available on other side.
c) CRLs (Certificate revocation lists) are good example which are available for PULL, same for all clients and yet rarely anyone does it correctly or does it at all even though its in security domain. In fact they are simple files on webservers in most of the implementations.
2. The primary use case for Webhook is for triggering the side effect and allowing the customers to choose if they want to subscribe for that event. A customer subscribing for everything even if its non-actionable should just treat it as logging data.3. Logging data can and always have gaps, it should never be treated as source of truth. I might question the need for deduplication, usually there is a unique identifier and almost all databases support insert ignore kind of clause. Logging the event data just provides you with better availability and latency, the source of truth is still with the provider if a next step needs to happen.
4. If user cancelled the subscription in stripe and the event never arrived then its a system design issue or system availability issue on the client side. The complete data checksum or bulk imports at night are attempt to fix the problem in a hammerhead way . I understand it exists in lot of places, however it defeats the whole purpose.
Push-oriented models are much easier to reason about and should be preferred where possible (cybernetically they are a closed loop, vs. push models which could literally just be a barrage of UDP packets). But they do have a little bit of overhead which makes them the wrong tool for some cases, like live-streamed entertainment or massive telemetry flows which value performance (latency, throughput) over missing a few packets.
When things change, you don’t immediately update the balance. Instead it is written to a transaction journal aka a log. The thing is this log is the source of truth. State or the balance is derived from the log.
You don’t send a continuous stream of logs. Instead it is batched and sent asynchronously. It is also applied asynchronously. It also records if the batch was successful or not.
If you have multiple systems sending their logs to a central server. No problem. The central server orders them all before applying the batches.
Every so often. The books are “closed”. Meaning the central server won’t accept any more journal entries for things that happened older than X dates.
If that is a real requirement, it seems like it'd be easier to meet by giving customers a realtime-stream/log API whose history starts when they were most recently granted access, and providing them older historical events via a separate API of the classic "ask for a report and we'll get back to you within a day or two with an S3 presigned URL" variety, then synthesizing that huge historical report in batch code that's aware of the subtleties of the customer's visibility windows.
If all events need to arrive, then the problem is not "notification" (which would be solved by webhooks) but "database replication": subscribe to new events, fetch the full snapshot, fetch the updates in range, have the monotonic value to establish the "range" in the first place. Reach the eventual consistency.
The proposed SCROLL handles half of these, which limits its use-cases.
Another benefit is that you get to exercise those disaster recovery mechanisms regularly as part of the normal functioning of the system, rather than a specialized path that is only rarely exercised (and thus may be broken when you need it the most).
If I offer my customers a source of ordered records, the "trust" in that system is the fact that they pay me to make sure records are ordered. If I sell a fast or slow log database, approximately zero customers in the world care to verify ordering cryptographically.
Or by "blockchain" do you just mean .... records with sequential IDs? Because sequential, guaranteed IDs surface gappiness/idempotency a lot easier than Markov chains over cryptographic primitives.
That also doesn't address the other core problems in the article: the replication (or data retrieval/polling) protocol is a lot more complex than a blockchain's "I can verify and replicate the entire chain state from the beginning of time to you" single behavior. People want more specificity than that.
Just one correction. My spec doesn't force /scroll/ URL's, just proposes it as a convention.
https://semiceu.github.io/LinkedDataEventStreams/releases/1....
And why formalize on HTTP rather than on a similar protocol over websockets?
1. Long Poll the cursor to pull down the latest events
2. Trigger a long-poll even if in exponential backoff because they shot you a webhook saying 'eTag changed!'
Especially since in most of the cases where it's not-totally-insane to use, the right solution is the classic distributed database which came first, where the ledger is among predefined/controlled node-membership... as opposed to a bloated mass of workarounds-upon-workarounds to make it survive being ungovernable.
I've seen some boosters pivot to saying "private blockchain", but that's a vapid marketing-lie, a contradiction like selling "single-user Twitter" when really it's just a blog.
"It’s a jigsaw puzzle where the manufacturer had the original picture, cut it up, mailed me the pieces one at a time, lost a few in the post, mailed some twice, and printed nothing on the box."
"and that’s the entire problem: nothing announces a gap."
"None of this is any provider’s bug. Their webhooks work exactly as documented. The problem is what a webhook is: a notification, “something happened, here’s a POST about it.”"
Almost this entire section is clearly written by an LLM
I didn't get that reading this (I didn't read the whole piece but I had read the parts you quote before reading your comment).
Often the LLM-beloved constructions are good usage in the right contexts.
But I wrote my commment after reading the article then the spec (https://welidev.github.io/scroll/), and so the spec was "top of mind".
The spec is just awash in LLM-isms. The cadence and rhetorical style are very Claudish. The visual style is basically "Claude's artifact plugin" (it may not be exactly that but it is an incredibly distinct signature). So the experience of reading the spec is very much an "AI slop" experience.
The reason I object to this is that the way these LLMs write is really well-tuned to gloss over small but critical details. And "small but critical details" are sort of the whole field of distributed systems.
This seems to be most true for Anthropic models (I am assuming there is some cultural defect in the way they give feedback), but it seems to be pretty universal, unless you give them some really strong stylistic anchor to a different style.
(As an aside, I sometimes wonder if this is part of the reason that LLMs seem from the outside to be succeeding disproportionately at mathematics: mathematics papers and mathematical notation may be a strong enough cultural force to override Anthropic's lack of taste and unlock the true power of the model).
I'm not saying there might not have been plenty of human guidance, but either way I don't think there's quite enough substance to this (based on everything I wrote in my comment) for this to feel like "a solution" either way.
I’ve built the same system three times now, at three different companies, for three different providers. It never has a name and it never appears on a roadmap, but it always goes the same way: the truth about your own customers lives in someone else’s database. The users live in an identity provider, the subscriptions in Stripe, the bounces in whatever sends your email, and your product needs that truth locally. So you subscribe to webhooks and keep a copy.
The first time, I thought I was building an endpoint: one route that parses the JSON and updates a row, an afternoon of work.
The afternoon grew into a week. First came signature verification, because an open endpoint that mutates your database is a hole. Then the dedup table, because deliveries arrive twice and the docs cheerfully call this “at-least-once.” Then the handler got a buffer, because a membership.created sometimes shows up before the user.created it points to. Then the bootstrap importer, because webhooks only tell you what happens after you subscribe, and it raced the live events, so it grew a locking scheme. And finally came the reconciliation cron: a job that crawls the provider’s list APIs at 3 a.m., diffs them against our tables, and quietly fixes what disagrees.
I want to be honest about what that cron is. It’s a written confession. It says: I do not trust the copy I built, and I have no way to know when it’s wrong, so I will re-derive it from scratch every night, forever.
The trust was gone for a reason. The drift never announces itself; ours was found by a support ticket. A customer had cancelled months earlier and our database still said active; some customer.subscription.deleted had evaporated between Stripe and us, and nothing anywhere was capable of noticing: not their dashboard, which showed the delivery as retried and eventually dropped, and not our logs, which cannot log a request that never arrived.
And that’s only the code. Every provider also brings its own dashboard. Three providers, three webhook configuration pages, each with its own idea of how endpoints are registered, which events exist, how test and live environments are kept apart, and where the signing secret lives. When something breaks, debugging is a tour: their delivery log in one tab, our logs in another, and a third tab for whichever dashboard I currently suspect. None of them look alike, and all of them have to be checked.
By the third time I built this system, I had stopped pretending. I budgeted for the whole stack up front (signatures, dedup, buffering, bootstrap, cron) and somewhere in the middle of writing my third dedup table, I finally asked the question I should have asked the first time.
What exactly am I reconstructing here?
I was reconstructing an ordered log. Every one of those integrations was an attempt to turn a stream of notifications back into the ordered, complete, current history it came from.
And here’s the absurd part: that history exists. It has to, because it’s sitting inside the provider; it’s how they render their dashboards, their event pages, and their webhook replay tools. The provider takes their ordered log, shreds it into individual HTTP POSTs, fires them at my endpoint over a channel that guarantees neither order nor delivery, and then I reassemble the log on my side. So does every other consumer, independently, each with their own bugs.
It’s a jigsaw puzzle where the manufacturer had the original picture, cut it up, mailed me the pieces one at a time, lost a few in the post, mailed some twice, and printed nothing on the box. And when my assembled puzzle doesn’t match the original, their support team asks me which pieces I’m missing. I don’t know, and that’s the entire problem: nothing announces a gap.
None of this is any provider’s bug. Their webhooks work exactly as documented. The problem is what a webhook is: a notification, “something happened, here’s a POST about it.” Notifications are a fine way to trigger a side effect and a terrible way to transfer a dataset, and somewhere along the way we started using them for the second thing without noticing we’d changed jobs.
Nobody decided this. The term “webhook” was coined by Jeff Lindsay in 2007, and the early uses were genuinely good fits: GitHub’s post-receive hooks kicking off a CI build, or a payment event pinging your server so it could email a receipt. The job was to do a thing when a thing happens, and for that a POST is perfect: fire-and-forget is fine when forgetting is fine.
Webhooks spread because they were the cheapest thing a provider could ship (one HTTP POST) and the cheapest thing a consumer could receive (you already had a web server, so you just added a route). By the early 2010s, “we have webhooks” was a checkbox on every API’s landing page, and the checkbox never distinguished between two very different jobs:
Job one is what webhooks were born for. Job two is what I was doing all three times, and job two is the one where every property webhooks lack (ordering, completeness, bootstrap, verifiability) is precisely the property you need.
We picked the tool that was lying on the table in 2007, and then we spent fifteen years compensating.
There’s a concept in evolutionary biology I can’t stop thinking about: the fitness landscape. Peaks are good designs, valleys are bad ones, and populations climb whatever slope they happen to be standing on. The trap is the local optimum: a small hill that’s better than its immediate surroundings, so evolution parks there, even when a much higher peak exists across the valley. Getting to the higher peak means crossing through designs that are temporarily worse, and evolution doesn’t do temporarily worse.
Webhooks-for-replication are a local optimum, and the proof is the pile of workarounds on the valley floor: signature schemes, dedup stores, idempotent handlers, retry queues with exponential backoff on the provider side and dead-letter queues behind them, webhook logs with replay tooling because consumers keep asking for replays, and my 3 a.m. cron.
The pile has an economy on top of it. Svix exists so providers don’t have to build webhook delivery; Hookdeck exists so consumers don’t have to build webhook ingestion. AWS will sell you the valley as managed services, with EventBridge to ingest your SaaS partners’ events, SQS to queue them, and Lambda to retry your handler, and you get to assemble the pipeline yourself. And an entire industry of connector platforms (Fivetran, Airbyte, every “unified API” startup) is, at bottom, pseudo-CDC: change data capture reconstructed from webhooks and polled list APIs, one bespoke connector at a time, sold as a product. Inside a database, capturing changes is a solved problem: it’s called replication, and it works because there’s a log. Between companies, we rebuild it out of doorbells.
My favorite workaround of them all is the local tunnel. Many providers ship a CLI like stripe listen that opens a tunnel to your laptop, because a webhook cannot reach localhost. Think about what that is: a product, built and maintained by the provider, reinvented multiple times, whose entire purpose is to work around the delivery direction of their own primitive. When multiple providers all need to ship a local tunnel so developers can develop, the primitive is answering the wrong question.
None of this tooling is bad engineering; it’s excellent engineering. That’s what a local optimum looks like: so much excellent engineering poured into the valley floor that the valley becomes comfortable, and nobody looks up.
But some providers have looked up. Stripe retains thirty days of events and exposes /v1/events, an ordered, listable log, and recommends reconciling against it. WorkOS ships an Events API, an ordered cursor-paginated log, and their own docs recommend it over webhooks when data consistency matters. The log keeps escaping, and each escape mints its own bespoke cursor semantics, its own bootstrap story, no way to verify a replica, and no shared contract, but the direction is unmistakable. This is convergent evolution: unrelated organisms, same environmental pressure, same wing.
The log exists everywhere, but the contract doesn’t.
Before reaching for a new design, it’s worth asking what any replacement would actually have to provide. My three integrations suggest the list: order, so changes can be applied without buffering; a way to start from nothing, so bootstrap isn’t a separate import racing the live events; deletes as data, so absence stops being the failure mode; resumability, so my downtime is my problem instead of a data-loss event; and some way to verify the result, so trust doesn’t decay into a 3 a.m. cron.
Measured against that list, the obvious candidates come up short. Polling the list APIs harder is the reconciliation cron promoted to a whole strategy: it can rebuild current state, but it burns rate limits discovering that mostly nothing changed, it says nothing about order, and a deleted object looks identical to an object that never existed. Managed delivery, whether that’s Svix on the provider’s side or EventBridge and SQS on mine, makes the pushes more reliable, but they are still pushes: still no bootstrap, still no verification, still notifications pretending to be a dataset. That path hardens the valley floor without climbing anywhere.
The third candidate is the one the providers keep half-building on their own: stop pushing altogether, and let the consumer read the log itself.
So here’s the thought experiment. What if instead of the provider telling us when there is new information, we ask the provider what new information it has for us since we last checked?
Suppose a provider served one URL per collection, and that URL returned an ordered, cursor-addressed change log of full-state events. Ask without a cursor and you read from the beginning, which is your bootstrap, with no separate import and no race. Ask with a cursor and you resume where you left off. Your entire sync state is that cursor.
GET /feed/customers?cursor=01J9XQ4R
Prefer: stream
200 OK
Content-Type: application/x-ndjson
{"cursor":"01J9XR2M","operation":"upsert","object":{"id":"cus_123","plan":"pro"}}
{"cursor":"01J9XR2N","operation":"delete","object_id":"cus_099"}
Send Prefer: stream and the response never ends: each change arrives as it commits, over a connection you opened, using the same API key you use for the normal REST endpoints. Leave it off and you get a bounded page you can poll from a cron. It’s the same endpoint, the same events, the same cursors, and the same consumer code.
None of this is exotic; it’s a paginated GET. But walk back through my afternoon-that-grew and watch what it does to the stack:
active, because absence stopped being the failure mode.The feed could carry one more thing. When a read reaches the end of the log, the provider could tell you what should be there: a count and a checksum of current state, at the cursor you now hold. You compare the two, and you know your replica is right instead of assuming it. My 3 a.m. cron, the written confession, becomes a comparison I’ve already made by the time I would have thought to schedule one.
If a feed like that existed, the fourth time I build this system would be a loop: GET the feed, let “upsert” upsert the object into my db and “delete” delete it, and save the last cursor. That’s twenty lines with no route, no secrets to rotate, no queue, and no cron. The replica carries its own proof of correctness, and when someone asks which customers have an active subscription and a bouncing email address, the answer is a JOIN across local tables with no silent asterisk attached.
Nobody serves this today. That’s the catch, and it’s also the point.
I wanted to see whether the idea survives being written down precisely, so I drafted it as a protocol: SCROLL, short for Synchronized Change Replication Over Line Logs, at welidev.github.io/scroll. It’s draft-00 in the request-for-comments sense of the phrase. It pins down the feed, the cursors, the streaming and polling modes, the checkpoints, tombstones, and retention, and it marks the places where my own confidence is lowest. It also doesn’t require waiting for providers, since a shim can synthesize a feed from any provider’s existing webhooks and list APIs, which is how I plan to find out where the design is wrong.
If you’ve lived in the valley, if you’ve written a dedup table or debugged a reconciliation cron or watched a delete evaporate, read it and tell me where it breaks. Disagreement is the desired response; silence is the failure mode.