if any guardrails fails, we suspend probes till cooldown.
also 1. every probe is bounded by hits/expiry time (whichever comes earlier) 2. hit budgeting happens with a token bucket at a global level, per probe was an overkill (numbers are configurable) 3. we even measure the execution time that probes have when active and suspend if that that takes longer than threshold (again configurable) 4. we even have budgets for the network bandwidth it would take (approximated by the size of payloads) 5. collection itself is bounded by max no of total snapshots we can keep in memory. 6. every snapshot has a size limit as well, every variable has a size limit as well. 7. depth of objects, no of objects, size of lists is capped by default.
latency delta varies by platform under load but is mostly negligible
nodejs: ~7-10ms python: ~4-9ms java: 1-2ms
the main reason for this is guardrails suspending probes, having loosened guardrails will increase this under load
regarding localization of failures.. absolutely we even report the error in the probe snapshot (confirmed by adding side effects in an expression and commenting out the guardrails during testing)
huge payload size doesnt matter.. we limit the objects depth, list length, remove duplicate refs from data etc.. even string length is truncated., but even if it happens, your request would still survive.
also, even if the collector dies or there's a network failure, your service remains unaffected, we just are unable to collect telemetry
we are boring under extreme conditions :)
1. What makes "read-only" a guarantee rather than a convention? In Python a plain attribute read can hit a @property that lazy-loads from the DB; in Java a getter can mutate state or take a lock. If the capture expression permits attribute access at all, read-only becomes a property of the code being probed rather than of your SDK. Do you restrict the expression grammar, or is it best-effort?
2. What's the shape of what comes back through MCP? A captured frame can serialize into something enormous, and an agent will cheerfully spend its entire context on one request object. Can you project at capture time (user.id rather than user), or does trimming happen after the full payload is already built?
For people that don't have these neat observability tools (like me), I've been using https://shellshare.net (disclaimer: I made it).
This is a single command to share a terminal live with e2e encryption. Originally it was for teaching classes or helping colleagues, but it's also very helpful for agents. I SSH into prod and run:
> npx shellshare exec --json -- tail /var/log/my-app.log
This generates a URL, then I can tell any agent:
> monitor <URL>, instructions in https://shellshare.net/llms.txt
They can see the output live. No need to install anything in the agent's machine. Next shellshare version it will be just "monitor <URL>" and the agent's instructions will be in the URL itself.
Nothing even near what you've guys done, but it has been helpful for me. Best of luck in your startup!
When I looked into this a while back I explored using ptrace() to add breakpoints and even add functions at specific line numbers. But ptrace is so slow, and it doesn't work with bytecode-in-VM setups.
What were some of the requirements you guys had when building HyperProbe? I can see low latency was one.
> Every log-and-trace tool hands the agent data that already exists and asks it to reason backward to what probably happened
If the app is using a decent instrumentation tool, the data shows what 'actually' happened, not what 'probably' happened.
> "checkout returns 200 but some users are seeing their order fail, find out why."
Does this tool only exist to shore up poor system design? Failing orders at any e-commerce business I've worked with, large and small, are a huge red flag. Typically that is one of the first actions that is logged and traced (alongside onboarding/login), and the metrics are actively monitored. Returning 200 for failure and not catching that error is very bad API design.
Similarly, putting engineers in a situation where debugging requires accessing unknown amounts of live sensitive customer data is generally considered bad practice (even if it happens often IRL) -- in a hurry to debug, it's easy to miss that a property should have been redacted; by then it's too late and sensitive data is exposed. Plus, in most systems with significant usage the volume of trace data is prohibitive to individually examine and search through. That's why Rollbar etc aggregate errors and captured data to identify patterns before a human (or agent, or tool) ever takes a look at it. A single captured instance can also be very misleading as to the true cause.
How are you addressing these common concerns?
How does it work? Using the NodeJs inspector API or other language equivalent to drop breakpoints? Those APIs are unavailable in many serverless environments and are challenging to use alongside bundlers.
If you don’t know how it broke, and you don’t know how you fixed it, what exactly is it you think you understand about your application?
Consider putting this near the beginning rather than 2/3 of the way down your pitch. I nearly stopped reading because these dramatic 1-2 sentence paragraphs are unpleasantly like listening to TV commercials. I think your target audience should not be CTOs or their direct reports, but engineers themselves, and I think you need a more focused pitch that takes less time to get to the point.
Anyway, an MCP-managed passive debugger seems like a useful tool. Best of luck with it.
A useful adjunct to this kind of production in-memory debugging is a read-only agent locked down role in AWS or equivalent. I believe amazon has just set up some kind of a wizard for configuring a role like this. It really gives agents the ability to relatively safely look at the prod setup without exposing sensitive details or making changes, especially with e.g. terraform
Running an autonomous pipeline for eight months, the three incidents that cost me the most days all had the surface error naming the wrong subsystem:
- "x264: malloc of size N failed / incorrect parameters" — I read it as a codec or bad-args bug and went looking there. It was RAM exhaustion. The encoder was the victim, not the cause.
- A 22x slowdown in an LLM step that was indistinguishable from a hang. It was swap: the model no longer fit in RAM, and the page file did the rest.
- A 27-minute "freeze" in a background job. The process was healthy; the pipe was buffering, so nothing appeared until exit.
In all three the logs were complete and the metrics were green. The mistake was in the inference drawn from them — and an agent will produce that wrong inference far faster than I did, with better prose attached to it.
So: does HyperProbe ever return "I don't know — here are two competing hypotheses and the cheapest check that separates them"? The discriminating check is the part I'd pay for. A single confident answer that's wrong is worse than no answer, because it sends a human down a road with the agent's credibility behind it.
we work at the application layer by hooking into production grade tooling if available (inspector in v8, sys.monitoring in python) or bytecode manipulation(jvm)
since we arent controlling the application for a different process, we dont need to freeze the app to get the current app state
requirements we had in mind in order of importance
1. safety -> user app needs to function as usual no matter what happens, there shouldnt be an error in the user's app because of us
2. zero idle footprint -> if no probe is active, cpu/memory differency in the user app should be immeasurable
3. zero latency footprint at non probe paths while other probes are active
4. measure mem/cpu footprint directly or via a proxy like eventloop lag and have guardrails around it. suspend probes or even lose snapshot data if guardrail conditions meet
5. minimal mem/cpu footprint for active probes
6. minimal latency foot print for active probe paths
These work only on either uncaught exceptions or wrapping up caught exceptions with their sdk. These tools will not help you with silent failures, like logic bugs where code executes cleanly without throwing, but produces the wrong business state. If every problem in your app ends up as an exception, sure you'll be able to catch the symptoms of where the exception got thrown. we can deal with these too, but these tools cant deal with the messy bugs where no exception fires.
> Such very mature tools exist that auto-instrument, collect variables from the call stack, and pinpoint error causes.
That is true for python using frame.f_locals (we use this as well)
nodejs only gives it only till the lasy async boundary, after that v8 itself drops this data. java only gives you the current frame, to get variables beyond that you would needs JDI/JVMTI which would block your threads, usually unnacceptable in production
To get around this safely, we add multiple probes all across the call chain and collate collected data using the traceId from the context (or thread id as a fallback);
> Does this tool only exist to shore up poor system design?
Returning 200 OK on a silent failure is 100% bad system design, I completely agree. But real-world production systems are full of legacy edge cases. (if that weren't true, L1/L2/L3 support team shenanigans wouldn't exist)
Also, the exception will tell you that an exception occured in order service in GET /orders/{id}/payment, your trace will tell you payment service is giving 404 for that order ID
what it wont tell you it happened becuase the webhook endpoint that your payment gateway calls is now receiving a new payment state called 'PENDING' and that you dont handle but still mark the payment as 'processed' for idempotency check. and now your order service is calling the payment service and its giving 404 because it never got written
Bad design. 100% Agree, but has happened IRL.
> putting engineers in a situation where debugging requires accessing unknown amounts of live sensitive customer data is generally considered bad practice (even if it happens often IRL)
I think tells that teams would go to these extents to fix issues. Not ideal. I agree.
> in a hurry to debug, it's easy to miss that a property should have been redacted; by then it's too late and sensitive data is exposed
fair critique. we currently use in-process rule engines to filter known sensitive patterns, and users can add on to it. but we are also building out-of-process secondary checks (using NER/classifiers) to sanitize payloads before storage. It requires strict rules, but getting verified runtime evidence is far safer and faster than blindly guessing and shipping trial-and-error hotfixes to production. or waiting to be too sure.. a luxury that might not be possible everytime.
> Rollbar etc aggregate errors and captured data to identify patterns before a human (or agent, or tool) ever takes a look at it.
There is merit in that as well, if you are looking at so many logs/traces, you kinda have to do it. We have a different approach, we use hypothesis driven conditional probing instead. probes are dropped dynamically as the understanding of the bug evolves in a session
exmaple:
console.log('hello');
const x = await getThisValueSomehow();
if (condition A) {
console.log('i m in condition A');
// do something;
} else if (condtion B) {
console.log('i m in condition B');
// do something;
}
You can also place a probe before the branch to capture variable state when neither condition evaluates to true. You gather precise data on demand rather than paying to store petabytes of static trace data.
> A single captured instance can also be very misleading as to the true cause.
We collect multiple snapshots per probe run. However, because we capture full variable state at the exact execution line, a single snapshot frequently reveals the root cause for that specific failure path. If that snapshot raises new questions, you/your agent simply drops more probes deeper down the call chain
Thanks! This was very insightful
You're correct inpector API is not available in many non-v8 targets. Bun also has somewhat of a partial support for inpector API but at least has a programmable debugger interface. It's not going to be as fast as native inspector but its better than nothing i guess :P
for python sys.monitoring. for JVM, we do bytecode manipulation itself.
bundlers are not an issue because we support sourcemaps. We just need mappings, not code in the sourcemaps and we do sourcemap resolutions out of process so that your app doesnt spend ~200 MB of memory for parsing sourcemaps
This is readonly and safe by default.
expressions come into play for conditonal probes
read only safety guarantees here depend on the runtime
NodeJS: handled implicitly by using `throwOnSideEffect: true` any side possible effects are prevented using this
Python and Java: As of now, we don't let conditions have method invocations at all and only allow a subset of comparator operators. no assignment allowed
usually property can invoke getter which usually should be safe to execute by design, but since we cant guarantee how it would have been written, we dont allow that as well for now.
order.total > 50 => not allowed
total > 50 => allowed
to get around this we use multiple probes, agrregated by the current context's traceId (if avaliable)
we plan to eliminate this problem by adding a custom DSL + AST parsing which can act as the policy layer to dissallow condtional probes
Audit trail is in our roadmap. As of now, you can delete the data that's collected by probes. the only problem we have with audit trail is what if you capture something sensitive and that remains in your audit trail.. so we need some immutability that registers audit trails.. but then have enough flexibility to remove the data collected.. can be done
Although, our primary sell is debugging, the context from production on how things work currently helps ai agents during feature development and code reviews as well.
But when the pipeline fails (bugs happen that's fine) re-running the exact same process may not be the solution.
Where does the additional intelligence that wasn't there before come from? We ran the pipeline that got it to prod on the same exact models you have access to. So the value prop is that you read the logs automatically instead of a developer directing a debug session?
trying to debug using our tool might even lookup some memleak candidates in your primary container, but there wont be conclusive evidence for it and it would say so.
for in app errors all we do is hypothesize and either prove/disprove that using data from running system.
and whenever we do report something we give have the evidence for it. its not fool proof but just asking does this hypothesis gets proved with this evidence in a subagent mostly does the trick
You're correct, serverless is a bit tricky. CPU gets suspended the moment your function returns. The way it works is that you wrap your functions with a wrapper in our sdk.
that wrapper is supposed to track if there's telemetry to be sent, if so.. it sends it, otherwise, return as usual
this makes sure that when there's no active probe, there's no latency added. But when there's an active probe.. ~100-200ms could be added in the worst case if the probe is just before the return.
again, this isnt a problem in non serverless worloads because the CPU is always on.
but since probes are bounded by time and count, this will go away as soon as the time or count condition meets. beats adding new logs and redeploying in my opinion
It could work, the technology isnt the limitation.
But we were clear from day one that we cant let our sdks change the memory. Even if it helped solve a real problem.. say for example resetting a bad env variable or a feature flag without redeployment.
I might be biased from my experience, but i would prefer having a bug in my system for longer that i can reliably reason with than having it solved dynamically within the app which adds another thing to keep in my mind.
for me, bug -> fails -> good bug + dynamic patch -> works -> bad
also we dont think that we ourselves wont have any downtime ever, so we design for it. we'd not want to become as critical for your app as say your database.
As of now, your app works even if our servers are down/blocked/slow, adding the ability to change memory on the fly could change this
what we wanted to convey is that sometimes people confuse "the symptom went away" with "the root cause was fixed"
I have seen that a rollback, a quick redeploy, or a temporary drop in tenant load makes the alerts go away and issue is considered resolved. specially true for larger teams with many engineers and services
a real example: a dev got OOMed after a release that coincided with a flash sale. he increased memory limits, and containers stopped crashing and it was "fixed". Actualy, a newly introduced internal module had a memory leak. adding RAM just hid the leak until the next traffic spike.
hyperprobe exists to capture actual in-memory runtime state during live traffic so you can prove the root cause before changing code or scaling infra in this case
Also thanks for the candid feedback! (And fair call on the design — we definitely prioritized shipping core functionality over UI polish, but point taken on the orange/brown palette, we'll change it)
your coding agent will have access to your code and can see what logs you have enabled, if indeed that would help in debugging, going that route helps. it can take your agents a bunch of retries but it might get there if the answer is in logs
what we provide your coding agent is a detailed snapshot of all your variables at any line it feels would help debugging. and not just in the current call frame.. even the variables of the callers of your current function, like a debugger.
suupose funcA() -> funcB() -> funcC() -> yourCurrentFn()
we'll provide all the variables that were set in all 4 functions to your agent. debugging using this would be a lot more accurate and you just one snapshot like this instead of looking at a thousand log lines to understand why something is not working the way you want to.
this kind of data is missing from your logs and and even your traces because it will be impractical for privacy and performance.
Are you saying that hyperprobe would have in fact caught that issue?
The class I never solved sits one level below that: a check that runs, passes, and is looking at the wrong object. My top-level health signal was green for three days while zero artifacts shipped. Sixteen daemons alive, backend responding, auth token valid — every organ it polled was genuinely healthy, and nothing measured the thing leaving the building. A missing k8s connector would not have helped; the data was all there and all correct.
Related one from the same eight months: 29 quality gates, written and unit-tested and committed, none of which was ever called, because nothing was a runner. The tests proved the gates worked. Nothing proved they were wired.
Do you see that shape at customers — monitoring correct, conclusion still wrong because it describes the process instead of the result? I ask because it decides what your diagnosis agent should be sceptical of: if the input signals can be individually true and jointly meaningless, evidence-checking inside the hypothesis does not catch it.
How does the probe function technically. Inspector API I believe is unavailable on CloudFlare etc.
I once wrote something like this which could work on serverless platforms without the Inspector API. It used Typescript AST transforms to insert no-op listeners at every line, so they would dynamically eval or dump breakpoint style if a listenToLine parameter equalled their line, otherwise no-op. So trivial but not technically zero runtime cost.
makes no difference if its cold or warm start. (The latency because of us, not latency in general)
Using AST, that's clever actually! I thought along somewhat similar lines. User tree-sitter. But it needs a build step! not sure how people feel about that :D It changes your source code itself so we'd need a 2 tiered source resolution. Dirty, but doable.
And instead of having this at every line, i did this at "lines of interest" before and after every scope ends.
so at the start/end of an if condition, start/end of fn definiton. it sort of worked, but it slowed down our synthetic benchmarks for "no effect when probes arent there" by more than what i wanted to tolerate and it depended of eval which i thought devs wont accept. and using node-vm slowed it further
But will give another try again. thanks for sharing this!
our snapshots give evidence, whether or not its helpful is determined by the engineer/ai agent
I'm not saying every issue would be diagnosed this way, sometimes, it might well be beyond your control like a VM on a noisy neighbour hogging shared CPU
we might get things wrong as well.
If you are transforming anyway you're looking at virtually zero dev time cost and runtime cost when no probe is active of less than 1 ms.
This approach survives any environment I know of and has almost zero runtime cost.
Y
Backed by Y Combinator · AI ON-CALL AGENT
02:47 AM · order-service · 847 failures / 10 min · @priya paged
TRIGGERED
Every hour they spend in a war room is an hour they're not building. HyperProbe works the incident for them, alert to confirmed root cause before they've opened their laptop.
Node.js · TypeScript · Java · Python · Works with Cursor, Claude Code, Codex, Opencode
01
Incidents don't just break production. They break your roadmap. Your best engineers become your on-call team, every hour spent debugging is an hour not building.
02
The hotfix was an educated guess. Nobody confirmed what actually caused it. If the same conditions appear next week, the same incident fires.
03
The incident costs the same every minute it stays open. The fix takes mins. Finding takes hours, because the value that explains failure is never logged.
The solution
HyperProbe makes your coding agents drop a read-only probe on the exact line where the problem happened in prod. It captures data your logs do not have, without redeployment or restarting the service.
Every other tool reasons hard over data you already have. HyperProbe captures exact evidence.
3 to 4 hrs→<10 minTime to root cause
3 to 4→0Redeployments per incident
2 to 3→0Senior engineers on the investigation
"Sync issues used to take us days to reproduce locally. HyperProbe caught the silent data mismatch in production on the first attempt."
"During peak traffic, our listing service was black-boxing failures. HyperProbe let us inspect the live memory state during the spike. We fixed the race condition in the same hour."
Bhagwan Bansal
SDE, Housing.com
How it works
01
Alert
Picks up the page from PagerDuty, Datadog, or Slack automatically.
02
Plan
Reads logs and traces, to automatically locate the file, line with the issue, and plan debugging flow.
03
Probe
Logs not enough? Places a read-only virtual breakpoint on the suspect line. No redeploy.
04
Capture
Breakpoint fires on live traffic. Exact variable state captured at that line.
05
Confirm
Diagnosis verified against real evidence. Confirmed RCA delivered.
What is a probe?
A probe is a read-only, non-blocking snapshot of the live variable state at a specific line in your running service. It fires on real traffic, captures the exact values at that moment, and disappears after capture. Your service never pauses. Zero user impact.
The agent captures state. It cannot write memory or execute code. Every probe is logged in an immutable audit trail. Approval-gated until you trust it.
Self-hosted or private VPC. Nothing leaves your environment. PII redacted at the agent before capture. Your security team defines what can be observed.
The breakpoint fires asynchronously. Requests complete at full speed. Users experience nothing. Less than 1% overhead at 3,000 RPS.
What we cover
No exception. No alert. HyperProbe shines even with problems hardest to find.
Returns 200 with the wrong body. The trace is green. The value was never logged.
Stack trace names line 82. The cause is at line 18, or in a different file.
Exception caught and swallowed. No alert. No error. The business metric just moves.
Needs thread state at the exact moment of overlap. Nothing logs that.
Vendor added a new field or status value. Your parser has no case for it.
Payments failing, orders dropping. No exception anywhere in the stack.
Shipping this month: Memory leak diagnosis · OOM root cause · CPU spike isolation · Latency spike tracing
One real incident, start to finish
Not a feature walkthrough. This is exactly what happens when HyperProbe works an incident on your behalf.
02:47 AMAlert fires
PagerDuty fires. GET /api/orders/{id}/status is returning 500 for nearly a quarter of requests. 847 failures in the last 10 minutes. No exception in the logs.
PagerDuty alertHIGH ERROR RATE · order-service
GET /api/orders/{id}/status · 500 · 23% error rate
847 failures / 10 min · threshold exceeded
02:48 AMScouting
HyperProbe reads the distributed traces and follows the failure chain. Order service is healthy. Payment service downstream is returning 404. Payments exist in the payment gateway but are not in the system.
Trace for failing requestGET /api/orders/{id}/status 500
|
└── GET payment-service/api/getPaymentsByOrder/{orderId} 404
Payments exist in the payment gateway. Not found in the system. A write failed silently somewhere upstream.
02:49 AMProbe placed
HyperProbe identifies payments are recorded when the payment gateway calls a webhook. A virtual breakpoint placed on the webhook handler at /src/api/webhooks.ts line 78. No redeploy. Service keeps running.
Probe activatedPOST /api/webhooks/payments
file: /src/api/webhooks.ts · line 78
Non-blocking · Read-only · No redeploy
02:50 AMBug found
Gateway is sending PENDING. The code has no case for it. Idempotency check marks payment as processed before confirming state. Payment never written to DB. No exception fires.
Live snapshot · webhooks.ts:78 · captured 02:50:14 UTCstatus = "PENDING" ← payment gateway sending this, no handler exists
duplicate = null ← first time seen, passes through
db.insert → never called
redis.set → called anyway, payment locked out permanently
Payment gateway started sending PENDING, a status your code never handled. Idempotency key written before state is checked. Payment marked processed, never recorded.
02:52 AMFix suggested
Payment gateway started sending PENDING, a status your code never handled. Idempotency key written before state is checked. Payment marked processed, never recorded.
Before and after
Your best engineers should not be your on-call team. HyperProbe handles the investigation so they can go back to building.
Without HyperProbe
02:47 AM
Alert fires. Engineer paged.
02:50 AM
Stack trace points to line 82. The variable that caused it was set several frames up, in a different file. No log captures it there.
03:10 AM
Frame located. Variable value not visible. Adds a log line to capture it.
03:40 AM
CI/CD deploys. 30 minutes gone. Waiting for the condition to reproduce in production.
04:15 AM
First log visible. Partial data. Not enough. Another log line. Another 30-minute deploy cycle.
05:20 AM
After 2 to 3 redeploy cycles, root cause confirmed. 2 hours 33 minutes.
With HyperProbe
02:47 AM
Alert fires. HyperProbe picks it up.
02:48 AM
HyperProbe uses your coding agent to locate the exact frame where the probe should go. All in background.
02:49 AM
HyperProbe activates a virtual breakpoint at that exact line. No redeploy.
02:53 AM
Breakpoint fires safely at next request. Exact variable value captured. Service keeps running.
02:56 AM
Root cause confirmed. 9 minutes from alert to evidence-backed diagnosis.
03:00 AM
Engineer commits the fix.
Pricing
Probes and captures are unlimited on every plan. You should never hit a wall in the middle of an incident.
Free
$0 forever
1 service · managed cloud
Install the SDK and see a real capture the same afternoon.
Most teams
Professional
$99 per service per month
$79 billed annually · 3 service minimum
For teams running real production traffic who want the whole stack instrumented, not one service.
Enterprise
Custom
Annual contract · volume pricing
For teams whose security review has to sign off before anything touches production.
First incident we work with you is free · Cancel any month · No seat counts · No host counts · No capture limits
See full pricing and what is in each plan →
This is built for you if
If you don't, we're probably not the right fit yet. If you do, let's talk.
If you have a specific memory of that night, HyperProbe is built for you. The data you needed was never in the logs. You grepped, guessed, redeployed, and hoped. It does not have to work that way.
When prod breaks, your most expensive hire gets paged to do work a machine should do. Every night on-call is a night they resent. And your best people have options.
Without the exact variable state at the moment of failure, every fix is a guess. Guesses hold until they don't. HyperProbe confirms root cause so the fix is final, not deferred.
Start the POC
30 minutes. Your service. Your incident. You'll see a confirmed root cause before the call ends — or there's nothing more to discuss.
Node.js · TypeScript · Java · Kotlin · Runs in your own infra · Up in 15 minutes