Selaa lähdekoodia

fix(office): validate strict and corrupt OOXML

yudshj 2 viikkoa sitten
vanhempi
sitoutus
82ab876331

+ 2 - 2
.agents/notes/implemented/feature/2026-09-15-bundled-office-skills.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-09-15-bundled-office-skills.md
-2026-09-15-bundled-office-skills.md: e041e52cbfc3d4d33870fa3c8faf5fac3c0b35a4
-2026-09-15-bundled-office-skills.zh.md: 3f36a0e6883f08df66c6781cbf8784a6a4c358b0
+2026-09-15-bundled-office-skills.md: 089ca0326a7c53d4ee023608010efd49a8616db2
+2026-09-15-bundled-office-skills.zh.md: cd10aa863b1149594e9eb6b6a8e479511424f20d

+ 1 - 1
.agents/notes/implemented/feature/2026-09-15-bundled-office-skills.md

@@ -12,7 +12,7 @@ Office tasks need format-specific editing guidance and dependable file checks. R
 
 The [Office provider](../../../../packages/skill/skill-office/README.md) contributes three independently discoverable skills at the bundled rank. The default workflow uses `load_workspace_dependencies` and its Python executable; explicit user and AGENTS.md environment choices take precedence. A configurable absolute asset root lets Desktop expose Python-readable resources outside its application archive. Registration validates the required resources and YAML descriptions; loaded instructions exclude the metadata. Disposal removes every candidate.
 
-One standard-library checker validates ZIP/XML integrity and internal relationships, reports format-specific structure, and checks only explicit text or count assertions. DOCX table summaries count logical grid columns, including merged cells. Section geometry is reported rather than judged against the final section; font filenames do not establish glyph coverage. XLSX formula counts never imply recalculation. Text assertions follow section and note references and worksheet string indices, so retained headers, unused note definitions, comments, glossary entries, and unused strings cannot satisfy requested wording.
+One standard-library checker recognizes Transitional and Strict OOXML namespaces, validates ZIP/XML integrity and internal relationships, reports format-specific structure, and checks only explicit text or count assertions. Corrupt or encrypted ZIP members produce the same JSON package-failure report as other invalid documents. DOCX table summaries count logical grid columns, including merged cells. Section geometry is reported rather than judged against the final section; font filenames do not establish glyph coverage. XLSX formula counts never imply recalculation. Text assertions follow section and note references and worksheet string indices, so retained headers, unused note definitions, comments, glossary entries, and unused strings cannot satisfy requested wording.
 
 Desktop mounts the skill provider and runtime query independently of document rendering. Word uses python-docx, PowerPoint creation and editing use python-pptx, and Excel uses openpyxl and pandas. The managed payload and the ordinary creation examples require neither a rendering engine nor a separate presentation authoring library.
 

+ 1 - 1
.agents/notes/implemented/feature/2026-09-15-bundled-office-skills.zh.md

@@ -12,7 +12,7 @@ Office 任务需要针对文件格式的编辑指引和可靠的文件检查。
 
 [Office 提供方](../../../../packages/skill/skill-office/README.zh.md)以内置优先级提供三个可独立发现的 skill(技能)。默认工作流使用 `load_workspace_dependencies` 及其 Python 可执行文件;用户和 AGENTS.md 明确指定的环境优先。可配置的绝对资源根目录让 Desktop 在应用归档之外暴露 Python 可读取的资源。注册时验证必需资源与 YAML 描述;加载后的指令不包含元数据。卸载时移除全部候选项。
 
-一个仅依赖标准库的检查器验证 ZIP/XML 完整性与包内引用,报告各格式的结构,并只检查明确指定的文本或数量断言。DOCX 表格摘要计算逻辑网格列数,包含合并单元格。分节几何信息仅报告,不以最终分节判断全文;字体文件名不能证明字形覆盖范围。XLSX 公式数量不代表已执行重算。文本断言跟随分节与脚注/尾注引用及工作表字符串索引,因此保留的页眉、未使用的脚注/尾注定义、批注、词库条目和未使用的字符串不能满足措辞要求。
+一个仅依赖标准库的检查器识别 Transitional 与 Strict OOXML 命名空间,验证 ZIP/XML 完整性与包内引用,报告各格式的结构,并只检查明确指定的文本或数量断言。损坏或加密的 ZIP 成员与其他无效文档一样生成 JSON 包失败报告。DOCX 表格摘要计算逻辑网格列数,包含合并单元格。分节几何信息仅报告,不以最终分节判断全文;字体文件名不能证明字形覆盖范围。XLSX 公式数量不代表已执行重算。文本断言跟随分节与脚注/尾注引用及工作表字符串索引,因此保留的页眉、未使用的脚注/尾注定义、批注、词库条目和未使用的字符串不能满足措辞要求。
 
 Desktop 独立于文档渲染挂载技能提供者和运行时查询。Word 使用 python-docx,PowerPoint 创建和编辑使用 python-pptx,Excel 使用 openpyxl 和 pandas。受管理产物和普通创建示例都不要求渲染引擎或单独的演示文稿创作库。
 

+ 1 - 0
apps/desktop-host/src/office.ts

@@ -11,6 +11,7 @@ export const name = 'desktop-office'
 export interface Config {
   /** Bundled payload directory. Missing sibling `office-skills` resources fail Host startup. */
   readonly source: string
+  /** Harness-home directory where workspace dependencies are installed. */
   readonly root: string
 }
 

+ 2 - 2
docs/module-graph.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write docs/module-graph.md
-module-graph.md: 719273c14784fcf5dd98b2fddabfb7666c666139
-module-graph.zh.md: b1098cc39fdc1d44f9d7edc3555d18a2f2c26bc5
+module-graph.md: ae72394aa8ab8cf0d00d9d6cf707a17281292517
+module-graph.zh.md: 65fac7e5ccf176d65a2b600c8652820c95884788

+ 3 - 0
docs/module-graph.md

@@ -64,6 +64,7 @@ flowchart TD
     pkg_skill["skill"]
     pkg_skill_badge["skill-badge"]
     pkg_skill_filesystem["skill-filesystem"]
+    pkg_skill_office["skill-office"]
     pkg_tool_skill["tool-skill"]
   end
   subgraph group_subagent["packages/subagent"]
@@ -456,6 +457,7 @@ flowchart TD
   pkg_subprocess_local --> pkg_subprocess
   pkg_subprocess_local --> pkg_timeout
   pkg_skill_badge --> pkg_skill
+  pkg_skill_office --> pkg_skill
   pkg_spill --> pkg_brand
   pkg_spill --> pkg_llm
   pkg_spill --> pkg_session
@@ -1418,6 +1420,7 @@ flowchart TD
 | [`sandbox-windows-acl`](../packages/sandbox/sandbox-windows-acl) | `sandbox` | [`subprocess`](../packages/subprocess/subprocess) |
 | [`subprocess-local`](../packages/subprocess/subprocess-local) | `subprocess` | [`subprocess`](../packages/subprocess/subprocess), [`timeout`](../packages/util/timeout) |
 | [`skill-badge`](../packages/skill/skill-badge) | `skill` | [`skill`](../packages/skill/skill) |
+| [`skill-office`](../packages/skill/skill-office) | `skill` | [`skill`](../packages/skill/skill) |
 | [`spill`](../packages/spill/spill) | `spill` | [`brand`](../packages/util/brand), [`llm`](../packages/llm/llm), [`session`](../packages/core/session) |
 | [`app-boot`](../packages/boot/app-boot) | `boot` | [`home-paths`](../packages/util/home-paths), [`launch-environment`](../packages/util/launch-environment), [`system-prompt`](../packages/core/system-prompt) |
 | [`persona`](../packages/preset/persona) | `preset` | [`system-prompt`](../packages/core/system-prompt) |

+ 3 - 0
docs/module-graph.zh.md

@@ -66,6 +66,7 @@ flowchart TD
     pkg_skill["skill"]
     pkg_skill_badge["skill-badge"]
     pkg_skill_filesystem["skill-filesystem"]
+    pkg_skill_office["skill-office"]
     pkg_tool_skill["tool-skill"]
   end
   subgraph group_subagent["packages/subagent"]
@@ -458,6 +459,7 @@ flowchart TD
   pkg_subprocess_local --> pkg_subprocess
   pkg_subprocess_local --> pkg_timeout
   pkg_skill_badge --> pkg_skill
+  pkg_skill_office --> pkg_skill
   pkg_spill --> pkg_brand
   pkg_spill --> pkg_llm
   pkg_spill --> pkg_session
@@ -1420,6 +1422,7 @@ flowchart TD
 | [`sandbox-windows-acl`](../packages/sandbox/sandbox-windows-acl) | `sandbox` | [`subprocess`](../packages/subprocess/subprocess) |
 | [`subprocess-local`](../packages/subprocess/subprocess-local) | `subprocess` | [`subprocess`](../packages/subprocess/subprocess), [`timeout`](../packages/util/timeout) |
 | [`skill-badge`](../packages/skill/skill-badge) | `skill` | [`skill`](../packages/skill/skill) |
+| [`skill-office`](../packages/skill/skill-office) | `skill` | [`skill`](../packages/skill/skill) |
 | [`spill`](../packages/spill/spill) | `spill` | [`brand`](../packages/util/brand), [`llm`](../packages/llm/llm), [`session`](../packages/core/session) |
 | [`app-boot`](../packages/boot/app-boot) | `boot` | [`home-paths`](../packages/util/home-paths), [`launch-environment`](../packages/util/launch-environment), [`system-prompt`](../packages/core/system-prompt) |
 | [`persona`](../packages/preset/persona) | `preset` | [`system-prompt`](../packages/core/system-prompt) |

+ 2 - 2
docs/subsystems/skills.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write docs/subsystems/skills.md
-skills.md: 84165578d37c0d947f20435d9605bb685a5b673b
-skills.zh.md: 02fd41d5bc3a93ea4aa6de2a9996bc91bd648859
+skills.md: 7e45418b8b63955619b1d964e0cfc40ee7264632
+skills.zh.md: ead07c09fe7174d960a36173d842c652fa544c07

+ 2 - 2
docs/subsystems/skills.md

@@ -2,9 +2,9 @@
 
 English | [中文](skills.zh.md)
 
-The [skill capability family](../../packages/skill) includes the Service Definition ([dsh-skill](../../packages/skill/skill), `ctx.skills`), the local Service Provider ([dsh-skill-filesystem](../../packages/skill/skill-filesystem)), the optional packaged badge provider ([dsh-skill-badge](../../packages/skill/skill-badge)), and the Consumer ([dsh-tool-skill](../../packages/skill/tool-skill)). The registry merges provider catalogs across its host and per-scope layers; providers contribute local or packaged skills; the Consumer owns the initial and replacement catalogs plus the model-facing `skill` tool. Skills are optional instructions, not session events, so their vocabulary lives here rather than in [core.md](core.md).
+The [skill capability family](../../packages/skill) includes the Service Definition ([dsh-skill](../../packages/skill/skill), `ctx.skills`), the local Service Provider ([dsh-skill-filesystem](../../packages/skill/skill-filesystem)), optional packaged providers ([dsh-skill-badge](../../packages/skill/skill-badge) and [dsh-skill-office](../../packages/skill/skill-office)), and the Consumer ([dsh-tool-skill](../../packages/skill/tool-skill)). The registry merges provider catalogs across its host and per-scope layers; providers contribute local or packaged skills; the Consumer owns the initial and replacement catalogs plus the model-facing `skill` tool. Skills are optional instructions, not session events, so their vocabulary lives here rather than in [core.md](core.md).
 
-Source: [`packages/skill/skill/src/index.ts`](../../packages/skill/skill/src/index.ts), [`packages/skill/skill-filesystem/src/index.ts`](../../packages/skill/skill-filesystem/src/index.ts), [`packages/skill/skill-badge/src/index.ts`](../../packages/skill/skill-badge/src/index.ts), and [`packages/skill/tool-skill/src/index.ts`](../../packages/skill/tool-skill/src/index.ts).
+Source: [`packages/skill/skill/src/index.ts`](../../packages/skill/skill/src/index.ts), [`packages/skill/skill-filesystem/src/index.ts`](../../packages/skill/skill-filesystem/src/index.ts), [`packages/skill/skill-badge/src/index.ts`](../../packages/skill/skill-badge/src/index.ts), [`packages/skill/skill-office/src/index.ts`](../../packages/skill/skill-office/src/index.ts), and [`packages/skill/tool-skill/src/index.ts`](../../packages/skill/tool-skill/src/index.ts).
 
 ## Provider registry
 

+ 2 - 2
docs/subsystems/skills.zh.md

@@ -2,9 +2,9 @@
 
 [English](skills.md) | 中文
 
-[skill(技能)能力族](../../packages/skill) 包含 Service Definition([dsh-skill](../../packages/skill/skill),`ctx.skills`)、本地 Service Provider([dsh-skill-filesystem](../../packages/skill/skill-filesystem))、可选的随包徽章提供方([dsh-skill-badge](../../packages/skill/skill-badge))和 Consumer([dsh-tool-skill](../../packages/skill/tool-skill))。注册表在其宿主层与各 scope 层之间合并各提供方的目录;提供方贡献本地或随包 skill;Consumer 拥有初始目录和替换目录,以及面向模型的 `skill` 工具。skill 是可选的指令而非会话事件,因此其词汇定义在此处而非 [core.md](core.zh.md)。
+[skill(技能)能力族](../../packages/skill) 包含 Service Definition([dsh-skill](../../packages/skill/skill),`ctx.skills`)、本地 Service Provider([dsh-skill-filesystem](../../packages/skill/skill-filesystem))、可选的随包提供方([dsh-skill-badge](../../packages/skill/skill-badge) 与 [dsh-skill-office](../../packages/skill/skill-office))和 Consumer([dsh-tool-skill](../../packages/skill/tool-skill))。注册表在其宿主层与各 scope 层之间合并各提供方的目录;提供方贡献本地或随包 skill;Consumer 拥有初始目录和替换目录,以及面向模型的 `skill` 工具。skill 是可选的指令而非会话事件,因此其词汇定义在此处而非 [core.md](core.zh.md)。
 
-源码:[`packages/skill/skill/src/index.ts`](../../packages/skill/skill/src/index.ts)、[`packages/skill/skill-filesystem/src/index.ts`](../../packages/skill/skill-filesystem/src/index.ts)、[`packages/skill/skill-badge/src/index.ts`](../../packages/skill/skill-badge/src/index.ts) 与 [`packages/skill/tool-skill/src/index.ts`](../../packages/skill/tool-skill/src/index.ts)。
+源码:[`packages/skill/skill/src/index.ts`](../../packages/skill/skill/src/index.ts)、[`packages/skill/skill-filesystem/src/index.ts`](../../packages/skill/skill-filesystem/src/index.ts)、[`packages/skill/skill-badge/src/index.ts`](../../packages/skill/skill-badge/src/index.ts)、[`packages/skill/skill-office/src/index.ts`](../../packages/skill/skill-office/src/index.ts) 与 [`packages/skill/tool-skill/src/index.ts`](../../packages/skill/tool-skill/src/index.ts)。
 
 ## 提供方注册表
 

+ 2 - 2
packages/skill/skill-office/README.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write packages/skill/skill-office/README.md
-README.md: 7ed8f0c0d839db3c62718f5ce97a55875bd65ddc
-README.zh.md: 6b6e84bbd7db1a08a4c19c073759523b6cb4d1fd
+README.md: 41150b81daa4a38fcf5905f4d1e07d2745a7616f
+README.zh.md: 8b9342e137e526060fd072f0bfece904d2b7c1dc

+ 1 - 1
packages/skill/skill-office/README.md

@@ -41,7 +41,7 @@ Relative paths, missing resources, and skill files without a YAML frontmatter de
 
 ### Structural checks
 
-The shared Python checker reads DOCX, PPTX, or XLSX without modifying the source. It validates ZIP/XML and internal package relationships, reports document structure, and optionally checks required text or slide/sheet count. DOCX text assertions cover the main body, section-referenced headers and footers, and body-referenced footnotes and endnotes; comments, glossary text, and unreferenced parts or notes do not satisfy them. It uses only the Python standard library. Invalid packages, encrypted ZIP members, and report-file write failures produce a JSON failure report on stdout. A passing report does not establish appearance, feature preservation, or calculated formula results.
+The shared Python checker reads DOCX, PPTX, or XLSX without modifying the source. It recognizes Transitional and Strict OOXML namespaces, validates ZIP/XML and internal package relationships, reports document structure, and optionally checks required text or slide/sheet count. DOCX text assertions cover the main body, section-referenced headers and footers, and body-referenced footnotes and endnotes; comments, glossary text, and unreferenced parts or notes do not satisfy them. It uses only the Python standard library. Invalid packages, corrupt or encrypted ZIP members, and report-file write failures produce a JSON failure report on stdout. A passing report does not establish appearance, feature preservation, or calculated formula results.
 
 -----
 

+ 1 - 1
packages/skill/skill-office/README.zh.md

@@ -41,7 +41,7 @@ Agent(智能体)可以加载 Word、PowerPoint 和 Excel 工作流,默认
 
 ### 结构检查
 
-共享 Python 检查器读取 DOCX、PPTX 或 XLSX,不修改源文件。它验证 ZIP/XML 和包内引用关系,报告文档结构,并可检查必需文本或幻灯片/工作表数量。DOCX 文本断言覆盖正文、分节引用的页眉和页脚,以及正文引用的脚注和尾注;批注、词库文本与未引用的部件或脚注/尾注不能满足断言。它只使用 Python 标准库。无效包、加密 ZIP 成员和报告文件写入失败都会在标准输出中产生 JSON 失败报告。检查通过不代表外观、特性保留或公式计算结果已得到验证。
+共享 Python 检查器读取 DOCX、PPTX 或 XLSX,不修改源文件。它识别 Transitional 与 Strict OOXML 命名空间,验证 ZIP/XML 和包内引用关系,报告文档结构,并可检查必需文本或幻灯片/工作表数量。DOCX 文本断言覆盖正文、分节引用的页眉和页脚,以及正文引用的脚注和尾注;批注、词库文本与未引用的部件或脚注/尾注不能满足断言。它只使用 Python 标准库。无效包、损坏或加密的 ZIP 成员和报告文件写入失败都会在标准输出中产生 JSON 失败报告。检查通过不代表外观、特性保留或公式计算结果已得到验证。
 
 -----
 

+ 91 - 40
packages/skill/skill-office/assets/scripts/check_office.py

@@ -16,15 +16,31 @@ import json
 import posixpath
 import sys
 import zipfile
+import zlib
 from pathlib import Path
 from urllib.parse import unquote, urlsplit
 from xml.etree import ElementTree as ET
 
-W = "{http://schemas.openxmlformats.org/wordprocessingml/2006/main}"
-A = "{http://schemas.openxmlformats.org/drawingml/2006/main}"
-P = "{http://schemas.openxmlformats.org/presentationml/2006/main}"
-S = "{http://schemas.openxmlformats.org/spreadsheetml/2006/main}"
-R = "{http://schemas.openxmlformats.org/officeDocument/2006/relationships}"
+W_NAMESPACES = (
+    "http://schemas.openxmlformats.org/wordprocessingml/2006/main",
+    "http://purl.oclc.org/ooxml/wordprocessingml/main",
+)
+A_NAMESPACES = (
+    "http://schemas.openxmlformats.org/drawingml/2006/main",
+    "http://purl.oclc.org/ooxml/drawingml/main",
+)
+P_NAMESPACES = (
+    "http://schemas.openxmlformats.org/presentationml/2006/main",
+    "http://purl.oclc.org/ooxml/presentationml/main",
+)
+S_NAMESPACES = (
+    "http://schemas.openxmlformats.org/spreadsheetml/2006/main",
+    "http://purl.oclc.org/ooxml/spreadsheetml/main",
+)
+R_NAMESPACES = (
+    "http://schemas.openxmlformats.org/officeDocument/2006/relationships",
+    "http://purl.oclc.org/ooxml/officeDocument/relationships",
+)
 MAIN_PARTS = {
     ".docx": ("word/document.xml", "wordprocessingml.document.main+xml"),
     ".pptx": ("ppt/presentation.xml", "presentationml.presentation.main+xml"),
@@ -32,6 +48,34 @@ MAIN_PARTS = {
 }
 
 
+def namespace(root: ET.Element, supported: tuple[str, ...], part: str) -> str:
+    """Return the main XML namespace after checking its OOXML variant."""
+    uri = root.tag[1:].split("}", 1)[0] if root.tag.startswith("{") else ""
+    if uri not in supported:
+        raise ValueError(f"{part} uses unsupported XML namespace: {uri or '(none)'}")
+    return "{" + uri + "}"
+
+
+def relationship_id(node: ET.Element, part: str) -> str:
+    """Read an office-document relationship id from Transitional or Strict OOXML."""
+    for uri in R_NAMESPACES:
+        value = node.get("{" + uri + "}id")
+        if value is not None:
+            return value
+    raise ValueError(f"{part} has a reference without a relationship id")
+
+
+def relationship_types(kind: str) -> set[str]:
+    """Return the Transitional and Strict relationship type names for one role."""
+    return {f"{uri}/{kind}" for uri in R_NAMESPACES}
+
+
+def iter_namespaces(root: ET.Element, namespaces: tuple[str, ...], local_name: str):
+    """Iterate matching elements across Transitional and Strict namespaces."""
+    for uri in namespaces:
+        yield from root.iter("{" + uri + "}" + local_name)
+
+
 def relationship_target(part: str, target: str) -> str:
     """Resolve a package relationship without fetching external resources."""
     path = unquote(urlsplit(target).path)
@@ -62,78 +106,85 @@ def related_xml(part: str, reference: str, links: dict[str, str], xml: dict[str,
 def inspect_docx(xml: dict[str, ET.Element]) -> tuple[dict, str]:
     part = "word/document.xml"
     root = xml[part]
-    body = root.find(f"{W}body")
+    w = namespace(root, W_NAMESPACES, part)
+    body = root.find(f"{w}body")
     if body is None:
         raise ValueError("word/document.xml has no document body")
     tables = []
-    for table in body.iter(f"{W}tbl"):
-        grid = table.findall(f"{W}tblGrid/{W}gridCol")
-        rows = table.findall(f"{W}tr")
+    for table in body.iter(f"{w}tbl"):
+        grid = table.findall(f"{w}tblGrid/{w}gridCol")
+        rows = table.findall(f"{w}tr")
         # Merged cells span logical grid columns; counting physical cells loses them.
         columns = len(grid) if grid else max((sum(
-            int(cell.find(f"{W}tcPr/{W}gridSpan").get(f"{W}val", "1"))
-            if cell.find(f"{W}tcPr/{W}gridSpan") is not None else 1
-            for cell in row.findall(f"{W}tc")
+            int(cell.find(f"{w}tcPr/{w}gridSpan").get(f"{w}val", "1"))
+            if cell.find(f"{w}tcPr/{w}gridSpan") is not None else 1
+            for cell in row.findall(f"{w}tc")
         ) for row in rows), default=0)
         tables.append({"rows": len(rows), "columns": columns})
     sections = []
-    for section in body.iter(f"{W}sectPr"):
-        size = section.find(f"{W}pgSz")
-        margins = section.find(f"{W}pgMar")
+    for section in body.iter(f"{w}sectPr"):
+        size = section.find(f"{w}pgSz")
+        margins = section.find(f"{w}pgMar")
         sections.append({
-            "page_twips": {} if size is None else {key.removeprefix(W): value for key, value in size.attrib.items()},
-            "margins_twips": {} if margins is None else {key.removeprefix(W): value for key, value in margins.attrib.items()},
+            "page_twips": {} if size is None else {key.removeprefix(w): value for key, value in size.attrib.items()},
+            "margins_twips": {} if margins is None else {key.removeprefix(w): value for key, value in margins.attrib.items()},
         })
     text_parts = [body]
     for kind in ("header", "footer"):
-        links = relationships(part, xml, {f"{R[1:-1]}/{kind}"})
-        for section in body.iter(f"{W}sectPr"):
-            for reference in section.findall(f"{W}{kind}Reference"):
-                text_parts.append(related_xml(part, reference.attrib[f"{R}id"], links, xml))
+        links = relationships(part, xml, relationship_types(kind))
+        for section in body.iter(f"{w}sectPr"):
+            for reference in section.findall(f"{w}{kind}Reference"):
+                text_parts.append(related_xml(part, relationship_id(reference, part), links, xml))
     for kind in ("footnote", "endnote"):
-        references = {node.attrib[f"{W}id"] for node in body.iter(f"{W}{kind}Reference")}
+        references = {node.attrib[f"{w}id"] for node in body.iter(f"{w}{kind}Reference")}
         if not references:
             continue
-        links = relationships(part, xml, {f"{R[1:-1]}/{kind}s"})
+        links = relationships(part, xml, relationship_types(f"{kind}s"))
         for reference in links:
             tree = related_xml(part, reference, links, xml)
-            text_parts.extend(note for note in tree.findall(f"{W}{kind}") if note.get(f"{W}id") in references)
-    text = "\n".join("".join(node.text or "" for node in paragraph.iter(f"{W}t")) for tree in text_parts
-                     for paragraph in tree.iter(f"{W}p"))
-    return {"paragraphs": len(list(body.iter(f"{W}p"))), "tables": tables, "sections": sections}, text
+            text_parts.extend(note for note in tree.findall(f"{w}{kind}") if note.get(f"{w}id") in references)
+    text = "\n".join("".join(node.text or "" for node in paragraph.iter(f"{w}t")) for tree in text_parts
+                     for paragraph in tree.iter(f"{w}p"))
+    return {"paragraphs": len(list(body.iter(f"{w}p"))), "tables": tables, "sections": sections}, text
 
 
 def inspect_pptx(xml: dict[str, ET.Element]) -> tuple[dict, str]:
     part = "ppt/presentation.xml"
     links = relationships(part, xml)
-    slides = xml[part].findall(f"{P}sldIdLst/{P}sldId")
+    root = xml[part]
+    p = namespace(root, P_NAMESPACES, part)
+    slides = root.findall(f"{p}sldIdLst/{p}sldId")
     texts = []
     for slide in slides:
-        reference = slide.attrib[f"{R}id"]
+        reference = relationship_id(slide, part)
         tree = related_xml(part, reference, links, xml)
-        texts.append("\n".join("".join(node.text or "" for node in paragraph.iter(f"{A}t"))
-                               for paragraph in tree.iter(f"{A}p")))
+        texts.append("\n".join("".join(node.text or "" for node in iter_namespaces(paragraph, A_NAMESPACES, "t"))
+                               for paragraph in iter_namespaces(tree, A_NAMESPACES, "p")))
     return {"slides": len(slides)}, "\n".join(texts)
 
 
 def inspect_xlsx(xml: dict[str, ET.Element]) -> tuple[dict, str]:
     part = "xl/workbook.xml"
     links = relationships(part, xml)
+    root = xml[part]
+    s = namespace(root, S_NAMESPACES, part)
     sheets = []
     texts = []
     shared = xml.get("xl/sharedStrings.xml")
+    shared_s = s if shared is None else namespace(shared, S_NAMESPACES, "xl/sharedStrings.xml")
     shared_strings = [] if shared is None else [
-        "".join(node.text or "" for node in item.iter(f"{S}t")) for item in shared.iter(f"{S}si")
+        "".join(node.text or "" for node in item.iter(f"{shared_s}t")) for item in shared.iter(f"{shared_s}si")
     ]
-    for sheet in xml[part].findall(f"{S}sheets/{S}sheet"):
-        reference = sheet.attrib[f"{R}id"]
+    for sheet in root.findall(f"{s}sheets/{s}sheet"):
+        reference = relationship_id(sheet, part)
         tree = related_xml(part, reference, links, xml)
-        cells = list(tree.iter(f"{S}c"))
-        formulas = sum(cell.find(f"{S}f") is not None for cell in cells)
+        sheet_s = namespace(tree, S_NAMESPACES, links[reference])
+        cells = list(tree.iter(f"{sheet_s}c"))
+        formulas = sum(cell.find(f"{sheet_s}f") is not None for cell in cells)
         sheets.append({"name": sheet.attrib["name"], "cells": len(cells), "formulas": formulas})
         for cell in cells:
             if cell.get("t") == "s":
-                value = cell.findtext(f"{S}v", "")
+                value = cell.findtext(f"{sheet_s}v", "")
                 location = f"{links[reference]} cell {cell.get('r', '(no reference)')}"
                 try:
                     index = int(value)
@@ -142,8 +193,8 @@ def inspect_xlsx(xml: dict[str, ET.Element]) -> tuple[dict, str]:
                 if not 0 <= index < len(shared_strings):
                     raise ValueError(f"{location}: shared string index out of range: {index}")
                 texts.append(shared_strings[index])
-        texts.extend("".join(node.text or "" for node in cell.iter(f"{S}t")) for cell in cells)
-        texts.extend(cell.findtext(f"{S}v", "") for cell in cells if cell.get("t") == "str")
+        texts.extend("".join(node.text or "" for node in cell.iter(f"{sheet_s}t")) for cell in cells)
+        texts.extend(cell.findtext(f"{sheet_s}v", "") for cell in cells if cell.get("t") == "str")
         texts.append(sheet.attrib["name"])
     return {"sheets": sheets, "formulas_evaluated": False}, "\n".join(texts)
 
@@ -206,7 +257,7 @@ def main() -> int:
         if args.count is not None:
             actual = summary.get("slides", len(summary.get("sheets", [])))
             checks.append({"id": "count", "status": "pass" if actual == args.count else "fail", "expected": args.count, "actual": actual})
-    except (OSError, ValueError, KeyError, RuntimeError, ET.ParseError, zipfile.BadZipFile) as error:
+    except (OSError, ValueError, KeyError, RuntimeError, ET.ParseError, zipfile.BadZipFile, zlib.error) as error:
         checks.append({"id": "package", "status": "fail", "detail": str(error)})
     failed = any(check["status"] == "fail" for check in checks)
     report = {"format": args.input.suffix.lower()[1:], "verdict": "fail" if failed else "pass", "checks": checks, "summary": summary}

+ 52 - 2
packages/skill/skill-office/tests/check_office_test.py

@@ -14,6 +14,11 @@ P = "http://schemas.openxmlformats.org/presentationml/2006/main"
 A = "http://schemas.openxmlformats.org/drawingml/2006/main"
 S = "http://schemas.openxmlformats.org/spreadsheetml/2006/main"
 R = "http://schemas.openxmlformats.org/officeDocument/2006/relationships"
+STRICT_W = "http://purl.oclc.org/ooxml/wordprocessingml/main"
+STRICT_P = "http://purl.oclc.org/ooxml/presentationml/main"
+STRICT_A = "http://purl.oclc.org/ooxml/drawingml/main"
+STRICT_S = "http://purl.oclc.org/ooxml/spreadsheetml/main"
+STRICT_R = "http://purl.oclc.org/ooxml/officeDocument/relationships"
 PKG = "http://schemas.openxmlformats.org/package/2006/relationships"
 
 
@@ -23,14 +28,14 @@ class OfficeCheckTest(unittest.TestCase):
         self.addCleanup(self.directory.cleanup)
         self.root = Path(self.directory.name)
 
-    def package(self, suffix, parts):
+    def package(self, suffix, parts, compression=zipfile.ZIP_STORED):
         primary, mime = {
             "docx": ("word/document.xml", "wordprocessingml.document.main+xml"),
             "pptx": ("ppt/presentation.xml", "presentationml.presentation.main+xml"),
             "xlsx": ("xl/workbook.xml", "spreadsheetml.sheet.main+xml"),
         }[suffix]
         path = self.root / ("中文 document." + suffix)
-        with zipfile.ZipFile(path, "w") as archive:
+        with zipfile.ZipFile(path, "w", compression=compression) as archive:
             archive.writestr("[Content_Types].xml", '<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types">'
                              f'<Override PartName="/{primary}" ContentType="application/vnd.openxmlformats-officedocument.{mime}"/></Types>')
             for name, content in parts.items():
@@ -154,6 +159,33 @@ class OfficeCheckTest(unittest.TestCase):
         self.assertEqual(code, 0)
         self.assertEqual(report["summary"], {"sheets": [{"name": "Data", "cells": 3, "formulas": 1}], "formulas_evaluated": False})
 
+    def test_strict_ooxml_namespaces_are_inspected(self):
+        docx = self.package("docx", {
+            "word/document.xml": f'<w:document xmlns:w="{STRICT_W}"><w:body><w:p><w:r><w:t>Strict Word</w:t></w:r></w:p></w:body></w:document>',
+        })
+        self.assertEqual(self.run_check(docx, "--contains", "Strict Word")[0], 0)
+
+        pptx = self.package("pptx", {
+            "ppt/presentation.xml": f'<p:presentation xmlns:p="{STRICT_P}" xmlns:r="{STRICT_R}"><p:sldIdLst><p:sldId id="256" r:id="r1"/></p:sldIdLst></p:presentation>',
+            "ppt/_rels/presentation.xml.rels": f'<Relationships xmlns="{PKG}"><Relationship Id="r1" Target="slides/slide1.xml"/></Relationships>',
+            "ppt/slides/slide1.xml": f'<p:sld xmlns:p="{STRICT_P}" xmlns:a="{STRICT_A}"><a:p><a:r><a:t>Strict Slides</a:t></a:r></a:p></p:sld>',
+        })
+        self.assertEqual(self.run_check(pptx, "--contains", "Strict Slides", "--count", "1")[0], 0)
+
+        xlsx = self.package("xlsx", {
+            "xl/workbook.xml": f'''<workbook xmlns="{STRICT_S}" xmlns:r="{STRICT_R}"><sheets>
+              <sheet name="First" sheetId="1" r:id="r1"/><sheet name="Second" sheetId="2" r:id="r2"/>
+            </sheets></workbook>''',
+            "xl/_rels/workbook.xml.rels": f'''<Relationships xmlns="{PKG}">
+              <Relationship Id="r1" Target="worksheets/sheet1.xml"/><Relationship Id="r2" Target="worksheets/sheet2.xml"/>
+            </Relationships>''',
+            "xl/worksheets/sheet1.xml": f'<worksheet xmlns="{STRICT_S}"><sheetData><row r="1"><c r="A1" t="inlineStr"><is><t>Strict Sheet</t></is></c></row></sheetData></worksheet>',
+            "xl/worksheets/sheet2.xml": f'<worksheet xmlns="{STRICT_S}"><sheetData/></worksheet>',
+        })
+        code, report = self.run_check(xlsx, "--contains", "Strict Sheet", "--count", "2")
+        self.assertEqual(code, 0)
+        self.assertEqual([sheet["name"] for sheet in report["summary"]["sheets"]], ["First", "Second"])
+
     def test_xlsx_checks_only_cell_referenced_shared_strings(self):
         parts = {
             "xl/workbook.xml": f'<workbook xmlns="{S}" xmlns:r="{R}"><sheets><sheet name="Data" sheetId="1" r:id="r1"/></sheets></workbook>',
@@ -212,6 +244,24 @@ class OfficeCheckTest(unittest.TestCase):
         self.assertEqual(code, 1)
         self.assertIn("encrypted", report["checks"][0]["detail"])
 
+    def test_corrupt_deflate_member_returns_json_package_failure(self):
+        path = self.package("docx", {
+            "word/document.xml": f'<w:document xmlns:w="{W}"><w:body/></w:document>',
+        }, compression=zipfile.ZIP_DEFLATED)
+        with zipfile.ZipFile(path) as archive:
+            member = archive.getinfo("word/document.xml")
+        data = bytearray(path.read_bytes())
+        name_length = int.from_bytes(data[member.header_offset + 26:member.header_offset + 28], "little")
+        extra_length = int.from_bytes(data[member.header_offset + 28:member.header_offset + 30], "little")
+        compressed = member.header_offset + 30 + name_length + extra_length
+        data[compressed] = 0x07  # BTYPE=3 is reserved and invalid in a DEFLATE block.
+        path.write_bytes(data)
+        code, report = self.run_check(path)
+        self.assertEqual(code, 1)
+        self.assertEqual(report["checks"][0]["id"], "package")
+        self.assertEqual(report["checks"][0]["status"], "fail")
+        self.assertTrue(report["checks"][0]["detail"])
+
     def test_invalid_zip_and_xml_fail_and_output_cannot_overwrite_input(self):
         path = self.root / "broken.docx"
         path.write_bytes(b"not an Office archive")