Jelajahi Sumber

docs(session): align SQLite schema evidence

Tianyi Cui 1 bulan lalu
induk
melakukan
0d551b9f5c

+ 2 - 2
.agents/notes/implemented/architecture/2026-08-18-sqlite-physical-chunk-row-compression.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-08-18-sqlite-physical-chunk-row-compression.md
-2026-08-18-sqlite-physical-chunk-row-compression.md: 6892c362fa0b81853b393b7c2fa00e5356ce768c
-2026-08-18-sqlite-physical-chunk-row-compression.zh.md: e8d03f9f1ab553cd68d02275463ce1daeb755f26
+2026-08-18-sqlite-physical-chunk-row-compression.md: 34aac2f183d386ffe22f86a6b62fe5e3105b3dfa
+2026-08-18-sqlite-physical-chunk-row-compression.zh.md: 1845185d543f565b55ace6adac973dad5535ad7b

+ 1 - 1
.agents/notes/implemented/architecture/2026-08-18-sqlite-physical-chunk-row-compression.md

@@ -58,7 +58,7 @@ The repository regression guard writes 1,000 streamed deltas in 40-event durable
 
 **Compress every payload.** Rejected because small independent Zstandard frames add headers and synchronous CPU work while losing the cross-record dictionary opportunity of a whole-file stream. On the 105-session comparison corpus, a threshold sweep produced 75.01 MB at 4 KiB, versus 93.87 MB at 16 KiB and 60.92 MB at 1 KiB. The writer fixes level 3 rather than inheriting a library default, matching the moderate level used by [Codex cold-rollout compression](https://github.com/openai/codex/blob/main/codex-rs/rollout/src/compression.rs) while retaining independent row access.
 
-The final frozen comparison used 105 sessions, 2,507,860 logical events, 512-event durable batches, three independent builds per backend, and three read passes per build. SQLite used 75.01 MB, wrote in 8.58 s, read complete sessions at 3.95/21.58 ms p50/p95, read 50-event tails at 0.253/0.378 ms, and forked every session in 13.10 s. Zstandard JSONL used 30.65 MB and measured 28.21 s, 4.49/23.36 ms, 10.58/80.90 ms, and 14.48 s. The predecessor scalar SQLite layout used 709.57 MB and measured 10.64 s, 9.02/69.16 ms, 0.189/0.293 ms, and 19.30 s. The packed layout is 89.4% smaller than the predecessor, writes 19.4% faster, improves complete-read p50/p95 by 56.2%/68.8%, and reduces 2,507,860 physical event rows to 65,810. Scalar tail-50 and list micro-latency are lower, but the packed provider remains materially faster than JSONL on those paths and wins the dominant size, write, full-read, and fork costs. The 4 KiB threshold is the accepted balance rather than a strict dominance claim.
+The final frozen comparison used 105 sessions, 2,507,860 logical events, 512-event durable batches, three independent builds per backend, and three read passes per build. SQLite used 75.01 MB, wrote in 8.58 s, read complete sessions at 3.95/21.58 ms p50/p95, read 50-event tails at 0.253/0.378 ms, and forked every session in 13.10 s. Zstandard JSONL used 30.65 MB and measured 28.21 s, 4.49/23.36 ms, 10.58/80.90 ms, and 14.48 s. The predecessor scalar SQLite layout used 709.57 MB and measured 10.64 s, 9.02/69.16 ms, 0.189/0.293 ms, and 19.30 s. The packed layout is 89.4% smaller than the predecessor, writes 19.4% faster, improves complete-read p50/p95 by 56.2%/68.8%, and reduces 2,507,860 physical event rows to 65,810. Scalar tail-50 and list micro-latency are lower, but the packed provider remains materially faster than JSONL on those paths and wins the dominant size, write, full-read, and fork costs. The 4 KiB threshold is the accepted balance rather than a strict dominance claim. This comparison measured schema 17; schema 18 retains the chunk codec and bounds but changes the row discriminator, so the exact size and timing values remain schema-17 evidence until schema 18 is remeasured.
 
 **Store packed payloads under the logical `assistant/chunk` type.** Rejected because payload heuristics make malformed rows ambiguous and couple physical decoding to future logical payload fields. Explicit tags fail loudly.
 

+ 1 - 1
.agents/notes/implemented/architecture/2026-08-18-sqlite-physical-chunk-row-compression.zh.md

@@ -58,7 +58,7 @@ SQLite 在 schema 18 包内拥有分片编码和验证。字段完全匹配的
 
 **压缩每个 payload。** 不予采用,因为小型独立 Zstandard frame 会增加 header 和同步 CPU 工作,也无法利用整文件流的跨记录字典。在 105 个会话的对比语料上,阈值扫描结果为:4 KiB 生成 75.01 MB,16 KiB 为 93.87 MB,1 KiB 为 60.92 MB。写入方固定使用 level 3,而不是继承库默认值;这与 [Codex 冷 rollout 压缩](https://github.com/openai/codex/blob/main/codex-rs/rollout/src/compression.rs)所用的适中级别一致,同时保留独立行访问。
 
-最终冻结对比包含 105 个会话、2,507,860 个逻辑事件,以 512 个事件为持久批次;每个后端独立构建三次,每次构建执行三轮读取。SQLite 使用 75.01 MB,写入耗时 8.58 秒,完整读取 p50/p95 为 3.95/21.58 毫秒,读取最后 50 个事件为 0.253/0.378 毫秒,对所有会话执行 fork 为 13.10 秒。Zstandard JSONL 使用 30.65 MB,对应指标为 28.21 秒、4.49/23.36 毫秒、10.58/80.90 毫秒和 14.48 秒。此前的标量 SQLite 布局使用 709.57 MB,对应指标为 10.64 秒、9.02/69.16 毫秒、0.189/0.293 毫秒和 19.30 秒。打包布局比此前布局小 89.4%,写入快 19.4%,完整读取 p50/p95 改善 56.2%/68.8%,并把 2,507,860 个物理事件行减少到 65,810 行。标量布局的最后 50 个事件读取与 list 微延迟更低,但打包提供方在这些路径上仍明显快于 JSONL,并改善主要的空间、写入、完整读取和 fork 成本。4 KiB 阈值是接受的平衡点,而不是严格支配所有指标的结论。
+最终冻结对比包含 105 个会话、2,507,860 个逻辑事件,以 512 个事件为持久批次;每个后端独立构建三次,每次构建执行三轮读取。SQLite 使用 75.01 MB,写入耗时 8.58 秒,完整读取 p50/p95 为 3.95/21.58 毫秒,读取最后 50 个事件为 0.253/0.378 毫秒,对所有会话执行 fork 为 13.10 秒。Zstandard JSONL 使用 30.65 MB,对应指标为 28.21 秒、4.49/23.36 毫秒、10.58/80.90 毫秒和 14.48 秒。此前的标量 SQLite 布局使用 709.57 MB,对应指标为 10.64 秒、9.02/69.16 毫秒、0.189/0.293 毫秒和 19.30 秒。打包布局比此前布局小 89.4%,写入快 19.4%,完整读取 p50/p95 改善 56.2%/68.8%,并把 2,507,860 个物理事件行减少到 65,810 行。标量布局的最后 50 个事件读取与 list 微延迟更低,但打包提供方在这些路径上仍明显快于 JSONL,并改善主要的空间、写入、完整读取和 fork 成本。4 KiB 阈值是接受的平衡点,而不是严格支配所有指标的结论。该对比测量 schema 17;schema 18 保留分片 codec 与上限,但改变行判别值,因此在重新测量 schema 18 前,精确的大小与时延值仍是 schema 17 证据。
 
 **把打包 payload 存在逻辑 `assistant/chunk` 类型下。** 不予采用,因为 payload 启发式判断会使畸形行产生歧义,并把物理解码耦合到未来逻辑 payload 字段。显式标签会明确失败。
 

+ 2 - 2
docs/subsystems/persistence.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write docs/subsystems/persistence.md
-persistence.md: 5da8e11f8b4932c732db24d2de996f0419cb3ff0
-persistence.zh.md: d69c551bfcb994e1e16e04515ad84af175bf5cfd
+persistence.md: 402836cb727fb99d92cea5e2a0d242b010aad2c9
+persistence.zh.md: 424fc928d5a8ef18b403ea31e5a5d26d3fd3fdc7

+ 1 - 1
docs/subsystems/persistence.md

@@ -233,7 +233,7 @@ interface SessionPersistenceSnapshot {
 All implement the same abstract `SessionPersistence` (locate/create/append/prepare/load/inspect/readFrom/list/listSnapshots over `SessionEvent`, with optional cancellation on observation methods) and pass the shared `runPersistenceContract` suite:
 
 - **[dsh-session-persistence-jsonl](../../packages/session/session-persistence-jsonl)** — an append-only logical JSONL log per session, stored as checksummed concatenated Zstandard frames by default or raw lines by configuration, with crash-safe atomic writes, interrupted-turn recovery, and a read/replay path.
-- **[dsh-session-persistence-sqlite](../../packages/session/session-persistence-sqlite)** — an opt-in `node:sqlite` backend using schema 17 to store exact same-block delta runs in bounded physical `text-chunks`, `reasoning-chunks`, and `tool-call-chunks` rows. It reconstructs the complete logical event stream before returning it, packs only newly durable batches, and rejects older schemas rather than migrating them.
+- **[dsh-session-persistence-sqlite](../../packages/session/session-persistence-sqlite)** — an opt-in `node:sqlite` backend using schema 18 to store exact same-block delta runs in bounded physical `text-chunks`, `reasoning-chunks`, and `tool-call-chunks` rows. It reconstructs the complete logical event stream before returning it, packs only newly durable batches, and rejects older schemas rather than migrating them.
 
 <!-- BEGIN GENERATED cordis-surface (gen-cordis-catalog.ts) — do not edit between markers -->
 

+ 1 - 1
docs/subsystems/persistence.zh.md

@@ -233,7 +233,7 @@ interface SessionPersistenceSnapshot {
 两者都实现同一个抽象 `SessionPersistence`(在 `SessionEvent` 上执行 locate/create/append/prepare/load/inspect/readFrom/list/listSnapshots,观察方法可选支持取消),并通过共享的 `runPersistenceContract` 套件:
 
 - **[dsh-session-persistence-jsonl](../../packages/session/session-persistence-jsonl)**——逐会话仅追加的逻辑 JSONL 日志,默认存储为带 checksum 的连续 Zstandard frame,也可配置为原始行;支持崩溃安全的原子写入、被中断轮次的恢复以及读取/回放路径。
-- **[dsh-session-persistence-sqlite](../../packages/session/session-persistence-sqlite)**:一个可选启用的 `node:sqlite` 后端,使用 schema 17 把同一分片块中字段完全匹配的 delta 连续段存为有界物理 `text-chunks`、`reasoning-chunks` 与 `tool-call-chunks` 行。它在返回前重建完整逻辑事件流,只打包新增的持久批次,并拒绝旧 schema,而不是执行迁移。
+- **[dsh-session-persistence-sqlite](../../packages/session/session-persistence-sqlite)**:一个可选启用的 `node:sqlite` 后端,使用 schema 18 把同一分片块中字段完全匹配的 delta 连续段存为有界物理 `text-chunks`、`reasoning-chunks` 与 `tool-call-chunks` 行。它在返回前重建完整逻辑事件流,只打包新增的持久批次,并拒绝旧 schema,而不是执行迁移。
 
 <!-- BEGIN GENERATED cordis-surface (gen-cordis-catalog.ts) — do not edit between markers -->
 

+ 2 - 2
packages/session/session-persistence-sqlite/README.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write packages/session/session-persistence-sqlite/README.md
-README.md: 11552519a47afaa44c05a4b16f6fc0dd48661991
-README.zh.md: 2c98da293b726114b319336635f086145fc55324
+README.md: 5f8d4cebf1a189cd768673359686ab2774550285
+README.zh.md: 12900cc8783430b8716612cdda55d2622cc7f7aa

+ 3 - 3
packages/session/session-persistence-sqlite/README.md

@@ -33,7 +33,7 @@ Choose this backend when a local deployment benefits from one queryable database
 
 ### Disk footprint and performance
 
-The packed layout trades disk space for speed and structure. On the benchmark corpus behind schema 17 — 105 sessions, about 2.5 million events — the SQLite database used 75 MB against 31 MB for the default compressed JSONL logs: roughly 2.5× the on-disk size.
+The packed layout trades disk space for speed and structure. The available benchmark measures schema 17, the packed predecessor with the same chunk codec but the former row discriminator; schema 18 has not been remeasured. On its corpus — 105 sessions, about 2.5 million events — the SQLite database used 75 MB against 31 MB for the default compressed JSONL logs: roughly 2.5× the on-disk size.
 
 The same measurements show writes finishing about 3× faster, 50-event suffix reads about 40× faster, full-session reads comparable or slightly faster, and about 2.5 million physical rows shrinking to roughly 66 thousand. Expect 2–3× the compressed JSONL footprint depending on session content; the full numbers and method live in the [SQLite physical chunk-row decision](../../../.agents/notes/implemented/architecture/2026-08-18-sqlite-physical-chunk-row-compression.md).
 
@@ -191,7 +191,7 @@ This Dev Note is working context for maintainers: measured artifacts, open desig
 
 #### Benchmark artifact
 
-The numbers below are the benchmark behind schema 17; the [SQLite physical chunk-row decision](../../../.agents/notes/implemented/architecture/2026-08-18-sqlite-physical-chunk-row-compression.md) is the authoritative record, and this table is an annotated digest.
+The numbers below are the frozen schema-17 benchmark. Schema 18 changes the row discriminator and has not been remeasured; the [SQLite physical chunk-row decision](../../../.agents/notes/implemented/architecture/2026-08-18-sqlite-physical-chunk-row-compression.md) is the authoritative record, and this table is an annotated digest.
 
 | Metric | **JSONL (zstd)** | **SQLite (legacy)** | **SQLite (new)** |
 |---|---|---|---|
@@ -202,7 +202,7 @@ The numbers below are the benchmark behind schema 17; the [SQLite physical chunk
 | Event rows | 2,507,860 (logical) | 2,507,860 | **65,810** |
 | Fork of all sessions | 14.48 s | 19.30 s | **13.10 s** |
 
-The corpus was 105 sessions with 2,507,860 logical events appended in 512-event durable batches, so the ratios depend on session content, stream density, and batch boundaries. `SQLite (legacy)` is the scalar layout — one physical row per logical event, no packing — whose 709.57 MB footprint motivated the packed rows. Against JSONL, schema 17 uses ≈2.5× the disk space but writes ≈3.3× faster, reads complete sessions faster at both percentiles, and reads 50-event tails ≈40× faster; against the scalar layout it is ≈89% smaller, faster to write, and shrinks 2,507,860 rows to 65,810, while scalar tail reads remain marginally faster (0.189 vs 0.253 ms p50). Re-run or extend this benchmark whenever the write path or the schema changes.
+The corpus was 105 sessions with 2,507,860 logical events appended in 512-event durable batches, so the ratios depend on session content, stream density, and batch boundaries. `SQLite (legacy)` is the scalar layout — one physical row per logical event, no packing — whose 709.57 MB footprint motivated the packed rows. In the measured schema-17 layout, SQLite uses ≈2.5× the JSONL disk space but writes ≈3.3× faster, reads complete sessions faster at both percentiles, and reads 50-event tails ≈40× faster; against the scalar layout it is ≈89% smaller, faster to write, and shrinks 2,507,860 rows to 65,810, while scalar tail reads remain marginally faster (0.189 vs 0.253 ms p50). Re-run or extend this benchmark whenever the write path or the schema changes.
 
 #### Future: multi-backend RDB persistence (Drizzle)
 

+ 3 - 3
packages/session/session-persistence-sqlite/README.zh.md

@@ -33,7 +33,7 @@ kind: "package-reference"
 
 ### 磁盘占用与性能
 
-打包布局以磁盘空间换取速度与结构。在 schema 17 的基准语料上——105 个会话、约 250 万个事件——SQLite 数据库占用 75 MB,而默认压缩 JSONL 日志为 31 MB:磁盘占用约为后者的 2.5 倍。
+打包布局以磁盘空间换取速度与结构。现有基准测量 schema 17,该打包前身使用相同的分片 codec,但行判别值不同;schema 18 尚未重新测量。在该语料上——105 个会话、约 250 万个事件——SQLite 数据库占用 75 MB,而默认压缩 JSONL 日志为 31 MB:磁盘占用约为后者的 2.5 倍。
 
 同一组测量显示,写入快约 3 倍,50 个事件的后缀读取快约 40 倍,完整会话读取相当或略快,约 250 万个物理行缩减到约 6.6 万个。按会话内容不同,磁盘占用约为压缩 JSONL 的 2–3 倍;完整数据与方法见 [SQLite 物理分片行决策](../../../.agents/notes/implemented/architecture/2026-08-18-sqlite-physical-chunk-row-compression.zh.md)。
 
@@ -191,7 +191,7 @@ await ctx.sessionPersistence.append(id, events)
 
 #### 基准产物
 
-以下数字是 schema 17 背后的基准;[SQLite 物理分片行决策](../../../.agents/notes/implemented/architecture/2026-08-18-sqlite-physical-chunk-row-compression.zh.md) 是权威记录,本表只是带注的摘要。
+以下数字是冻结的 schema 17 基准。Schema 18 改变了行判别值,尚未重新测量;[SQLite 物理分片行决策](../../../.agents/notes/implemented/architecture/2026-08-18-sqlite-physical-chunk-row-compression.zh.md) 是权威记录,本表只是带注的摘要。
 
 | 指标 | **JSONL(zstd)** | **SQLite(legacy)** | **SQLite(new)** |
 |---|---|---|---|
@@ -202,7 +202,7 @@ await ctx.sessionPersistence.append(id, events)
 | 事件行数 | 2,507,860(逻辑) | 2,507,860 | **65,810** |
 | 全部会话 fork | 14.48 s | 19.30 s | **13.10 s** |
 
-语料为 105 个会话、2,507,860 个逻辑事件,按 512 个事件一批追加,因此具体比例取决于会话内容、流密度与批次边界。`SQLite(legacy)` 是标量布局——每个逻辑事件一行、不打包——其 709.57 MB 的占用正是打包行的动机。相对 JSONL,schema 17 磁盘占用约为 2.5 倍,但写入快约 3.3 倍,完整会话读取在两个分位上都更快,50 个事件尾部读取快约 40 倍;相对标量布局,它缩小约 89%、写入更快,并把 2,507,860 行缩减到 65,810 行,只有标量尾部读取仍略快(0.189 对 0.253 ms p50)。写入路径或 schema 变化时,请重跑或扩展该基准。
+语料为 105 个会话、2,507,860 个逻辑事件,按 512 个事件一批追加,因此具体比例取决于会话内容、流密度与批次边界。`SQLite(legacy)` 是标量布局——每个逻辑事件一行、不打包——其 709.57 MB 的占用正是打包行的动机。在已测量的 schema 17 布局中,SQLite 磁盘占用约为 JSONL 的 2.5 倍,但写入快约 3.3 倍,完整会话读取在两个分位上都更快,50 个事件尾部读取快约 40 倍;相对标量布局,它缩小约 89%、写入更快,并把 2,507,860 行缩减到 65,810 行,只有标量尾部读取仍略快(0.189 对 0.253 ms p50)。写入路径或 schema 变化时,请重跑或扩展该基准。
 
 #### 未来:多后端 RDB 持久化(Drizzle)
 

+ 1 - 1
packages/session/session-persistence-sqlite/src/codec.ts

@@ -1,5 +1,5 @@
 /**
- * Schema-17 physical chunk-row codec. This package owns the durable tags,
+ * Schema-18 physical chunk-row codec. This package owns the durable tags,
  * validation, and row-size limits independently from other persistence formats.
  * @module @deepseek-ai/dsh-session-persistence-sqlite/codec
  */