Bladeren bron

Clarify token counts after aspect-preserving image projection

creatixchu 4 weken geleden
bovenliggende
commit
f0f1988a7a

+ 2 - 2
.agents/notes/implemented/bug-fix/2026-09-10-deepseek-v41-request-image-projection.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/bug-fix/2026-09-10-deepseek-v41-request-image-projection.md
-2026-09-10-deepseek-v41-request-image-projection.md: 87f79f244dcf61039fe27ef170ba654ee7a1a62a
-2026-09-10-deepseek-v41-request-image-projection.zh.md: a5d8a4e99a3c7b4f3fc26f385ae1e764b535fe15
+2026-09-10-deepseek-v41-request-image-projection.md: 2bf7ac3f7f09ff7aea1067030f4474306d3f9a8e
+2026-09-10-deepseek-v41-request-image-projection.zh.md: 4f0a8cfd3a71d54a54bf74d6819235ed7d0749f8

+ 2 - 2
.agents/notes/implemented/bug-fix/2026-09-10-deepseek-v41-request-image-projection.md

@@ -22,10 +22,10 @@ The route chooses each request image's dimensions; the attachment provider only
 
 **A `token-grid` projection kind on the attachment policy, with the solver in `dsh-attachment`.** This was built first: the policy became a closed union of `pixel-budget` and `token-grid`, the solver moved into `request-projection.ts`, and `deepSeekImageTokens` imported it back. It put one provider's layout formula and patch constants into the provider-neutral package under a generic-looking name, needed an `unscaled` flag so the store could tell "send the source" from "send the solved size", and would grow a new union member for every provider rule. Handing the store a finished target keeps the provider rule beside the provider's pricing and leaves the store with no projection vocabulary at all.
 
-**Send the solver's exact grid dimensions with a fill resize.** The solved grid edges are whole patches and differ from the source aspect ratio by under one patch. Filling that box would distort the image slightly even though the provider does the same on its side; the issue requires the aspect ratio preserved, the padding on the provider side does not change the token count, and pricing reproduces the provider from the sent dimensions either way.
+**Send the solver's exact grid dimensions with a fill resize.** The solved grid edges are whole patches and differ from the source aspect ratio by under one patch. Filling that box would distort the image slightly even though the provider does the same on its side; preserving the source aspect ratio can change how many token cells the rounded short edge covers. A 1224×1429 source is sent as 1187×1386: the published calculator gives 959 tokens for the source and 992 for the sent dimensions. Request generation and pricing share the target dimensions, and pricing applies the published calculator to that target.
 
 **Keep the 1 MiB target.** The target is not a cap: an output over it is still sent at the smallest ladder quality. At 1302×1302 a JPEG photograph at quality 85 lands between 400 KB and 1.2 MB, so 2 MiB keeps most images at the top quality and lets more PNG screenshots pass through losslessly, while the inline base64 fallback still holds about seven such images under its 20 MiB bound.
 
 ## Consequences
 
-A square source now reaches the model at up to 1302×1302 pixels and 994 tokens instead of 800×800 and 422, so image-heavy sessions reach compaction pressure sooner and the estimator matches provider usage for the sent version. Every existing request-image cache entry and DeepSeek Files API mapping is regenerated on the next request. Thin images keep their full grid until the per-side cap applies: an 8192×78 source costs 396 tokens under the grid but is sent as 4096×39. `llm-replay` does not project images, so keyless snapshots cannot record the sent dimensions; the `llm-deepseek` adapter tests pin the resolved targets and the projected handle text against a mock server, and the local store tests resize real images to targets.
+A square source now reaches the model at up to 1302×1302 pixels and 994 tokens instead of 800×800 and 422, so image-heavy sessions reach compaction pressure sooner and the estimator applies the published token rules to the sent target dimensions. Every existing request-image cache entry and DeepSeek Files API mapping is regenerated on the next request. Thin images keep their full grid until the per-side cap applies: an 8192×78 source costs 396 tokens under the grid but is sent as 4096×39. `llm-replay` does not project images, so keyless snapshots cannot record the sent dimensions; the `llm-deepseek` adapter tests pin the resolved targets and the projected handle text against a mock server, and the local store tests resize real images to targets.

+ 2 - 2
.agents/notes/implemented/bug-fix/2026-09-10-deepseek-v41-request-image-projection.zh.md

@@ -22,10 +22,10 @@ harness 此前把每张 DeepSeek 请求图片投影到 640,000 总像素预算
 
 **在附件策略上加 `token-grid` 投影种类,求解器放进 `dsh-attachment`。** 最初就是这样做的:策略变成 `pixel-budget` 和 `token-grid` 的封闭联合,求解器搬进 `request-projection.ts`,`deepSeekImageTokens` 再从那里引回来。这把一家提供方的布局公式和 patch 常量放进了提供方无关的包,还起了个看似通用的名字;存储层需要一个 `unscaled` 标志来区分「发源图」和「发求解尺寸」;以后每多一家提供方规则,联合就要多长一个分支。把算好的目标交给存储层,提供方规则和它的计价放在一起,存储层不需要任何投影词汇。
 
-**用填充缩放发送求解器的精确网格尺寸。** 求解出的网格边长是整数个 patch,与源图宽高比相差不到一个 patch。填充到这个框会轻微变形,尽管提供方那侧也会这样做;issue 要求保持宽高比,提供方那侧的补齐不改变 token 数,而计价无论如何都从发送尺寸复现提供方。
+**用填充缩放发送求解器的精确网格尺寸。** 求解出的网格边长是整数个 patch,与源图宽高比相差不到一个 patch。填充到这个框会轻微变形,尽管提供方那侧也会这样做;保持源图宽高比可能改变取整后短边覆盖的 token 格数。1224×1429 的源图以 1187×1386 发送,官方计算器对源图计 959 token,对发送尺寸计 992 token。请求生成和定价共享目标尺寸,定价按目标尺寸应用官方计算规则。
 
 **保留 1 MiB 目标。** 目标不是上限:超过它的输出仍会以阶梯最小质量发送。1302×1302 的 JPEG 照片在质量 85 时约 400 KB 到 1.2 MB,2 MiB 让多数图片停在最高质量,也让更多 PNG 截图无损直发,而内联 base64 回退在 20 MiB 上界内仍能容纳约七张这样的图片。
 
 ## 后果
 
-正方形源图现在最多以 1302×1302 像素、994 token 到达模型,而不是 800×800 和 422,因此图片密集的会话更早触及 compaction 压力,预估器对发送版本的计价与提供方 usage 一致。所有已有的请求图片缓存条目和 DeepSeek Files API 映射在下次请求时重新生成。细长图在单边上限生效前保留完整网格:8192×78 的源图在网格下计 396 token,但以 4096×39 发送。`llm-replay` 不投影图片,keyless 快照记录不到发送尺寸;`llm-deepseek` 适配器测试对着 mock 服务器固定了解析出的目标和投影后的句柄文本,本地存储测试把真实图片缩放到目标尺寸。
+正方形源图现在最多以 1302×1302 像素、994 token 到达模型,而不是 800×800 和 422,因此图片密集的会话更早触及 compaction 压力,预估器按发送目标尺寸应用官方 token 计算规则。所有已有的请求图片缓存条目和 DeepSeek Files API 映射在下次请求时重新生成。细长图在单边上限生效前保留完整网格:8192×78 的源图在网格下计 396 token,但以 4096×39 发送。`llm-replay` 不投影图片,keyless 快照记录不到发送尺寸;`llm-deepseek` 适配器测试对着 mock 服务器固定了解析出的目标和投影后的句柄文本,本地存储测试把真实图片缩放到目标尺寸。

+ 10 - 0
packages/attachment/attachment-local/tests/request-image.spec.ts

@@ -137,6 +137,16 @@ describe('local request-image cache', () => {
     expect(enlarged.variantId).not.toBe(thinRequest.variantId)
   })
 
+  it('encodes the rounded short edge at the route target', async () => {
+    const attachments = await store()
+    const attachment = await attachments.saveImage({ data: await image(1224, 1429), mediaType: 'image/png' })
+    const request = await attachments.readImageRequest(attachment, {
+      width: 1187, height: 1386, maxBytes: 2 * 1024 * 1024,
+    })
+    expect(request).toMatchObject({ width: 1187, height: 1386 })
+    await expect(sharp(request.data).metadata()).resolves.toMatchObject({ width: 1187, height: 1386 })
+  })
+
   it('keeps the smallest ladder output when the encoded-byte target is unreachable', async () => {
     const attachments = await store()
     const attachment = await attachments.saveImage({ data: await image(1, 1), mediaType: 'image/png' })

+ 3 - 2
packages/llm/llm-deepseek/src/image-tokens.ts

@@ -122,8 +122,9 @@ function sameResize(a: GridResize, b: GridResize): boolean {
  * Dimensions the harness sends so the provider keeps the whole image: the
  * source itself when its patch-padded grid fits the token cap, otherwise the
  * source aspect ratio at the solved grid's long edge. The provider pads the
- * short edge to whole patches on its side, so the token count equals the
- * solved grid's. Small images are never enlarged.
+ * short edge to whole patches on its side. Rounding the aspect-preserving
+ * short edge can change the token count from the source's solved grid;
+ * request pricing uses the sent dimensions. Small images are never enlarged.
  * @param width - positive integer source width in pixels.
  * @param height - positive integer source height in pixels.
  * @returns the request dimensions to encode.

+ 7 - 0
packages/llm/llm-deepseek/tests/image-tokens.spec.ts

@@ -49,6 +49,13 @@ describe('DeepSeek image tokens', () => {
 })
 
 describe('DeepSeek request image dimensions', () => {
+  it('can cross a token-cell boundary when preserving the source aspect ratio', () => {
+    const sent = deepSeekRequestImageDimensions(1224, 1429)
+    expect(sent).toEqual({ width: 1187, height: 1386 })
+    expect(deepSeekImageTokens(1224, 1429)).toBe(959)
+    expect(deepSeekImageTokens(sent.width, sent.height)).toBe(992)
+  })
+
   it.each([
     [800, 800, 800, 800],
     [1302, 1302, 1302, 1302],

+ 9 - 0
packages/llm/llm-deepseek/tests/request-pricing.spec.ts

@@ -61,6 +61,15 @@ describe('DeepSeek request-image pricing', () => {
     }])
   })
 
+  it('prices the sent dimensions when aspect-preserving projection changes the token grid', () => {
+    const image = ref('portrait', 1224, 1429)
+    const prices = deepSeekImageRequestPricing(connection(), 'vision').priceImages([image])
+    expect(prices).toEqual([{
+      visualTokens: 992,
+      text: requestImageHandleText(image, { width: 1187, height: 1386 }),
+    }])
+  })
+
   it('honors a numeric pixel budget override', () => {
     const image = ref('photo', 4096, 4096)
     const options = resolveAdapterOptions({