瀏覽代碼

fix(llm-deepseek): 图片 token 预估器改用官方计算器 v41 配置

DeepSeek 图像理解指南改为 544×544 放大下限、约 1300×1300 缩小目标、
单图 1024 token 上限;官方图片 Token 计算器以 v41 配置承载该规则。
重写 image-tokens.ts 为 v41 的逐句移植:取消对齐 pad、宽高比钳制和
奇偶校正网格布局,估算值从最坏情况上界变为精确值。测试向量按官方
计算器重新固定,request-pricing 期望值、README 与 Agent Note 同步更新。
creatixchu 3 周之前
父節點
當前提交
a64dc3a690

+ 6 - 0
.agents/notes/implemented/bug-fix/2026-09-10-deepseek-image-token-calculator-v41.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/bug-fix/2026-09-10-deepseek-image-token-calculator-v41.md
+2026-09-10-deepseek-image-token-calculator-v41.md: d9a8b7e72143793f13aaf3a4db26468dc0f4c162
+2026-09-10-deepseek-image-token-calculator-v41.zh.md: 26cdf23ccf2b693813648c015b5a9926cd362ab2

+ 27 - 0
.agents/notes/implemented/bug-fix/2026-09-10-deepseek-image-token-calculator-v41.md

@@ -0,0 +1,27 @@
+# Agent Note: DeepSeek image-token estimator on the v41 calculator
+
+Status: implemented
+
+English | [中文](2026-09-10-deepseek-image-token-calculator-v41.zh.md)
+
+## Problem
+
+`deepSeekImageTokens()` in `llm-deepseek` ported the provider's published image-token calculator in its `v4` configuration: a 384×384 scale-up floor, a 384-token cap, an 8:1 width clamp, a grid layout that adds a row for odd row counts and parity corrections, and a pad-to-4 alignment charged at its worst case. The provider's Vision guide now documents a different projection for the current Flash model: images below roughly 544×544 total pixels scale up, larger images scale down to roughly 1300×1300 total pixels, and one image costs at most 1024 tokens. The published calculator carries this as a `v41` configuration and the docs page instantiates that one. The old port therefore underpriced every retained image, an 800×800 request image by 73 tokens, so automatic compaction on image-dense sessions triggered late.
+
+## Decision
+
+`image-tokens.ts` is rewritten as a verbatim port of the `v41` configuration. The constants are a 14px patch, 3:1 per-axis downsampling, a 544×544 total-pixel floor, and a 1024-token cap. The grid formula is `rows × (cols + 1) + 2` with no odd-row extra row, no parity correction, and no even-row trimming in the solver. There is no alignment pad, so the estimate is exact rather than a worst-case upper bound, and there is no aspect-ratio clamp, so extreme aspect ratios reach the cap through the solver's one-row and one-column branches. The over-budget path is a single closed-form solve followed by the published assertion; the decrementing retry loop existed only for the odd-row layout. The provider's fixpoint iteration over the projected dimensions is unchanged.
+
+The test vectors are re-pinned from the published calculator. The request-pricing tests, package README, and this note carry the new numbers; the pixel budget the harness applies before pricing (`DEFAULT_REQUEST_IMAGE_PIXEL_BUDGET`, 640,000 total pixels) and the catalog model ids are unchanged.
+
+## Alternatives considered
+
+**Keep both configurations and select by model id.** The provider states that requests to the retired `deepseek-v4-flash-vision-exp` id are served by the current Flash model, so no reachable route prices under the old configuration. Two configurations would keep dead branches and their tests alive.
+
+**Keep the generic class with the `isNLayout`, pad, and ratio-clamp switches.** A one-configuration port has fewer unreachable branches to exclude from coverage and states the shipped rule directly; a future provider revision changes this one module and its pinned vectors either way.
+
+**Raise the harness pixel budget in the same change.** The provider now accepts roughly 1300×1300 total pixels per image, so the harness's 640,000-pixel projection discards detail the model could use. That is a request-content change with its own snapshot impact, separate from pricing what is actually sent.
+
+## Consequences
+
+Retained request images price higher, so image-dense DeepSeek sessions compact earlier and closer to the provider's real pressure. Under the harness's 640,000-pixel projection the ceiling for one image is 422 tokens at 800×800 rather than 349. The estimate no longer carries a three-token conservative margin; provider usage remains the authoritative anchor once a request completes. Sessions replayed through `llm-replay` use their fixture's `imageRequestTokens` and are unaffected.

+ 27 - 0
.agents/notes/implemented/bug-fix/2026-09-10-deepseek-image-token-calculator-v41.zh.md

@@ -0,0 +1,27 @@
+# Agent Note: DeepSeek 图片 token 估算器改用 v41 计算器
+
+Status: implemented
+
+[English](2026-09-10-deepseek-image-token-calculator-v41.md) | 中文
+
+## 问题
+
+`llm-deepseek` 中的 `deepSeekImageTokens()` 移植的是提供方公开图片 token 计算器的 `v4` 配置:384×384 放大下限、384 token 上限、8:1 宽度钳制、奇数行数额外加一行并做奇偶校正的网格布局,以及按最坏情况计价的 pad-to-4 对齐。提供方的图像理解指南现在为当前 Flash 模型记录了另一套投影规则:总像素小于约 544×544 的图片放大,更大的图片缩小到约 1300×1300 总像素,单张图片最多 1024 token。公开计算器以 `v41` 配置承载这套规则,文档页实例化的也是它。旧移植因此对每张保留图片都估低,一张 800×800 的请求图片低估 73 token,图片密集会话的自动压缩触发过晚。
+
+## 决策
+
+`image-tokens.ts` 重写为 `v41` 配置的逐句移植。常量为 14px patch、每轴 3:1 降采样、544×544 总像素下限、1024 token 上限。网格公式为 `rows × (cols + 1) + 2`,没有奇数行额外行、没有奇偶校正、求解器也不再把行数截成偶数。没有对齐 pad,所以估算值是精确值而非最坏情况上界;没有宽高比钳制,所以极端长宽比会经求解器的单行和单列分支到达上限。超预算路径是一次闭式求解加上公开的断言;逐步递减的重试循环只服务于奇数行布局。提供方对投影尺寸的定点迭代保持不变。
+
+测试向量按公开计算器重新固定。request-pricing 测试、包 README 和本 note 使用新数字;harness 在定价前应用的像素预算(`DEFAULT_REQUEST_IMAGE_PIXEL_BUDGET`,640,000 总像素)和 catalog 模型 id 不变。
+
+## 备选方案
+
+**保留两套配置并按模型 id 选择。** 提供方说明发往已下线的 `deepseek-v4-flash-vision-exp` 的请求由当前 Flash 模型承接,所以没有可达路由会按旧配置计价。两套配置会保留死分支及其测试。
+
+**保留带 `isNLayout`、pad 和宽高比钳制开关的通用类。** 单配置移植需要从覆盖率中排除的不可达分支更少,并直接陈述已上线的规则;提供方未来再修订时,两种写法都只改这一个模块和它固定的向量。
+
+**在同一改动中提高 harness 像素预算。** 提供方现在每张图接受约 1300×1300 总像素,harness 的 640,000 像素投影会丢弃模型本可利用的细节。那是请求内容的改动,有自己的快照影响,与为实际发送内容计价是两件事。
+
+## 后果
+
+保留的请求图片计价更高,图片密集的 DeepSeek 会话会更早压缩,更接近提供方的真实压力。在 harness 的 640,000 像素投影下,单张图片的上限是 800×800 时的 422 token,而非 349。估算值不再带 3 token 的保守余量;请求完成后,提供方 usage 仍是权威锚点。经 `llm-replay` 回放的会话使用各自 fixture 的 `imageRequestTokens`,不受影响。

+ 2 - 2
packages/llm/llm-deepseek/README.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write packages/llm/llm-deepseek/README.md
-README.md: 73992b3d5e6b72215cb3994881c361147c6885d8
-README.zh.md: 47d6cad48cde5dec8139d6981436bc6023f276b1
+README.md: 67c700bec9fbe6b9df7ab95ded5c991bf01dba1c
+README.zh.md: 43670a3947519ae1b7598900b8ae329becf9ad6a

+ 1 - 1
packages/llm/llm-deepseek/README.md

@@ -156,7 +156,7 @@ The selected DeepSeek model receives the harness system prompt, message history,
 
 #### Token effect
 
-Provider tokenization governs exact text and image-token input. The adapter declares per-route `imageRequestPricing`: it reproduces oldest-first image offload from durable byte lengths and prices each retained image at its projected dimensions with the published v4 vision accounting (14px patch grid, 3:1 downsampling, 384-token cap, worst-case alignment pad). This lets the token meter price image pressure before a request; reported usage remains authoritative. Reasoning passback carries every reasoned turn's chain of thought into later requests, while dropping over-budget images avoids paying those tokens again. Cache-read usage is reported when available. `totalTokens` is the exact `prompt_tokens + completion_tokens` aggregate and is omitted if a supplied `total_tokens` disagrees.
+Provider tokenization governs exact text and image-token input. The adapter declares per-route `imageRequestPricing`: it reproduces oldest-first image offload from durable byte lengths and prices each retained image at its projected dimensions with the published vision accounting (14px patch grid, 3:1 downsampling, 544×544 scale-up floor, 1024-token cap). This lets the token meter price image pressure before a request; reported usage remains authoritative. Reasoning passback carries every reasoned turn's chain of thought into later requests, while dropping over-budget images avoids paying those tokens again. Cache-read usage is reported when available. `totalTokens` is the exact `prompt_tokens + completion_tokens` aggregate and is omitted if a supplied `total_tokens` disagrees.
 
 #### KV Cache effect
 

+ 1 - 1
packages/llm/llm-deepseek/README.zh.md

@@ -156,7 +156,7 @@ Files 模式通过 `maxRequestFilesBytes` 与 `maxImagesPerRequest` 限制保留
 
 #### Token 影响
 
-提供方分词决定精确的文本与图片 token 输入。适配器声明按路由的 `imageRequestPricing`:它根据持久记录中的字节长度复现最旧优先的图片 offload,并按投影后的尺寸使用公开的 v4 视觉计量规则(14 px patch 网格、3:1 降采样、单图 384 token 上限、最坏情况下的对齐 pad)为每张保留图片计价。这使 token 计量服务可以在请求发出前为图片压力定价;上报的 usage 仍是权威值。推理回传会把每个推理轮次的思维链带进后续请求,而丢弃超预算图片会避免再次为它们付费。可用时报告缓存读取用量。`totalTokens` 是精确的 `prompt_tokens + completion_tokens` 汇总值;提供方给出的 `total_tokens` 不一致时省略该值。
+提供方分词决定精确的文本与图片 token 输入。适配器声明按路由的 `imageRequestPricing`:它根据持久记录中的字节长度复现最旧优先的图片 offload,并按投影后的尺寸使用公开的视觉计量规则(14 px patch 网格、3:1 降采样、544×544 放大下限、单图 1024 token 上限)为每张保留图片计价。这使 token 计量服务可以在请求发出前为图片压力定价;上报的 usage 仍是权威值。推理回传会把每个推理轮次的思维链带进后续请求,而丢弃超预算图片会避免再次为它们付费。可用时报告缓存读取用量。`totalTokens` 是精确的 `prompt_tokens + completion_tokens` 汇总值;提供方给出的 `total_tokens` 不一致时省略该值。
 
 #### KV Cache 影响
 

+ 52 - 70
packages/llm/llm-deepseek/src/image-tokens.ts

@@ -1,10 +1,12 @@
 /**
- * DeepSeek v4 vision-token accounting: the provider's published image-token
- * calculator (api-docs.deepseek.com, Token & Token Usage) ported verbatim.
- * The provider resizes every request image onto a 14px-patch grid, downsamples
- * 3:1 per axis, and caps one image at 384 tokens; the port prices the
- * pad-to-4 alignment at its 3-token upper bound because request pricing has
- * no preceding-token position. Actual usage remains authoritative.
+ * DeepSeek vision-token accounting: the provider's published image-token
+ * calculator (api-docs.deepseek.com, Token & Token Usage) ported verbatim in
+ * its current `v41` configuration. The provider scales an image below
+ * 544×544 total pixels up, aligns it to a 14px-patch grid, downsamples 3:1
+ * per axis into token cells, and caps one image at 1024 tokens by solving the
+ * largest aspect-preserving grid inside that budget. The count is exact: this
+ * configuration has no alignment pad and no aspect-ratio clamp. Actual usage
+ * remains authoritative.
  *
  * @module dsh-llm-deepseek/image-tokens
  */
@@ -14,13 +16,11 @@ const PATCH_SIZE = 14
 /** Per-axis patch-to-token downsampling ratio. */
 const DOWNSAMPLE_RATIO = 3
 /** Provider cap on tokens for one request image. */
-const MAX_IMAGE_TOKENS = 384
-/** Token-alignment quantum; pricing charges its worst-case `QUANTUM - 1` pad. */
-const COMPRESS_PAD_TO = 4
-/** Width is clamped to this multiple of height before grid projection. */
-const MAX_WIDTH_HEIGHT_RATIO = 8
+const MAX_IMAGE_TOKENS = 1024
 /** Total-pixel floor; smaller images are scaled up before grid projection. */
-const MIN_PIXELS = 384 * 384
+const MIN_PIXELS = 544 * 544
+/** Pixels covered by one token cell along either axis. */
+const CELL_SIZE = PATCH_SIZE * DOWNSAMPLE_RATIO
 
 const intDiv = (value: number, divisor: number): number => Math.floor(value / divisor)
 const ceilDiv = (value: number, divisor: number): number => Math.floor((value + divisor - 1) / divisor)
@@ -33,12 +33,14 @@ interface GridResize {
   readonly numTokens: number
 }
 
-/** Token count of one grid, including row separators and framing. */
+/** Token count of one grid: every row carries a separator, plus two framing tokens. */
 function gridTokens(gridHeight: number, gridWidth: number): number {
-  let tokens = gridHeight * (gridWidth + 1) + 2
-  if (gridHeight % 2 === 1) tokens += gridWidth + 1
-  tokens += (ceilDiv(gridHeight, 2) * (gridWidth + 1) % 2) * 2
-  return tokens
+  return gridHeight * (gridWidth + 1) + 2
+}
+
+/** Token grid the padded pixel dimensions project onto. */
+function gridCells(paddedPixels: number): number {
+  return ceilDiv(intDiv(paddedPixels, PATCH_SIZE), DOWNSAMPLE_RATIO)
 }
 
 /** Solve the largest grid within `budget` tokens preserving the aspect ratio. */
@@ -50,80 +52,60 @@ function solveResizeRatio(height: number, width: number, budget: number): GridRe
   let bestWidth: number
   if (idealGridWidth < 1) {
     const solvedGridWidth = 1
-    let solvedGridHeight = intDiv(budget - 2, solvedGridWidth + 1)
-    // v8 ignore: at the provider budget the one-column solve always lands on
-    // the odd 189-row grid, so the even path is unreachable; kept for parity
-    // with the published solver.
-    /* v8 ignore next */
-    if (solvedGridHeight % 2 === 1) solvedGridHeight -= 1
-    bestWidth = solvedGridWidth * PATCH_SIZE * DOWNSAMPLE_RATIO
-    bestHeight = solvedGridHeight * PATCH_SIZE * DOWNSAMPLE_RATIO
-  /* v8 ignore start -- unreachable at the provider budget: idealGridWidth >= 1
-     bounds the aspect at (budget - 2) / 2, making idealGridHeight >= 2 for
-     every budget this module solves; kept for parity with the published
-     solver. */
-  } else if (idealGridHeight < 2) {
-    const solvedGridHeight = 2
+    const solvedGridHeight = intDiv(budget - 2, solvedGridWidth + 1)
+    bestWidth = solvedGridWidth * CELL_SIZE
+    bestHeight = solvedGridHeight * CELL_SIZE
+  } else if (idealGridHeight < 1) {
+    const solvedGridHeight = 1
     const solvedGridWidth = intDiv(budget - 2, solvedGridHeight) - 1
-    if (!(solvedGridWidth > 1)) throw new Error('deepseek image tokens: no grid fits the token budget')
-    bestWidth = solvedGridWidth * PATCH_SIZE * DOWNSAMPLE_RATIO
-    bestHeight = solvedGridHeight * PATCH_SIZE * DOWNSAMPLE_RATIO
-  /* v8 ignore stop */
+    bestWidth = solvedGridWidth * CELL_SIZE
+    bestHeight = solvedGridHeight * CELL_SIZE
   } else {
     const solvedGridWidth = Math.trunc(idealGridWidth)
-    let solvedGridHeight = Math.trunc(idealGridHeight)
-    if (solvedGridHeight % 2 === 1) solvedGridHeight -= 1
-    const widthScale = solvedGridWidth * PATCH_SIZE * DOWNSAMPLE_RATIO / width
-    const heightScale = solvedGridHeight * PATCH_SIZE * DOWNSAMPLE_RATIO / height
-    const scale = Math.min(widthScale, heightScale)
+    const solvedGridHeight = Math.trunc(idealGridHeight)
+    const scale = Math.min(solvedGridWidth * CELL_SIZE / width, solvedGridHeight * CELL_SIZE / height)
     bestWidth = Math.trunc(width * scale / PATCH_SIZE) * PATCH_SIZE
     bestHeight = Math.trunc(height * scale / PATCH_SIZE) * PATCH_SIZE
   }
-  const gridHeight = ceilDiv(intDiv(bestHeight, PATCH_SIZE), DOWNSAMPLE_RATIO)
-  const gridWidth = ceilDiv(intDiv(bestWidth, PATCH_SIZE), DOWNSAMPLE_RATIO)
+  const gridHeight = gridCells(bestHeight)
+  const gridWidth = gridCells(bestWidth)
   return { gridHeight, gridWidth, bestHeight, bestWidth, numTokens: gridTokens(gridHeight, gridWidth) }
 }
 
 /** Project padded pixel dimensions onto the largest in-budget token grid. */
 function safeResize(height: number, width: number, paddedHeight: number, paddedWidth: number): GridResize {
-  const gridHeight = ceilDiv(intDiv(paddedHeight, PATCH_SIZE), DOWNSAMPLE_RATIO)
-  const gridWidth = ceilDiv(intDiv(paddedWidth, PATCH_SIZE), DOWNSAMPLE_RATIO)
-  const pad = COMPRESS_PAD_TO - 1
-  const budget = MAX_IMAGE_TOKENS - pad
-  let result: GridResize = {
+  const gridHeight = gridCells(paddedHeight)
+  const gridWidth = gridCells(paddedWidth)
+  const direct: GridResize = {
     gridHeight,
     gridWidth,
     bestHeight: paddedHeight,
     bestWidth: paddedWidth,
     numTokens: gridTokens(gridHeight, gridWidth),
   }
-  if (result.numTokens > budget) {
-    result = solveResizeRatio(height, width, budget)
-    /* v8 ignore next 4 -- the published solver's safety net; the closed-form
-       solve stays within budget for every geometry the clamps admit. */
-    for (let reduced = budget; result.numTokens > budget; reduced -= 1) {
-      result = solveResizeRatio(height, width, reduced)
-    }
+  if (direct.numTokens <= MAX_IMAGE_TOKENS) return direct
+  const solved = solveResizeRatio(height, width, MAX_IMAGE_TOKENS)
+  /* v8 ignore next 3 -- the published solver's assertion; the closed-form
+     solve stays within the budget for every positive geometry. */
+  if (solved.numTokens > MAX_IMAGE_TOKENS) {
+    throw new Error(`deepseek image tokens: no grid fits the token budget for ${width}x${height}`)
   }
-  return { ...result, numTokens: result.numTokens + pad }
+  return solved
 }
 
-/** One clamp-scale-pad-project pass; the caller iterates it to a fixpoint. */
+/** One scale-pad-project pass; the caller iterates it to a fixpoint. */
 function resizeOnce(width: number, height: number): GridResize {
-  let clampedWidth = width
-  let clampedHeight = height
-  if (clampedWidth > clampedHeight * MAX_WIDTH_HEIGHT_RATIO) {
-    clampedWidth = clampedHeight * MAX_WIDTH_HEIGHT_RATIO
-  }
-  const pixels = clampedWidth * clampedHeight
+  let scaledWidth = width
+  let scaledHeight = height
+  const pixels = scaledWidth * scaledHeight
   if (pixels < MIN_PIXELS && pixels > 0) {
     const scale = Math.sqrt(MIN_PIXELS / pixels)
-    clampedWidth = Math.trunc(clampedWidth * scale)
-    clampedHeight = Math.trunc(clampedHeight * scale)
+    scaledWidth = Math.trunc(scaledWidth * scale)
+    scaledHeight = Math.trunc(scaledHeight * scale)
   }
-  const paddedWidth = ceilDiv(clampedWidth, PATCH_SIZE) * PATCH_SIZE
-  const paddedHeight = ceilDiv(clampedHeight, PATCH_SIZE) * PATCH_SIZE
-  return safeResize(clampedHeight, clampedWidth, paddedHeight, paddedWidth)
+  const paddedWidth = ceilDiv(scaledWidth, PATCH_SIZE) * PATCH_SIZE
+  const paddedHeight = ceilDiv(scaledHeight, PATCH_SIZE) * PATCH_SIZE
+  return safeResize(scaledHeight, scaledWidth, paddedHeight, paddedWidth)
 }
 
 function sameResize(a: GridResize, b: GridResize): boolean {
@@ -135,11 +117,11 @@ function sameResize(a: GridResize, b: GridResize): boolean {
 }
 
 /**
- * Vision tokens DeepSeek v4 charges for one request image of the given
- * dimensions, at the worst-case alignment pad.
+ * Vision tokens DeepSeek charges for one request image of the given
+ * dimensions.
  * @param width - positive integer request-image width in pixels.
  * @param height - positive integer request-image height in pixels.
- * @returns the provider vision-token price, at most 384.
+ * @returns the provider vision-token price, at most 1024.
  */
 export function deepSeekImageTokens(width: number, height: number): number {
   let result = resizeOnce(width, height)

+ 28 - 29
packages/llm/llm-deepseek/tests/image-tokens.spec.ts

@@ -1,53 +1,52 @@
 import { describe, expect, it } from 'vitest'
 import { deepSeekImageTokens } from '../src/image-tokens.ts'
 
-describe('DeepSeek v4 image tokens', () => {
+describe('DeepSeek image tokens', () => {
   // Reference values from the provider's published image token calculator
-  // (api-docs.deepseek.com, Token & Token Usage), at the worst-case pad.
+  // (api-docs.deepseek.com, Token & Token Usage) in its v41 configuration.
   it.each([
-    [100, 100, 117],
-    [384, 384, 117],
-    [640, 480, 209],
-    [800, 800, 349],
-    [1024, 768, 357],
-    [1920, 1080, 369],
-    [2000, 2000, 349],
-    [5000, 5000, 349],
-    [300, 50, 101],
+    [100, 100, 184],
+    [544, 544, 184],
+    [640, 480, 206],
+    [800, 800, 422],
+    [1024, 768, 496],
+    [1066, 600, 407],
+    [1300, 1300, 994],
+    [1920, 1080, 968],
+    [2000, 2000, 994],
+    [5000, 5000, 994],
+    [300, 50, 200],
   ])('prices %sx%s as %s tokens', (width, height, expected) => {
     expect(deepSeekImageTokens(width, height)).toBe(expected)
   })
 
-  it('caps every image at 384 tokens regardless of source size', () => {
-    for (const [width, height] of [[2000, 2000], [5000, 5000], [8192, 8192], [16, 8192]]) {
-      expect(deepSeekImageTokens(width!, height!)).toBeLessThanOrEqual(384)
+  it('caps every image at 1024 tokens regardless of source size', () => {
+    for (const [width, height] of [[2000, 2000], [5000, 5000], [8192, 8192], [16, 8192], [9000, 1], [1, 9000]]) {
+      expect(deepSeekImageTokens(width!, height!)).toBeLessThanOrEqual(1024)
     }
   })
 
   it('prices small images at the documented scale-up floor', () => {
-    // Below roughly 384x384 total pixels the provider scales up, so a tiny
-    // square costs the same as a 384x384 one.
-    expect(deepSeekImageTokens(100, 100)).toBe(deepSeekImageTokens(384, 384))
+    // Below roughly 544x544 total pixels the provider scales up, so a tiny
+    // square costs the same as a 544x544 one.
+    expect(deepSeekImageTokens(100, 100)).toBe(deepSeekImageTokens(544, 544))
   })
 
-  it('clamps extreme width by the aspect-ratio bound', () => {
-    // Width beyond 8x height projects onto the same clamped grid.
-    expect(deepSeekImageTokens(9000, 1)).toBe(113)
-    expect(deepSeekImageTokens(8192, 100)).toBe(113)
+  it('solves a one-row grid for an extremely wide image', () => {
+    // Width-dominant aspect drives the solver's single-row branch; no
+    // aspect-ratio clamp applies, so the row fills the whole budget.
+    expect(deepSeekImageTokens(9000, 1)).toBe(1024)
+    expect(deepSeekImageTokens(8192, 100)).toBe(593)
   })
 
   it('solves a one-column grid for an extremely tall image', () => {
     // Height-dominant aspect drives the solver's single-column branch.
-    expect(deepSeekImageTokens(16, 8192)).toBe(381)
-    expect(deepSeekImageTokens(1, 9000)).toBe(381)
-  })
-
-  it('trims an odd solved grid height to the even row count', () => {
-    expect(deepSeekImageTokens(100, 4036)).toBe(253)
+    expect(deepSeekImageTokens(1, 9000)).toBe(1024)
+    expect(deepSeekImageTokens(16, 8192)).toBe(590)
   })
 
   it('converges through a second projection pass when the first is not a fixpoint', () => {
-    expect(deepSeekImageTokens(4921, 353)).toBe(289)
-    expect(deepSeekImageTokens(97, 7289)).toBe(245)
+    expect(deepSeekImageTokens(12, 1123)).toBe(380)
+    expect(deepSeekImageTokens(89, 2076)).toBe(254)
   })
 })

+ 6 - 6
packages/llm/llm-deepseek/tests/request-pricing.spec.ts

@@ -44,7 +44,7 @@ describe('DeepSeek request-image pricing', () => {
     const image = ref('photo', 1920, 1080)
     const prices = deepSeekImageRequestPricing(connection(), 'vision').priceImages([image])
     expect(prices).toEqual([{
-      visualTokens: 369,
+      visualTokens: 407,
       text: requestImageHandleText(image, { width: 1066, height: 600 }),
     }])
   })
@@ -55,7 +55,7 @@ describe('DeepSeek request-image pricing', () => {
       models: [{ ...VISION_MODEL, imagePixelBudget: 'low' as const }],
     })
     const prices = deepSeekImageRequestPricing(options, 'vision').priceImages([image])
-    expect(prices[0]!.visualTokens).toBe(201)
+    expect(prices[0]!.visualTokens).toBe(184)
   })
 
   it('builds handle and placeholder text through the supplied access resolution', () => {
@@ -68,7 +68,7 @@ describe('DeepSeek request-image pricing', () => {
     ).priceImages(images)
     expect(prices[0]).toEqual({ visualTokens: 0, text: offloadedImageText(images[0]!, access) })
     expect(prices[1]).toEqual({
-      visualTokens: 349,
+      visualTokens: 422,
       text: requestImageHandleText(images[1]!, { width: 800, height: 800 }, access),
     })
     expect(prices[1]?.text).toContain('/world/attachments/photo.png')
@@ -82,8 +82,8 @@ describe('DeepSeek request-image pricing', () => {
     ).priceImages(images)
     expect(prices).toEqual([
       { visualTokens: 0, text: offloadedImageText(images[0]!) },
-      { visualTokens: 349, text: requestImageHandleText(images[1]!, { width: 800, height: 800 }) },
-      { visualTokens: 349, text: requestImageHandleText(images[2]!, { width: 800, height: 800 }) },
+      { visualTokens: 422, text: requestImageHandleText(images[1]!, { width: 800, height: 800 }) },
+      { visualTokens: 422, text: requestImageHandleText(images[2]!, { width: 800, height: 800 }) },
     ])
   })
 
@@ -100,7 +100,7 @@ describe('DeepSeek request-image pricing', () => {
       connection({ maxRequestFilesBytes: 2 * 1024 * 1024, imageOffloadByteQuantum: 1 }),
       'vision',
     ).priceImages(images)
-    expect(prices.map(price => price.visualTokens)).toEqual([0, 349, 349])
+    expect(prices.map(price => price.visualTokens)).toEqual([0, 422, 422])
     expect(prices[0]!.text).toBe(offloadedImageText(images[0]!))
   })
 })