Forráskód Böngészése

feat(web): turn speed metrics and composer context meter

Assistant footers and the stats line gain TTFT/tok-per-second readings
folded from step timings; context occupancy moves off the stats line onto
a composer ring whose panel shows a heuristic system/tools/messages
breakdown from the new token-meter contextBreakdown session projection.
Yif 2 hónapja
szülő
commit
0073d6aaa1
45 módosított fájl, 1650 hozzáadás és 239 törlés
  1. 6 0
      .agents/notes/implemented/feature/2026-08-04-web-latency-throughput-metrics.i18n.yaml
  2. 29 0
      .agents/notes/implemented/feature/2026-08-04-web-latency-throughput-metrics.md
  3. 29 0
      .agents/notes/implemented/feature/2026-08-04-web-latency-throughput-metrics.zh.md
  4. 6 0
      .agents/notes/implemented/feature/2026-08-05-composer-context-meter-breakdown.i18n.yaml
  5. 31 0
      .agents/notes/implemented/feature/2026-08-05-composer-context-meter-breakdown.md
  6. 31 0
      .agents/notes/implemented/feature/2026-08-05-composer-context-meter-breakdown.zh.md
  7. 4 3
      docs/cordis-catalog/services.md
  8. 2 2
      docs/core-data-structures/session.i18n.yaml
  9. 2 9
      docs/core-data-structures/session.md
  10. 2 9
      docs/core-data-structures/session.zh.md
  11. 4 42
      packages/client/runtime/src/client/session-history/history-fold.ts
  12. 84 0
      packages/client/runtime/src/client/sessions/assistant-timing.ts
  13. 10 1
      packages/client/runtime/src/client/sessions/transcript-adapter.ts
  14. 44 0
      packages/client/runtime/tests/transcript-adapter.spec.ts
  15. 2 2
      packages/client/ui-conversation/README.i18n.yaml
  16. 0 1
      packages/client/ui-conversation/README.md
  17. 0 1
      packages/client/ui-conversation/README.zh.md
  18. 7 1
      packages/client/ui-conversation/src/client/chat/AssistantMarkdown.tsx
  19. 7 0
      packages/client/ui-conversation/src/client/chat/ChatView.tsx
  20. 18 2
      packages/client/ui-conversation/src/client/chat/MessageIconActions.tsx
  21. 66 18
      packages/client/ui-conversation/src/client/chat/StatsLine.tsx
  22. 21 0
      packages/client/ui-conversation/src/client/chat/message-chrome.ts
  23. 97 0
      packages/client/ui-conversation/src/client/chat/turn-metrics.ts
  24. 14 0
      packages/client/ui-conversation/src/client/locales.ts
  25. 141 0
      packages/client/ui-conversation/src/client/skeleton/ContextMeter.module.css
  26. 131 0
      packages/client/ui-conversation/src/client/skeleton/ContextMeter.tsx
  27. 2 0
      packages/client/ui-conversation/src/client/skeleton/InputBar.tsx
  28. 9 0
      packages/client/ui-conversation/tests/chat-branch-tails.spec.tsx
  29. 85 37
      packages/client/ui-conversation/tests/chat-stats-bash-sample.spec.tsx
  30. 39 0
      packages/client/ui-conversation/tests/chat-view.spec.tsx
  31. 93 0
      packages/client/ui-conversation/tests/context-meter.spec.tsx
  32. 13 2
      packages/client/ui-conversation/tests/gate-branch-tails.spec.tsx
  33. 154 0
      packages/client/ui-conversation/tests/turn-metrics.spec.ts
  34. 1 1
      packages/cordis/tool-cordis/src/api-catalog.ts
  35. 5 46
      packages/core/session/src/index.ts
  36. 51 0
      packages/core/session/src/surface.ts
  37. 2 2
      packages/llm/token-meter/README.i18n.yaml
  38. 4 2
      packages/llm/token-meter/README.md
  39. 4 2
      packages/llm/token-meter/README.zh.md
  40. 87 0
      packages/llm/token-meter/src/breakdown-projection.ts
  41. 87 0
      packages/llm/token-meter/src/estimate.ts
  42. 11 55
      packages/llm/token-meter/src/index.ts
  43. 19 0
      packages/llm/token-meter/src/projection.ts
  44. 1 1
      packages/llm/token-meter/src/types.ts
  45. 195 0
      packages/llm/token-meter/tests/context-breakdown-projection.spec.ts

+ 6 - 0
.agents/notes/implemented/feature/2026-08-04-web-latency-throughput-metrics.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-08-04-web-latency-throughput-metrics.md
+2026-08-04-web-latency-throughput-metrics.md: f891ce8d2f772c66d5affba2d89694d0eeb681b6
+2026-08-04-web-latency-throughput-metrics.zh.md: e62f5ad6341cffe052118871b57e27f5380c2fd1

+ 29 - 0
.agents/notes/implemented/feature/2026-08-04-web-latency-throughput-metrics.md

@@ -0,0 +1,29 @@
+# Agent Note: Web turn and window latency/throughput metrics
+
+Status: implemented
+
+English | [中文](2026-08-04-web-latency-throughput-metrics.zh.md)
+
+## Problem
+
+The Web chat records per-step LLM timing (`stepStartTime` / `firstTokenTime` / `completedTime`) and per-step usage, and the trajectory view exposes them per step, but the chat surface answers neither "how responsive was this turn" nor "how fast is this session going": the assistant footer shows only the turn wall time, and the stats line folds only wall-time totals.
+
+## Decision
+
+A package-local fold, `ui-conversation`'s `chat/turn-metrics.ts`, is the single derivation from assistant nodes to latency/throughput readings. `assistantStepReading` turns one node into a step reading: TTFT needs both `stepStartTime` and `firstTokenTime`, decode span needs `firstTokenTime`, negative spans clamp to zero, and output tokens come from the untrusted `usage` value only when they are finite and non-negative. `deriveTurnMetrics` folds readings per turn: the lowest-numbered step owns the turn's TTFT slot, and throughput divides the summed output tokens by the summed decode spans over exactly the steps carrying both, so an unsampled step drops out instead of skewing the ratio; a turn with neither figure emits no entry.
+
+The assistant footer appends the readings to the existing hover-revealed time chrome after `Ran for`, as `TTFT {s}s · {tps} tok/s`, each omitted independently when unrecorded. ChatView shows a turn's readings only when that turn's `turnTimings` entry has an `endTime`: the loaded window is a contiguous log suffix, so an in-window settled turn carries every one of its steps and the first-step TTFT is genuine rather than a window artifact. `formatLatencySeconds` is unit-less so each locale template owns its second suffix (`TTFT {seconds}s` / `首 token {seconds}秒`).
+
+The stats line reuses the same step reading in its window fold: `deriveStats` accumulates TTFT sum/count and decode span/tokens, rendering a `TTFT avg … · … tok/s` group beside the LLM/tool wall times. Like those wall times the group is window-scoped and folds no billing; token accounting stays on the token-meter projections.
+
+## Alternatives considered
+
+**A durable session projection (token-meter shape).** A `ProjectionDefinition` folding step timings host-side would survive compaction and window paging and cover the whole log. Deferred, not rejected: projection state must stay O(1) (averages, not percentiles), it needs a host change plus a schema, and the chat stats line is already documented as window-scoped for its duration facts — the new group joins that scope. A later PR can add the durable projection without moving these readings.
+
+**Per-step footer chrome.** Showing each assistant message its own TTFT would attach chrome to mid-turn narration nodes, which the footer design deliberately keeps chrome-free; the trajectory view already exposes per-step timing detail.
+
+**Gating footer metrics on node presence instead of `turn/end` timing.** Rendering whatever steps happen to be loaded would show a plausible-looking TTFT that is actually the first *loaded* step after paging. The `endTime` gate plus the suffix-window invariant makes the displayed figure the turn's true first-step latency or nothing.
+
+## Consequences
+
+A settled in-window turn's footer reveals `TTFT`/`tok/s` on hover after the wall time, and the stats line shows window-average latency and throughput beside its wall times, all without new session events or host changes. Metrics degrade by omission: providers or steps without timing or usage samples drop individual figures rather than rendering zeros. Older history outside the loaded window stays uncounted, recorded in the package README's stats-line limitation.

+ 29 - 0
.agents/notes/implemented/feature/2026-08-04-web-latency-throughput-metrics.zh.md

@@ -0,0 +1,29 @@
+# Agent Note: Web 轮次与窗口级延迟/吞吐指标
+
+Status: implemented
+
+[English](2026-08-04-web-latency-throughput-metrics.md) | 中文
+
+## 问题
+
+Web 聊天已经记录了逐步骤的 LLM 计时(`stepStartTime`/`firstTokenTime`/`completedTime`)和逐步骤 usage,trajectory 视图也按步骤展示它们,但聊天界面既回答不了「这一轮响应有多快」,也回答不了「这个会话跑得有多快」:assistant 页脚只显示轮次实际耗时,统计行也只折算墙钟时间总量。
+
+## 决策
+
+包内折算 `ui-conversation` 的 `chat/turn-metrics.ts` 是从 assistant 节点推导延迟/吞吐读数的唯一位置。`assistantStepReading` 把一个节点转成一次步骤读数:TTFT(首 token 延迟)需要 `stepStartTime` 与 `firstTokenTime` 同时存在,解码时长需要 `firstTokenTime`,负时长收敛为零,输出 token 数只在不可信的 `usage` 值有限且非负时才采纳。`deriveTurnMetrics` 按轮次折算读数:编号最小的步骤拥有该轮次的 TTFT 槽位,吞吐用「同时携带两者的那些步骤」的输出 token 总和除以解码时长总和,因此缺采样的步骤直接退出而不是让比值失真;两个数字都没有的轮次不产生条目。
+
+assistant 页脚把读数追加到既有 hover 显示的时间附属元素中、`用时` 之后,形如 `首 token {s}秒 · {tps} tok/s`,未记录的数字各自省略。ChatView 仅在该轮次的 `turnTimings` 条目带有 `endTime` 时才显示读数:已加载窗口是日志的连续后缀,因此窗口内已结算的轮次必然带着它的全部步骤,首步 TTFT 是真实值而非窗口截断的产物。`formatLatencySeconds` 不带单位,各语言模板各自拥有秒后缀(`TTFT {seconds}s`/`首 token {seconds}秒`)。
+
+统计行在其窗口折算中复用同一份步骤读数:`deriveStats` 累计 TTFT 总和/计数与解码时长/token 数,在 LLM/工具墙钟时间旁渲染 `TTFT avg … · … tok/s` 分组。与那些墙钟时间一样,该分组是窗口作用域的,不折算任何计费;token 账目仍归 token-meter 投影。
+
+## 考虑过的替代方案
+
+**持久的会话投影(token-meter 形态)。** 在 host 侧用 `ProjectionDefinition` 折算步骤计时可以跨越压缩与窗口分页、覆盖整个日志。是暂缓而非否决:投影状态必须保持 O(1)(只能均值,不能分位数),它需要 host 改动加 schema,而聊天统计行的耗时事实本就被记录为窗口作用域——新分组沿用该作用域。后续 PR 可以在不挪动这些读数的情况下补上持久投影。
+
+**逐步骤页脚附属元素。** 让每条 assistant 消息显示自己的 TTFT,会给轮次中段的叙述节点挂上附属元素,而页脚设计刻意让它们保持无 chrome;trajectory 视图已经暴露逐步骤计时细节。
+
+**用节点是否在场而非 `turn/end` 计时做页脚门控。** 直接渲染碰巧加载到的步骤,会展示一个貌似合理、实为分页后「首个已加载步骤」的 TTFT。`endTime` 门控加上后缀窗口不变量,使显示的数字要么是该轮次真实的首步延迟,要么什么都不显示。
+
+## 后果
+
+窗口内已结算轮次的页脚在 hover 时于实际耗时之后显示 `首 token`/`tok/s`,统计行在墙钟时间旁显示窗口平均延迟与吞吐,全程不新增会话事件、不改 host。指标以省略的方式退化:没有计时或 usage 采样的提供方或步骤只是丢掉对应数字,而不会渲染成零。已加载窗口之外的更早历史仍不计入,已记录在包 README 的统计行限制中。

+ 6 - 0
.agents/notes/implemented/feature/2026-08-05-composer-context-meter-breakdown.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-08-05-composer-context-meter-breakdown.md
+2026-08-05-composer-context-meter-breakdown.md: 2879030aa7859385b2e6026fbf87da4548c392bd
+2026-08-05-composer-context-meter-breakdown.zh.md: 3aeff68ebd157c4104fc085ad19e1738ead8f983

+ 31 - 0
.agents/notes/implemented/feature/2026-08-05-composer-context-meter-breakdown.md

@@ -0,0 +1,31 @@
+# Agent Note: Composer context meter with heuristic composition breakdown
+
+Status: implemented
+
+English | [中文](2026-08-05-composer-context-meter-breakdown.zh.md)
+
+## Problem
+
+The Web chat's stats line showed context occupancy as one inline figure (`Context N% of X`) among its billing groups. That answers "how full" but not "what fills it": nothing showed how the window divides between the system prompt, tool schemas, and conversation, and the one-line row has no room for that detail. The available numbers also live in two vocabularies — the provider-exact billed prompt size from `contextPressure` versus the token-meter's fixed character heuristic — and no existing surface could present composition without conflating them.
+
+## Decision
+
+Three cooperating pieces, one per package boundary:
+
+`dsh-session` exports the pure `deriveEventMessage(event)` (previously reachable only as a `Session` method, which now delegates to it) so a host-side fold can price surface nodes without a `Session` instance.
+
+`dsh-token-meter` extracts its pricing heuristic into `src/estimate.ts` — shared verbatim by the measurement service — and registers a third session projection, `contextBreakdown`, carrying `systemTokens` / `toolsTokens` / `messageTokens`. Envelope figures reprice last-wins on each `request/header` through `canonicalHeader`; the message figure folds surface appends and positional replacements over a per-node `{seq, tokens}` list, so compaction shrinks it the same way it shrinks the next request. A replace range absent from the folded surface throws: committed logs are surface-validated at append time, so an unresolvable range is log corruption, not a skippable event.
+
+`ui-conversation` moves context occupancy off the stats line (one home per fact) onto a composer-trailing `ContextMeter`: a 14px occupancy ring after the model seat fed by `contextPressure`, click-opening a panel that pairs the provider-exact percent and `~used / capacity` header with a 4px color-segmented bar and `~`-prefixed composition rows. The two vocabularies deliberately never reconcile — the ring, header, and bar length stay provider-exact while the heuristic shares only proportion the bar's colored segments and rows, each marked `~` because the fixed 4-chars-per-token heuristic systematically underprices CJK text and code.
+
+## Alternatives considered
+
+**Deriving composition client-side from the loaded window.** The window is a contiguous log suffix: the `request/header` events carrying the system prompt and tool schemas may sit outside it, and paging would silently change the figures. Only a durable host-side projection survives paging and compaction, which is why the data crosses the wire as a third projection rather than a chat-window fold.
+
+**Scaling the heuristic rows to sum to `pressureTokens`.** Forced reconciliation fabricates precision: pressure lags one request, includes provider envelope overhead the estimator never models, and would make the rows move when nothing in the composition changed. Showing the estimator's real vocabulary with an explicit `~` was chosen instead.
+
+**Finer categories (rules, skills, MCP tools) as in Claude Code's `/context`.** Not separable here: the harness folds those contributions into the system text and the tools list before the request header exists, so three categories are the honest resolution.
+
+## Consequences
+
+Token-meter now registers three projection keys; unloading removes all three, and `contextBreakdown` restores from JSON checkpoints (`stateVersion` 1). The stats line dropped its Context group and the ring is the sole context UI. The panel's heuristic rows visibly disagree with the provider-exact header — accepted and signposted by the `~` prefix; improving estimate accuracy (for example CJK-aware weighting) is localized to `estimate.ts` and changes no seam. The legend's purple segment tint is a literal color because the design platform ships no purple static token.

+ 31 - 0
.agents/notes/implemented/feature/2026-08-05-composer-context-meter-breakdown.zh.md

@@ -0,0 +1,31 @@
+# Agent Note: composer 上下文占用圆环与启发式组成明细
+
+Status: implemented
+
+[English](2026-08-05-composer-context-meter-breakdown.md) | 中文
+
+## 问题
+
+Web 聊天的统计行把上下文占用率作为一个行内数字(`Context N% of X`)挤在计费分组之间。它回答了「有多满」,却回答不了「被什么占满」:没有任何地方展示窗口在系统提示词、工具 schema 与对话之间如何分配,而单行统计行也容纳不下这种明细。可用的数字还分属两套口径——来自 `contextPressure` 的提供方精确计费 prompt 规模,与 token-meter 的固定字符启发式——没有任何既有界面能在不混淆两者的前提下展示组成。
+
+## 决定
+
+三个协作部分,每个包边界一个:
+
+`dsh-session` 导出纯函数 `deriveEventMessage(event)`(此前只能通过 `Session` 方法访问,该方法现在委托给它),使 host 侧 fold 无需 `Session` 实例即可为表层节点计价。
+
+`dsh-token-meter` 把计价启发式抽取到 `src/estimate.ts`(与测量服务逐字共享),并注册第三个会话投影 `contextBreakdown`,携带 `systemTokens` / `toolsTokens` / `messageTokens`。envelope 数字在每条 `request/header` 上经 `canonicalHeader` 按后者胜重新计价;消息数字用逐节点 `{seq, tokens}` 列表折叠表层追加与位置替换,因此压缩会像缩小下一个请求那样缩小它。折叠表层中不存在的替换范围会直接抛出:已提交日志在追加时就经过表层校验,无法解析的范围是日志损坏,而不是可跳过的事件。
+
+`ui-conversation` 把上下文占用率从统计行移走(一个事实一个家),放到 composer 尾部的 `ContextMeter`:模型座位之后的一枚 14px 占用圆环,由 `contextPressure` 供数,点击弹出的面板把提供方精确的百分比与 `~已用 / 容量` 标题,与 4px 分色分段进度条及带 `~` 前缀的组成明细行并列。两套口径刻意永不对账——圆环、标题与进度条总长保持提供方精确值,启发式占比只用于切分进度条的彩色分段与明细行,且每个启发式数字都标 `~`,因为固定的「4 字符≈1 token」启发式会系统性低估 CJK 文本与代码。
+
+## 备选方案
+
+**在客户端从已加载窗口推导组成。** 窗口是日志的连续后缀:携带系统提示词与工具 schema 的 `request/header` 事件可能在窗口之外,翻页还会让数字悄悄变化。只有持久的 host 侧投影能在翻页与压缩后幸存,这正是数据以第三个投影而非聊天窗口 fold 的形式过线的原因。
+
+**把启发式明细行按比例缩放到与 `pressureTokens` 相加一致。** 强行对账是在捏造精度:压力滞后一个请求,还包含估算器从不建模的提供方封装开销,会让明细行在组成毫无变化时也跟着变动。最终选择以显式 `~` 展示估算器的真实口径。
+
+**更细的类别(rules、skills、MCP 工具,如 Claude Code 的 `/context`)。** 在这里不可分:harness 在请求标头存在之前就把这些贡献折入系统文本与工具列表,因此三个类别是诚实的分辨率。
+
+## 后果
+
+token-meter 现在注册三个投影键;卸载会移除全部三个,`contextBreakdown` 可从 JSON 检查点恢复(`stateVersion` 为 1)。统计行删除了 Context 分组,圆环成为唯一的上下文 UI。面板的启发式明细行与提供方精确的标题数字肉眼可见地不一致——已接受并以 `~` 前缀标示;提升估算精度(例如按 CJK 加权)只需改动 `estimate.ts`,不涉及任何 seam。图例的紫色分段色值是字面量,因为设计平台没有紫色静态 token。

+ 4 - 3
docs/cordis-catalog/services.md

@@ -1681,7 +1681,7 @@ fork(source: SessionForkSource, boundary?: number, childSessionId?: SessionId):
 
 Types: [CreateSessionOptions](../core-data-structures/persistence.md) · [Session](../core-data-structures/session.md) · [SessionId](../core-data-structures/core.md)
 
-Source: [`packages/core/session/src/index.ts:767`](../../packages/core/session/src/index.ts)
+Source: [`packages/core/session/src/index.ts:726`](../../packages/core/session/src/index.ts)
 
 ## `ctx.sessionTitle` — `SessionTitleService`
 
@@ -2307,7 +2307,8 @@ Replay owner for one service-wide estimator and isolated per-session folds.
 measure(session: Session, requestHeader?: EpochHeader): TokenMeasurement
 
 /**
- * Heuristically price one model-visible message.
+ * Heuristically price one model-visible message (instance face of the pure
+ * {@link estimateMessage}).
  * @param message - message to price without mutation.
  * @returns content and role-framing tokens under the fixed service heuristic.
  */
@@ -2316,7 +2317,7 @@ estimateMessage(message: Message): number
 
 Types: [EpochHeader](../core-data-structures/session.md) · [Message](../core-data-structures/core.md) · [Session](../core-data-structures/session.md) · [TokenMeasurement](../core-data-structures/token-meter.md)
 
-Source: [`packages/llm/token-meter/src/index.ts:85`](../../packages/llm/token-meter/src/index.ts)
+Source: [`packages/llm/token-meter/src/index.ts:78`](../../packages/llm/token-meter/src/index.ts)
 
 ## `ctx.toolResultPrune` — `ToolResultPruneService`
 

+ 2 - 2
docs/core-data-structures/session.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write docs/core-data-structures/session.md
-session.md: 6369c956df03c0d786db696c208b000300d5bfbc
-session.zh.md: c39382b7e6c9b14f91c311cc80526a6fd8898e4c
+session.md: 0d32b47585e01dd992167b6580245c5444f3891b
+session.zh.md: ad3fb474c4f9c68ef33bf9f1ddea0a1a64b31de6

+ 2 - 9
docs/core-data-structures/session.md

@@ -487,15 +487,8 @@ declare class Session {
    */
   deriveMessages(): Message[];
   /**
-   * Project a single event into the LLM message it derives to, or null when
-   * it produces none — a non-surface event (chunk, boundary, log-only record)
-   * or an empty-content assistant/message (which exists only to host usage).
-   * The per-node pure function {@link deriveMessages} folds over the surface;
-   * an external reconstructor (or the dev invariant) folds the same function
-   * over a log prefix's surface to rebuild the exact messages any request was
-   * built from (the reconstructability Agent Note). The returned message is
-   * the already frozen message nested in the event wrapper and shared by
-   * delivery, durable history, and model requests.
+   * Instance face of the pure per-node projection rule
+   * {@link deriveEventMessage} (see its contract in `surface.ts`).
    * @param event - the event to project.
    * @returns the derived message, or null when the event produces none.
    */

+ 2 - 9
docs/core-data-structures/session.zh.md

@@ -489,15 +489,8 @@ declare class Session {
    */
   deriveMessages(): Message[];
   /**
-   * Project a single event into the LLM message it derives to, or null when
-   * it produces none — a non-surface event (chunk, boundary, log-only record)
-   * or an empty-content assistant/message (which exists only to host usage).
-   * The per-node pure function {@link deriveMessages} folds over the surface;
-   * an external reconstructor (or the dev invariant) folds the same function
-   * over a log prefix's surface to rebuild the exact messages any request was
-   * built from (the reconstructability Agent Note). The returned message is
-   * the already frozen message nested in the event wrapper and shared by
-   * delivery, durable history, and model requests.
+   * Instance face of the pure per-node projection rule
+   * {@link deriveEventMessage} (see its contract in `surface.ts`).
    * @param event - the event to project.
    * @returns the derived message, or null when the event produces none.
    */

+ 4 - 42
packages/client/runtime/src/client/session-history/history-fold.ts

@@ -16,6 +16,8 @@ import type {
 } from '../sessions/conversation-context.ts'
 import type { ConversationPromptSnapshot } from '../sessions/request-inspection.ts'
 import { PartialAccumulator } from '../sessions/partial.ts'
+import type { AssistantStepMetadata } from '../sessions/assistant-timing.ts'
+import { indexAssistantStepTiming, settledAssistantTiming } from '../sessions/assistant-timing.ts'
 
 interface CallIndexEntry {
   name: string
@@ -30,11 +32,6 @@ interface FoldedContext {
   originSeq?: number
 }
 
-interface AssistantStepMetadata {
-  stepStartTime: number | null
-  firstTokenTime: number | null
-}
-
 /** Immutable conversation projections derived only from the history source. */
 export interface ConversationHistoryProjection {
   eventNodes: readonly ConversationNode[]
@@ -45,10 +42,6 @@ export interface ConversationHistoryProjection {
   codeDispatches: ReadonlyMap<string, readonly CodeSubCall[]>
 }
 
-function assistantStepKey(turn: number, step: number): string {
-  return `${turn}\u0000${step}`
-}
-
 // Trajectory owns surface-window reconstruction so its immutable ledger does
 // not depend on Chat's live fold adapter or Session's mutable state.
 /* jscpd:ignore-start */
@@ -72,18 +65,6 @@ function contextOriginKind(event: SessionEvent | undefined): ConversationContext
   return 'rewrite'
 }
 
-function isTokenDelta(chunk: SessionEvent<'assistant/chunk'>['data']['chunk']): boolean {
-  switch (chunk.type) {
-    case 'text-delta':
-    case 'reasoning-delta':
-      return chunk.text !== ''
-    case 'tool-call-delta':
-      return chunk.argumentsDelta !== '' || chunk.name !== undefined
-    default:
-      return false
-  }
-}
-
 function foldContexts(events: readonly SessionEvent[]): readonly FoldedContext[] {
   const replay: SessionEvent[] = []
   const surface = new SurfaceManager(replay)
@@ -362,6 +343,7 @@ export function projectConversationHistory(
       contextGeneration++
       if (activePrompt !== undefined) promptsByContext.set(contextGeneration, activePrompt)
     }
+    indexAssistantStepTiming(assistantSteps, event)
     if (event.type === 'request/header') {
       activeRequestConfig = event.data.header.config
       activePrompt = {
@@ -370,30 +352,10 @@ export function projectConversationHistory(
         tools: event.data.header.tools ?? [],
       }
       promptsByContext.set(contextGeneration, activePrompt)
-    } else if (event.type === 'step/start') {
-      assistantSteps.set(
-        assistantStepKey(event.data.turn, event.data.step),
-        { stepStartTime: event.time, firstTokenTime: null },
-      )
-    } else if (event.type === 'assistant/chunk' && isTokenDelta(event.data.chunk)) {
-      const key = assistantStepKey(event.data.turn, event.data.step)
-      const current = assistantSteps.get(key) ?? {
-        stepStartTime: null,
-        firstTokenTime: null,
-      }
-      if (current.firstTokenTime === null) {
-        assistantSteps.set(key, { ...current, firstTokenTime: event.time })
-      }
     } else if (event.type === 'assistant/message') {
       assistantTimings.set(
         event.seq,
-        {
-          ...(assistantSteps.get(assistantStepKey(event.data.turn, event.data.step)) ?? {
-            stepStartTime: null,
-            firstTokenTime: null,
-          }),
-          completedTime: event.time,
-        },
+        settledAssistantTiming(assistantSteps, event.data.turn, event.data.step, event.time),
       )
       if (activeRequestConfig !== undefined) {
         assistantRequestConfigs.set(event.seq, activeRequestConfig)

+ 84 - 0
packages/client/runtime/src/client/sessions/assistant-timing.ts

@@ -0,0 +1,84 @@
+// Shared assistant step-timing fold: both transcript projections (the live
+// window adapter and the trajectory history fold) derive AssistantTiming from
+// the same step/start -> first token delta -> assistant/message sequence, so
+// the derivation lives once here instead of drifting per projection.
+
+import type { SessionEvent } from '@deepseek-ai/dsh-session/types'
+import type { AssistantTiming } from './conversation.ts'
+
+/** Pre-finalize timing boundaries for one assistant step (start + first token). */
+export interface AssistantStepMetadata {
+  stepStartTime: number | null
+  firstTokenTime: number | null
+}
+
+/**
+ * Composite map key for one assistant step.
+ * @param turn - turn number from the event payload.
+ * @param step - step number from the event payload.
+ * @returns collision-free `turn`/`step` key (NUL separator).
+ */
+export function assistantStepKey(turn: number, step: number): string {
+  return `${turn}\u0000${step}`
+}
+
+/**
+ * Whether a chunk carries visible model output (first-token boundary). Empty
+ * deltas (heartbeats, empty tool-call frames) do not count as a first token.
+ * @param chunk - the assistant/chunk payload.
+ * @returns true when the chunk contains a non-empty text/reasoning/tool delta.
+ */
+export function isTokenDelta(chunk: SessionEvent<'assistant/chunk'>['data']['chunk']): boolean {
+  switch (chunk.type) {
+    case 'text-delta':
+    case 'reasoning-delta':
+      return chunk.text !== ''
+    case 'tool-call-delta':
+      return chunk.argumentsDelta !== '' || chunk.name !== undefined
+    default:
+      return false
+  }
+}
+
+/**
+ * Fold one event into the per-step timing index: step/start opens the entry,
+ * the first non-empty token delta stamps first-token time once. Other event
+ * types are no-ops.
+ * @param steps - the mutable per-step index, keyed by {@link assistantStepKey}.
+ * @param event - the raw window event.
+ */
+export function indexAssistantStepTiming(steps: Map<string, AssistantStepMetadata>, event: SessionEvent): void {
+  if (event.type === 'step/start') {
+    steps.set(
+      assistantStepKey(event.data.turn, event.data.step),
+      { stepStartTime: event.time, firstTokenTime: null },
+    )
+  } else if (event.type === 'assistant/chunk' && isTokenDelta(event.data.chunk)) {
+    const key = assistantStepKey(event.data.turn, event.data.step)
+    const current = steps.get(key) ?? { stepStartTime: null, firstTokenTime: null }
+    if (current.firstTokenTime === null) {
+      steps.set(key, { ...current, firstTokenTime: event.time })
+    }
+  }
+}
+
+/**
+ * Settle one finalized assistant message's timing from its step entry; a step
+ * whose start or first token fell outside the window yields null boundaries.
+ * @param steps - the per-step index built by {@link indexAssistantStepTiming}.
+ * @param turn - the assistant/message turn number.
+ * @param step - the assistant/message step number.
+ * @param completedTime - the assistant/message event timestamp (epoch ms).
+ * @returns the node-ready timing record.
+ */
+export function settledAssistantTiming(
+  steps: ReadonlyMap<string, AssistantStepMetadata>,
+  turn: number,
+  step: number,
+  completedTime: number,
+): AssistantTiming {
+  return {
+    ...(steps.get(assistantStepKey(turn, step)) ?? { stepStartTime: null, firstTokenTime: null }),
+    completedTime,
+  }
+}

+ 10 - 1
packages/client/runtime/src/client/sessions/transcript-adapter.ts

@@ -22,6 +22,8 @@ import type { COMPACT_CHECKPOINT_SOURCE } from '@deepseek-ai/dsh-compact/checkpo
 import type { ToolCallView, ToolEventView, ToolResultView } from '@deepseek-ai/dsh-client-connection/client'
 import type { CommandNode, CompactionSummaryNode, ConversationNode } from './conversation.ts'
 import { toAssistantBlocks } from './conversation.ts'
+import type { AssistantStepMetadata } from './assistant-timing.ts'
+import { indexAssistantStepTiming, settledAssistantTiming } from './assistant-timing.ts'
 
 /**
  * The compaction seam's checkpoint plugin, pinned to the seam's own declaration
@@ -50,6 +52,7 @@ function materializeNode(
   event: SessionEvent,
   callIndex: ReadonlyMap<string, CallIndexEntry>,
   resultView: ToolResultView | null,
+  stepTimings: ReadonlyMap<string, AssistantStepMetadata>,
 ): ConversationNode {
   switch (event.type) {
     case 'user/message':
@@ -71,6 +74,7 @@ function materializeNode(
         kind: 'assistant', seq: event.seq, time: event.time,
         turn: event.data.turn, step: event.data.step,
         blocks: toAssistantBlocks(event.data.message.content), usage: event.data.usage,
+        timing: settledAssistantTiming(stepTimings, event.data.turn, event.data.step, event.time),
       }
     case 'steering/message':
       return {
@@ -177,6 +181,8 @@ export class TranscriptAdapter {
   /** Transcript nodes in log order; copy-on-write so a published array never mutates. */
   private projected: ConversationNode[] = []
   private callIdx = new Map<string, CallIndexEntry>()
+  /** Per-step timing boundaries (step/start + first token delta), consumed when the step's assistant/message materializes. */
+  private stepTimings = new Map<string, AssistantStepMetadata>()
   /** Wire result views keyed by the tool/result event's seq (views ride the envelope, not the event). */
   private resultViews = new Map<number, ToolResultView>()
   /**
@@ -207,6 +213,7 @@ export class TranscriptAdapter {
     this.callIdx = new Map()
     this.resultViews.clear()
     this.commandIdx = new Map()
+    this.stepTimings = new Map()
     for (let i = 0; i < events.length; i++) {
       const event = events[i]
       /* v8 ignore next -- dense-array guard: i stays within events.length, so the undefined arm needs a sparse array no caller builds. */
@@ -214,6 +221,7 @@ export class TranscriptAdapter {
       this.eventIndex.set(event.seq, event)
       this.indexCall(event, views?.[i])
       this.indexCommand(event)
+      indexAssistantStepTiming(this.stepTimings, event)
     }
     // Indexes first, then project: a tool/result materializes against the
     // complete call index, and a checkpoint against the complete event index.
@@ -236,6 +244,7 @@ export class TranscriptAdapter {
   append(event: SessionEvent, view?: ToolEventView): void {
     this.eventIndex.set(event.seq, event)
     this.indexCall(event, view)
+    indexAssistantStepTiming(this.stepTimings, event)
     if (this.indexCommand(event)) this.rev++
     if (!isTranscriptEvent(event)) return
     this.projected = [...this.projected, this.materialize(event)]
@@ -274,7 +283,7 @@ export class TranscriptAdapter {
   private materialize(event: SessionEvent): ConversationNode {
     return isCompactCheckpoint(event)
       ? materializeCompaction(event, this.eventIndex)
-      : materializeNode(event, this.callIdx, this.resultViews.get(event.seq) ?? null)
+      : materializeNode(event, this.callIdx, this.resultViews.get(event.seq) ?? null, this.stepTimings)
   }
 
   /**

+ 44 - 0
packages/client/runtime/tests/transcript-adapter.spec.ts

@@ -415,4 +415,48 @@ describe('TranscriptAdapter', () => {
       expect(nodes[1]).toMatchObject({ name: 'compact', outcome: { kind: 'success', text: '已压缩' } })
     })
   })
+
+  describe('assistant timing', () => {
+    const base = 1_700_000_000_000
+
+    it('derives step timing across a window rebuild (start + first token + completion)', () => {
+      const adapter = new TranscriptAdapter()
+      adapter.reset([
+        ev.turnStart(0, 0),
+        ev.user(1, '问'),
+        ev.stepStart(2, 0),
+        ev.chunkStart(3, 0),
+        ev.chunkText(4, 0, '答'),
+        ev.chunkText(5, 0, '案'),
+        ev.assistant(6, 0, '答案'),
+        ev.turnEnd(7, 0),
+      ])
+      const assistant = adapter.nodes().find(n => n.kind === 'assistant')
+      expect(assistant).toMatchObject({
+        timing: { stepStartTime: base + 2, firstTokenTime: base + 4, completedTime: base + 6 },
+      })
+    })
+
+    it('derives the same timing on the live append path, first token winning once', () => {
+      const adapter = new TranscriptAdapter()
+      adapter.reset([ev.user(0, '问')])
+      adapter.append(ev.stepStart(1, 0))
+      adapter.append(ev.chunkText(2, 0, '首'))
+      adapter.append(ev.chunkText(3, 0, '次'))
+      adapter.append(ev.assistant(4, 0, '首次'))
+      const assistant = adapter.nodes().find(n => n.kind === 'assistant')
+      expect(assistant).toMatchObject({
+        timing: { stepStartTime: base + 1, firstTokenTime: base + 2, completedTime: base + 4 },
+      })
+    })
+
+    it('soft-falls to null boundaries when the step opening fell outside the window', () => {
+      const adapter = new TranscriptAdapter()
+      adapter.reset([ev.assistant(100, 0, '被切窗的答案')])
+      const assistant = adapter.nodes().find(n => n.kind === 'assistant')
+      expect(assistant).toMatchObject({
+        timing: { stepStartTime: null, firstTokenTime: null, completedTime: base + 100 },
+      })
+    })
+  })
 })

+ 2 - 2
packages/client/ui-conversation/README.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write packages/client/ui-conversation/README.md
-README.md: 7d3a4b5fe07cc8858c2f2059e6f65b6e27602b4d
-README.zh.md: d93af91381157bb4e8e4b6a14ad00edecd505246
+README.md: 737a53da1fa4541f81d89b41945cacd2b415ef4a
+README.zh.md: 800d2b3e762d410100c8c9eb64bfbcdb22cda62a

A különbségek nem kerülnek megjelenítésre, a fájl túl nagy
+ 0 - 1
packages/client/ui-conversation/README.md


A különbségek nem kerülnek megjelenítésre, a fájl túl nagy
+ 0 - 1
packages/client/ui-conversation/README.zh.md


+ 7 - 1
packages/client/ui-conversation/src/client/chat/AssistantMarkdown.tsx

@@ -30,6 +30,10 @@ export interface AssistantMarkdownProps {
   /** Turn wall time in ms for the IconActions run-time label; omitted when the
    *  turn's triggering input is outside the loaded window. */
   runMs?: number | undefined
+  /** Turn first-step TTFT in ms for the IconActions label; omitted when unrecorded. */
+  ttftMs?: number | undefined
+  /** Turn decode throughput for the IconActions label; omitted when unrecorded. */
+  tokensPerSecond?: number | undefined
   /** Event sequence used as the fork boundary; omitted while streaming. */
   seq?: number | undefined
   /** Fork the session through this finalized message's completed turn when eligible. */
@@ -82,7 +86,7 @@ function ThinkRow({ text, running, t }: { text: string; running: boolean; t: Ass
 }
 
 export const AssistantMarkdown = memo(function AssistantMarkdown({
-  blocks, streaming, interrupted, time, runMs, seq, onFork, forkUnavailable, t,
+  blocks, streaming, interrupted, time, runMs, ttftMs, tokensPerSecond, seq, onFork, forkUnavailable, t,
 }: AssistantMarkdownProps) {
   // Stable per locale revision (t identity changes on switch): a fresh object
   // per render would rebuild MarkdownText's component table every chunk.
@@ -125,6 +129,8 @@ export const AssistantMarkdown = memo(function AssistantMarkdown({
           text={copyText(blocks)}
           time={time}
           runMs={runMs}
+          ttftMs={ttftMs}
+          tokensPerSecond={tokensPerSecond}
           clock="end"
           onBranch={onFork === undefined || seq === undefined ? undefined : () => { onFork(seq) }}
           branchUnavailable={forkUnavailable}

+ 7 - 0
packages/client/ui-conversation/src/client/chat/ChatView.tsx

@@ -36,6 +36,7 @@ import { GenericCommandCard } from './GenericCommandCard.tsx'
 import { GenericToolCard } from './GenericToolCard.tsx'
 import { MessageItem, PendingSteeringBubble } from './MessageItem.tsx'
 import { formatRunDuration } from './message-chrome.ts'
+import { deriveTurnMetrics } from './turn-metrics.ts'
 import css from './ChatView.module.css'
 
 const FOLLOW_THRESHOLD = 24
@@ -362,6 +363,7 @@ export function ChatView({
   const actionSeqs = useMemo(() => assistantActionsSeqs(nodes), [nodes])
   const branchSeqs = useMemo(() => messageBranchSeqs(nodes, turnEnds), [nodes, turnEnds])
   const runningTurnStart = useMemo(() => runningTurnStartTime(turnTimings), [turnTimings])
+  const turnMetrics = useMemo(() => deriveTurnMetrics(nodes), [nodes])
 
   const listRef = useRef<HTMLDivElement | null>(null)
   const columnRef = useRef<HTMLDivElement | null>(null)
@@ -599,6 +601,9 @@ export function ChatView({
     const node: ConversationNode = item.node
     if (node.kind === 'assistant') {
       const timing = actionSeqs.has(node.seq) ? turnTimings.get(node.turn) : undefined
+      // Metrics gate on the settled in-window timing: turn/start loaded means
+      // every step of the turn is loaded, so first-step TTFT is genuine.
+      const metrics = timing?.endTime === undefined ? undefined : turnMetrics.get(node.turn)
       return (
         <AssistantMarkdown
           blocks={node.blocks}
@@ -608,6 +613,8 @@ export function ChatView({
           runMs={timing?.endTime === undefined
             ? undefined
             : Math.max(0, timing.endTime - timing.startTime)}
+          ttftMs={metrics?.ttftMs}
+          tokensPerSecond={metrics?.tokensPerSecond}
           seq={node.seq}
           onFork={forkAt}
           forkUnavailable={!branchSeqs.has(node.seq)}

+ 18 - 2
packages/client/ui-conversation/src/client/chat/MessageIconActions.tsx

@@ -6,7 +6,7 @@ import {
   IconBranchOutline16, IconCheckOutline16, IconCopyOutline16, Tooltip, writeClipboard,
 } from '@deepseek-ai/dsh-client-ui-primitives'
 import type { ChatViewSlotProps } from '../contract/slots.ts'
-import { formatMessageClock, formatRunDuration } from './message-chrome.ts'
+import { formatLatencySeconds, formatMessageClock, formatRunDuration, formatTokensPerSecond } from './message-chrome.ts'
 import { useCalendarDay } from './use-calendar-day.ts'
 import css from './MessageIconActions.module.css'
 
@@ -17,6 +17,10 @@ export interface MessageIconActionsProps {
   time?: number | undefined
   /** Turn wall time in ms, appended to the clock as `· Ran for 15s`; omitted when the turn's start is unknown. */
   runMs?: number | undefined
+  /** Turn first-step TTFT in ms, appended as `· TTFT 1.2s`; omitted when unrecorded. */
+  ttftMs?: number | undefined
+  /** Turn decode throughput, appended as `· 34 tok/s`; omitted when unrecorded. */
+  tokensPerSecond?: number | undefined
   /** Clock before icons (user) or after (assistant). */
   clock: 'start' | 'end'
   /** Fork the session at this message; omission hides the branch action. */
@@ -37,7 +41,7 @@ export interface MessageIconActionsProps {
  * @returns The actions row element.
  */
 export function MessageIconActions({
-  text, time, runMs, clock, onBranch, branchUnavailable = false, showBranch = true, className, t,
+  text, time, runMs, ttftMs, tokensPerSecond, clock, onBranch, branchUnavailable = false, showBranch = true, className, t,
 }: MessageIconActionsProps) {
   const day = useCalendarDay()
   const reasonId = useId()
@@ -76,6 +80,18 @@ export function MessageIconActions({
           {t('message.ranFor', { duration: formatRunDuration(runMs, t) })}
         </>
       )}
+      {ttftMs !== undefined && (
+        <>
+          <span className={css.runTimeDot} aria-hidden>·</span>
+          {t('message.ttft', { seconds: formatLatencySeconds(ttftMs) })}
+        </>
+      )}
+      {tokensPerSecond !== undefined && (
+        <>
+          <span className={css.runTimeDot} aria-hidden>·</span>
+          {t('message.tokensPerSecond', { tps: formatTokensPerSecond(tokensPerSecond) })}
+        </>
+      )}
     </span>
   )
   return (

+ 66 - 18
packages/client/ui-conversation/src/client/chat/StatsLine.tsx

@@ -2,10 +2,13 @@
 // Mounted on 'conversation.composer.dock' so it sticks with the composer in the
 // active conversation scrollport (see ConversationRoot data-conversation-scroll).
 
-import { Fragment, memo, useMemo } from 'react'
+import { Fragment, memo, useLayoutEffect, useMemo, useRef, useState } from 'react'
+import { Tooltip } from '@deepseek-ai/dsh-client-ui-primitives'
 import type { ConversationSnapshot, UseProjection } from '@deepseek-ai/dsh-client-runtime/client'
 import type { SnapshotSelectorHook } from '@deepseek-ai/dsh-client-ui-slots'
 import type { ContextPressureProjection, TokenUsageProjection } from '@deepseek-ai/dsh-token-meter/client'
+import { formatTokensPerSecond } from './message-chrome.ts'
+import { assistantStepReading } from './turn-metrics.ts'
 import css from './StatsLine.module.css'
 
 interface WindowStats {
@@ -15,6 +18,14 @@ interface WindowStats {
   llmMs: number
   /** Summed tool wall time (tool/call → tool/result); 0 when no pair is in-window. */
   toolMs: number
+  /** Summed first-token latency over `ttftSteps`; 0 when no step records it. */
+  ttftMs: number
+  /** Steps carrying a recorded TTFT. */
+  ttftSteps: number
+  /** Summed decode wall time over steps that also report output tokens. */
+  decodeMs: number
+  /** Summed output tokens over the same decode-timed steps. */
+  decodeTokens: number
 }
 
 /**
@@ -32,6 +43,10 @@ export function deriveStats(nodes: ConversationSnapshot['nodes']): WindowStats {
   let steps = 0
   let llmMs = 0
   let toolMs = 0
+  let ttftMs = 0
+  let ttftSteps = 0
+  let decodeMs = 0
+  let decodeTokens = 0
   for (const node of nodes) {
     if (node.kind === 'tool-result') {
       if (node.callTime !== null) toolMs += Math.max(0, node.time - node.callTime)
@@ -43,8 +58,17 @@ export function deriveStats(nodes: ConversationSnapshot['nodes']): WindowStats {
     if (node.timing !== undefined && node.timing.stepStartTime !== null) {
       llmMs += Math.max(0, node.timing.completedTime - node.timing.stepStartTime)
     }
+    const reading = assistantStepReading(node)
+    if (reading.ttftMs !== null) {
+      ttftMs += reading.ttftMs
+      ttftSteps += 1
+    }
+    if (reading.decodeMs !== null && reading.outputTokens !== null) {
+      decodeMs += reading.decodeMs
+      decodeTokens += reading.outputTokens
+    }
   }
-  return { turns: turns.size, steps, llmMs, toolMs }
+  return { turns: turns.size, steps, llmMs, toolMs, ttftMs, ttftSteps, decodeMs, decodeTokens }
 }
 
 /**
@@ -84,13 +108,18 @@ export function cacheHitPercent(usage: TokenUsageProjection): number | null {
     : Math.round(usage.cacheReadTokens / denominator * 100)
 }
 
-/** Sum the three disjoint prompt-side billing buckets. */
-function billedInputTokens(usage: TokenUsageProjection): number {
+/**
+ * Sum the three disjoint prompt-side billing buckets.
+ * @param usage - the session's token-usage projection value.
+ * @returns billed input tokens.
+ */
+export function billedInputTokens(usage: TokenUsageProjection): number {
   return usage.uncachedInputTokens + usage.cacheReadTokens + usage.cacheWriteTokens
 }
 
 interface ContextOccupancy {
   percent: number
+  pressureTokens: number
   contextWindow: number
 }
 
@@ -100,7 +129,7 @@ interface ContextOccupancy {
  * fields, so this is a reference figure rather than an exact measurement of one
  * request (see the token-meter README).
  * @param pressure - the session's context-pressure projection value.
- * @returns occupancy and its denominator, or null until both values are known.
+ * @returns occupancy with its numerator and denominator, or null until both values are known.
  */
 export function contextOccupancy(
   pressure: ContextPressureProjection | undefined,
@@ -108,6 +137,7 @@ export function contextOccupancy(
   if (pressure?.pressureTokens === undefined || pressure.contextWindow === undefined) return null
   return {
     percent: Math.min(100, Math.round(pressure.pressureTokens / pressure.contextWindow * 100)),
+    pressureTokens: pressure.pressureTokens,
     contextWindow: pressure.contextWindow,
   }
 }
@@ -121,7 +151,6 @@ export interface StatsLineProps {
 export const StatsLine = memo(function StatsLine({ useSession, useProjection }: StatsLineProps) {
   const nodes = useSession(s => s.nodes)
   const usage = useProjection('tokenUsage')
-  const pressure = useProjection('contextPressure')
   const stats = useMemo(() => deriveStats(nodes), [nodes])
   // Pipe-separated groups (figma stats strip); a group with no data drops out whole.
   const groups: string[] = []
@@ -131,11 +160,14 @@ export const StatsLine = memo(function StatsLine({ useSession, useProjection }:
     if (stats.llmMs > 0) durations.push(`LLM ${formatDuration(stats.llmMs)}`)
     if (stats.toolMs > 0) durations.push(`Tool call ${formatDuration(stats.toolMs)}`)
     if (durations.length > 0) groups.push(durations.join(' · '))
+    // Window-scoped like the wall times above: averages describe loaded steps.
+    const speeds: string[] = []
+    if (stats.ttftSteps > 0) speeds.push(`TTFT avg ${formatDuration(stats.ttftMs / stats.ttftSteps)}`)
+    if (stats.decodeMs > 0) speeds.push(`${formatTokensPerSecond(stats.decodeTokens / (stats.decodeMs / 1_000))} tok/s`)
+    if (speeds.length > 0) groups.push(speeds.join(' · '))
   }
-  const context = contextOccupancy(pressure)
-  if (context !== null) {
-    groups.push(`Context ${context.percent}% of ${formatTokens(context.contextWindow)}`)
-  }
+  // Context occupancy deliberately lives on the composer's ContextMeter ring,
+  // not here — one home per fact.
   // Billing rides the durable projection, so these survive paging and
   // compaction. Suppress the empty projection on a brand-new session.
   if (usage !== undefined
@@ -147,15 +179,31 @@ export const StatsLine = memo(function StatsLine({ useSession, useProjection }:
       + ` · Output ${formatTokens(usage.outputTokens)} tok`,
     )
   }
+  const line = groups.join(' | ')
+  // The row elides with ellipsis when overlong; a delayed hover tooltip carries
+  // the full line, enabled only while content is actually clipped.
+  const rootRef = useRef<HTMLDivElement | null>(null)
+  const [truncated, setTruncated] = useState(false)
+  useLayoutEffect(() => {
+    const el = rootRef.current
+    if (el === null) return
+    const measure = () => { setTruncated(el.scrollWidth > el.clientWidth) }
+    measure()
+    const observer = new ResizeObserver(measure)
+    observer.observe(el)
+    return () => { observer.disconnect() }
+  }, [line])
   if (groups.length === 0) return null
   return (
-    <div className={css.root}>
-      {groups.map((group, i) => (
-        <Fragment key={group}>
-          {i > 0 && <><span className={css.sep} aria-hidden>|</span>{' '}</>}
-          <span>{group}</span>
-        </Fragment>
-      ))}
-    </div>
+    <Tooltip label={line} side="top" delayMs={500} disabled={!truncated}>
+      <div ref={rootRef} className={css.root}>
+        {groups.map((group, i) => (
+          <Fragment key={group}>
+            {i > 0 && <><span className={css.sep} aria-hidden>|</span>{' '}</>}
+            <span>{group}</span>
+          </Fragment>
+        ))}
+      </div>
+    </Tooltip>
   )
 })

+ 21 - 0
packages/client/ui-conversation/src/client/chat/message-chrome.ts

@@ -48,6 +48,27 @@ export function formatRunDuration(ms: number, t: RunDurationTranslate): string {
     : t('duration.seconds', { seconds })
 }
 
+/**
+ * Sub-turn latency figure: one decimal under ten seconds, whole seconds
+ * beyond. Unit-less so the locale template owns the second suffix.
+ * @param ms - Latency in milliseconds (negatives clamp to zero).
+ * @returns Display number in seconds without unit.
+ */
+export function formatLatencySeconds(ms: number): string {
+  const s = Math.max(0, ms) / 1000
+  return s < 10 ? String(Math.round(s * 10) / 10) : String(Math.round(s))
+}
+
+/**
+ * Decode-throughput figure: whole tokens from ten up, one decimal below.
+ * @param tps - Tokens per second.
+ * @returns Display number without unit.
+ */
+export function formatTokensPerSecond(tps: number): string {
+  const clamped = Math.max(0, tps)
+  return clamped >= 10 ? String(Math.round(clamped)) : String(Math.round(clamped * 10) / 10)
+}
+
 /**
  * Compact local timestamp for message IconActions. Same calendar day →
  * `HH:mm`; earlier this year → the `clock.md` date template + clock; other

+ 97 - 0
packages/client/ui-conversation/src/client/chat/turn-metrics.ts

@@ -0,0 +1,97 @@
+// Latency/throughput folds shared by the settled turn footer and StatsLine.
+
+import type { ConversationSnapshot } from '@deepseek-ai/dsh-client-runtime/client'
+
+/** Latency and decode-throughput readings for one turn's footer. */
+export interface TurnMetrics {
+  /** First-step TTFT in ms; absent when that step carries no recorded timing. */
+  ttftMs?: number
+  /** Decode throughput over steps carrying both timing and provider usage. */
+  tokensPerSecond?: number
+}
+
+/** One assistant step's derivable latency facts; null marks an unrecorded part. */
+export interface StepReading {
+  /** step/start → first token delta, in ms. */
+  ttftMs: number | null
+  /** First token delta → final message, in ms. */
+  decodeMs: number | null
+  /** Provider-reported completion tokens. */
+  outputTokens: number | null
+}
+
+interface UsageLike {
+  outputTokens?: number
+}
+
+type AssistantNode = Extract<ConversationSnapshot['nodes'][number], { kind: 'assistant' }>
+
+function usageOutputTokens(usage: unknown): number | null {
+  if (typeof usage !== 'object' || usage === null) return null
+  const value = (usage as UsageLike).outputTokens
+  return typeof value === 'number' && Number.isFinite(value) && value >= 0 ? value : null
+}
+
+/**
+ * Read one assistant node's TTFT, decode wall time, and output tokens.
+ * @param node - A settled assistant node.
+ * @returns Per-part readings with `null` for unrecorded values.
+ */
+export function assistantStepReading(node: AssistantNode): StepReading {
+  const timing = node.timing
+  const ttftMs = timing !== undefined && timing.stepStartTime !== null && timing.firstTokenTime !== null
+    ? Math.max(0, timing.firstTokenTime - timing.stepStartTime)
+    : null
+  const decodeMs = timing !== undefined && timing.firstTokenTime !== null
+    ? Math.max(0, timing.completedTime - timing.firstTokenTime)
+    : null
+  return { ttftMs, decodeMs, outputTokens: usageOutputTokens(node.usage) }
+}
+
+interface TurnFold {
+  firstStep: number
+  firstStepTtftMs: number | null
+  decodeMs: number
+  outputTokens: number
+  sampled: boolean
+}
+
+/**
+ * Fold assistant nodes into per-turn footer metrics.
+ *
+ * TTFT is the turn's lowest-step reading — the user-perceived wait before
+ * output appeared — so it is only meaningful when the turn's start is inside
+ * the loaded window (the caller gates on `turnTimings`, which shares that
+ * window). Throughput divides summed output tokens by summed decode wall time,
+ * counting only steps that carry both.
+ * @param nodes - Snapshot nodes of the loaded window.
+ * @returns Turn number → available metrics; turns with none are absent.
+ */
+export function deriveTurnMetrics(nodes: ConversationSnapshot['nodes']): Map<number, TurnMetrics> {
+  const folds = new Map<number, TurnFold>()
+  for (const node of nodes) {
+    if (node.kind !== 'assistant') continue
+    const reading = assistantStepReading(node)
+    let fold = folds.get(node.turn)
+    if (fold === undefined) {
+      fold = { firstStep: node.step, firstStepTtftMs: reading.ttftMs, decodeMs: 0, outputTokens: 0, sampled: false }
+      folds.set(node.turn, fold)
+    } else if (node.step < fold.firstStep) {
+      fold.firstStep = node.step
+      fold.firstStepTtftMs = reading.ttftMs
+    }
+    if (reading.decodeMs !== null && reading.outputTokens !== null) {
+      fold.decodeMs += reading.decodeMs
+      fold.outputTokens += reading.outputTokens
+      fold.sampled = true
+    }
+  }
+  const metrics = new Map<number, TurnMetrics>()
+  for (const [turn, fold] of folds) {
+    const entry: TurnMetrics = {}
+    if (fold.firstStepTtftMs !== null) entry.ttftMs = fold.firstStepTtftMs
+    if (fold.sampled && fold.decodeMs > 0) entry.tokensPerSecond = fold.outputTokens / (fold.decodeMs / 1000)
+    if (entry.ttftMs !== undefined || entry.tokensPerSecond !== undefined) metrics.set(turn, entry)
+  }
+  return metrics
+}

+ 14 - 0
packages/client/ui-conversation/src/client/locales.ts

@@ -23,6 +23,11 @@ export const zh = {
   'input.stop': '停止生成',
   'input.send': '发送消息',
   'input.accessMode': '访问模式,当前:{name}',
+  'context.aria': '上下文已用 {percent}%',
+  'context.used': '上下文已用',
+  'context.system': '系统提示词',
+  'context.tools': '工具',
+  'context.messages': '对话消息',
   'settings.enter.title': '繁忙时 Enter 键行为',
   'settings.enter.description': '仅在智能体运行时生效;Cmd/Ctrl+Enter 使用另一行为',
   'settings.enter.queue': '排队发送',
@@ -71,6 +76,8 @@ export const zh = {
   'message.retry.failure': '失败原因:',
   'message.turnError': '本轮运行失败',
   'message.ranFor': '用时 {duration}',
+  'message.ttft': '首 token {seconds}秒',
+  'message.tokensPerSecond': '{tps} tok/s',
   'duration.seconds': '{seconds}秒',
   'duration.minutes': '{minutes}分{seconds}秒',
   'command.running': '执行中…',
@@ -136,6 +143,11 @@ export const en = {
   'input.stop': 'Stop generating',
   'input.send': 'Send message',
   'input.accessMode': 'Access mode, current: {name}',
+  'context.aria': '{percent}% of context used',
+  'context.used': 'of context used',
+  'context.system': 'System prompt',
+  'context.tools': 'Tools',
+  'context.messages': 'Messages',
   'settings.enter.title': 'Enter behavior while busy',
   'settings.enter.description': 'Busy only; Cmd/Ctrl+Enter uses the other behavior',
   'settings.enter.queue': 'Queue',
@@ -184,6 +196,8 @@ export const en = {
   'message.retry.failure': 'Failure reason: ',
   'message.turnError': 'This turn failed',
   'message.ranFor': 'Ran for {duration}',
+  'message.ttft': 'TTFT {seconds}s',
+  'message.tokensPerSecond': '{tps} tok/s',
   'duration.seconds': '{seconds}s',
   'duration.minutes': '{minutes}m {seconds}s',
   'command.running': 'Running…',

+ 141 - 0
packages/client/ui-conversation/src/client/skeleton/ContextMeter.module.css

@@ -0,0 +1,141 @@
+/* Context-occupancy ring beside the send button plus its click-open breakdown
+   panel (menu surface: r12, inverted hairline, shadow-lv3). */
+
+.root {
+  position: relative;
+  display: inline-flex;
+}
+
+/* Same 28px circular hit target family as the composer's attach button. */
+.trigger {
+  display: grid;
+  place-items: center;
+  flex: none;
+  width: 28px;
+  height: 28px;
+  border: none;
+  border-radius: 999px;
+  background: transparent;
+  color: var(--dsw-alias-label-secondary);
+  cursor: pointer;
+}
+
+.trigger:hover {
+  background: var(--dsw-alias-interactive-bg-hover);
+}
+
+.track {
+  fill: none;
+  stroke: var(--dsw-alias-border-l3);
+  stroke-width: 2;
+}
+
+.fill {
+  fill: none;
+  stroke: var(--dsw-alias-label-tertiary);
+  stroke-width: 2;
+  stroke-linecap: round;
+}
+
+.panel {
+  position: absolute;
+  bottom: calc(100% + 8px);
+  right: 0;
+  z-index: 100;
+  box-sizing: border-box;
+  width: 264px;
+  padding: 12px;
+  border: 1px solid var(--dsw-alias-border-inverted);
+  border-radius: 12px;
+  background: var(--dsw-specific-menu);
+  box-shadow: var(--dsw-shadow-lv3);
+  font-size: 12px;
+  line-height: 20px;
+  color: var(--dsw-alias-label-secondary);
+  cursor: default;
+}
+
+.header {
+  display: flex;
+  align-items: center;
+  gap: 6px;
+}
+
+.figures {
+  margin-left: auto;
+  font-weight: 500;
+  font-variant-numeric: tabular-nums;
+  color: var(--dsw-alias-label-primary);
+}
+
+.percent {
+  font-weight: 500;
+  color: var(--dsw-alias-label-primary);
+}
+
+.headline {
+  color: var(--dsw-alias-label-tertiary);
+}
+
+.bar {
+  display: flex;
+  gap: 1px;
+  margin: 10px 0 12px;
+  height: 4px;
+  border-radius: 999px;
+  background: var(--dsw-alias-interactive-bg-hover);
+  overflow: hidden;
+}
+
+.segment {
+  flex: none;
+  min-width: 2px;
+  height: 100%;
+  border-radius: 1px;
+  background: var(--meter-tint, var(--dsw-alias-label-tertiary));
+}
+
+.swatch {
+  display: inline-block;
+  margin-right: 6px;
+  width: 8px;
+  height: 8px;
+  border-radius: 2px;
+  background: var(--meter-tint);
+  vertical-align: baseline;
+}
+
+.colorSystem {
+  --meter-tint: var(--dsw-static-neutral-bluish-400);
+}
+
+.colorTools {
+  /* The design platform ships no purple static token; violet-400 literal. */
+  --meter-tint: rgb(167, 139, 250);
+}
+
+.colorMessages {
+  --meter-tint: var(--dsw-static-blue-450);
+}
+
+.rows {
+  margin: 6px 0 0;
+}
+
+.row {
+  display: flex;
+  align-items: center;
+  justify-content: space-between;
+  gap: 12px;
+  padding: 2px 0;
+}
+
+.row dt {
+  color: var(--dsw-alias-label-secondary);
+}
+
+.row dd {
+  margin: 0;
+  font-variant-numeric: tabular-nums;
+  color: var(--dsw-alias-label-primary);
+}

+ 131 - 0
packages/client/ui-conversation/src/client/skeleton/ContextMeter.tsx

@@ -0,0 +1,131 @@
+/** Composer context-occupancy meter: a ring beside the send button fed by the
+ * `contextPressure` projection, with a click-open panel of the heuristic
+ * `contextBreakdown` composition (system prompt, tools, conversation).
+ * Renders nothing until a provider reports both pressure and a route capacity
+ * (same gate as the stats row used). */
+
+import { useEffect, useRef, useState } from 'react'
+import type { UseProjection } from '@deepseek-ai/dsh-client-runtime/client'
+// Type-only: the `contextPressure` / `contextBreakdown` projection key merges.
+import type {} from '@deepseek-ai/dsh-token-meter/client'
+import { Tooltip } from '@deepseek-ai/dsh-client-ui-primitives'
+import type { ComposerBarProps } from '../contract/slots.ts'
+import { contextOccupancy, formatTokens } from '../chat/StatsLine.tsx'
+import css from './ContextMeter.module.css'
+
+/** Ring geometry: 14px viewBox, 2px stroke. */
+const RADIUS = 5.5
+const CIRCUMFERENCE = 2 * Math.PI * RADIUS
+
+/** Panel legend rows, in bar-segment order; each color class carries the shared swatch/segment tint. */
+const ROWS = [
+  { key: 'systemTokens', label: 'context.system', color: css.colorSystem },
+  { key: 'toolsTokens', label: 'context.tools', color: css.colorTools },
+  { key: 'messageTokens', label: 'context.messages', color: css.colorMessages },
+] as const
+
+export interface ContextMeterProps {
+  useProjection: UseProjection
+  /** The owning bar's locale seat, passed down as a plain prop. */
+  t: ComposerBarProps['t']
+}
+
+export function ContextMeter({ useProjection, t }: ContextMeterProps) {
+  const pressure = useProjection('contextPressure')
+  const breakdown = useProjection('contextBreakdown')
+  const [open, setOpen] = useState(false)
+  const rootRef = useRef<HTMLSpanElement | null>(null)
+
+  // Outside click / Escape close, one document listener while open (Menu's pattern).
+  useEffect(() => {
+    if (!open) return
+    const onPointerDown = (e: PointerEvent): void => {
+      if (e.target instanceof Node && rootRef.current?.contains(e.target) === true) return
+      setOpen(false)
+    }
+    const onKeyDown = (e: KeyboardEvent): void => {
+      if (e.key === 'Escape') setOpen(false)
+    }
+    document.addEventListener('pointerdown', onPointerDown)
+    document.addEventListener('keydown', onKeyDown)
+    return () => {
+      document.removeEventListener('pointerdown', onPointerDown)
+      document.removeEventListener('keydown', onKeyDown)
+    }
+  }, [open])
+
+  const context = contextOccupancy(pressure)
+  if (context === null) return null
+  const percent = context.percent
+
+  // The bar's overall length stays the provider-exact percent; the heuristic
+  // breakdown only proportions its colored segments.
+  const breakdownTotal = breakdown === undefined
+    ? 0
+    : breakdown.systemTokens + breakdown.toolsTokens + breakdown.messageTokens
+  const segments = breakdown === undefined || breakdownTotal === 0
+    ? null
+    : ROWS.map(row => ({ key: row.key, color: row.color, share: breakdown[row.key] / breakdownTotal }))
+
+  return (
+    <span ref={rootRef} className={css.root}>
+      <Tooltip label={t('context.aria', { percent })} side="top" delayMs={200} disabled={open}>
+        <button
+          type="button"
+          className={css.trigger}
+          aria-label={t('context.aria', { percent })}
+          aria-haspopup="dialog"
+          aria-expanded={open}
+          onClick={() => { setOpen(!open) }}
+        >
+          <svg viewBox="0 0 14 14" width="14" height="14" aria-hidden>
+            <circle className={css.track} cx="7" cy="7" r={RADIUS} />
+            <circle
+              className={css.fill}
+              cx="7"
+              cy="7"
+              r={RADIUS}
+              strokeDasharray={`${CIRCUMFERENCE * percent / 100} ${CIRCUMFERENCE}`}
+              transform="rotate(-90 7 7)"
+            />
+          </svg>
+        </button>
+      </Tooltip>
+      {open && (
+        <div className={css.panel} role="dialog" aria-label={t('context.used')}>
+          <div className={css.header}>
+            <span className={css.percent}>{`${percent}%`}</span>
+            <span className={css.headline}>{t('context.used')}</span>
+            <span className={css.figures}>
+              {`~${formatTokens(context.pressureTokens)} / ${formatTokens(context.contextWindow)}`}
+            </span>
+          </div>
+          <div className={css.bar}>
+            {segments === null
+              ? <div className={css.segment} style={{ width: `${percent}%` }} />
+              : segments.map(segment => (
+                <div
+                  key={segment.key}
+                  className={`${css.segment} ${segment.color}`}
+                  style={{ width: `${percent * segment.share}%` }}
+                />
+              ))}
+          </div>
+          {breakdown !== undefined && (
+            <dl className={css.rows}>
+              {ROWS.map(row => (
+                <div key={row.key} className={css.row}>
+                  <dt>
+                    <span className={`${css.swatch} ${row.color}`} aria-hidden />
+                    {t(row.label)}
+                  </dt>
+                  <dd>{`~${formatTokens(breakdown[row.key])}`}</dd>
+                </div>
+              ))}
+            </dl>
+          )}
+        </div>
+      )}
+    </span>
+  )
+}

+ 2 - 0
packages/client/ui-conversation/src/client/skeleton/InputBar.tsx

@@ -19,6 +19,7 @@ import type { Translate } from '@deepseek-ai/dsh-client-ui-slots'
 import type { ComposerBarProps } from '../contract/slots.ts'
 import { deriveDecorations } from '../input/decorations.ts'
 import type { DraftDecorations } from '../input/decorations.ts'
+import { ContextMeter } from './ContextMeter.tsx'
 import { PermissionSelect } from './PermissionSelect.tsx'
 import css from './InputBar.module.css'
 
@@ -512,6 +513,7 @@ export function InputBar({
           <div className={css.trailing}>
             {rightItems}
             {renderSlot('conversation.input.model', { locked })}
+            <ContextMeter useProjection={useProjection} t={t} />
             {/* {machineBusy && <span className={css.pending} data-input-pending aria-label="处理中" />} */}
             <Tooltip label={primaryLabel} side="top" delayMs={500}>
               <button

+ 9 - 0
packages/client/ui-conversation/tests/chat-branch-tails.spec.tsx

@@ -18,9 +18,18 @@ import { AssistantMarkdown } from '../src/client/chat/AssistantMarkdown.tsx'
 import { StatsLine, type StatsLineProps } from '../src/client/chat/StatsLine.tsx'
 import { zh } from '../src/client/locales.ts'
 
+/** jsdom has no ResizeObserver; StatsLine watches its row for ellipsis truncation through one. */
+class ResizeObserverStub {
+  observe(): void {}
+  unobserve(): void {}
+  disconnect(): void {}
+}
+
+beforeEach(() => { vi.stubGlobal('ResizeObserver', ResizeObserverStub) })
 afterEach(() => {
   cleanup()
   vi.useRealTimers()
+  vi.unstubAllGlobals()
 })
 
 // Mirrors the real lookup chain (conversation namespace, then common).

+ 85 - 37
packages/client/ui-conversation/tests/chat-stats-bash-sample.spec.tsx

@@ -3,8 +3,8 @@
 // hard acceptance — zero renders during streaming. Bash sample row: ToolRow
 // chrome (Bash · description) without a row click target.
 
-import { afterEach, describe, expect, it, vi } from 'vitest'
-import { act, cleanup, render } from '@testing-library/react'
+import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
+import { act, cleanup, fireEvent, render } from '@testing-library/react'
 import type {
   AssistantMessageNode, ConversationSnapshot, SessionId, SessionListState, ToolResultNode,
 } from '@deepseek-ai/dsh-client-runtime/client'
@@ -12,7 +12,7 @@ import { createSnapshotStore } from '@deepseek-ai/dsh-client-runtime/client'
 import { bindSnapshotSelector } from '@deepseek-ai/dsh-client-web-react'
 import { makeTranslate } from '@deepseek-ai/dsh-client-test-runtime'
 import { zh as commonZh } from '@deepseek-ai/dsh-client-locale/src/locales/zh.ts'
-import { StatsLine, deriveStats, formatDuration, formatTokens, type StatsLineProps } from '../src/client/chat/StatsLine.tsx'
+import { StatsLine, contextOccupancy, deriveStats, formatDuration, formatTokens, type StatsLineProps } from '../src/client/chat/StatsLine.tsx'
 import { BashRow } from '../src/client/toolviews/bash-sample.tsx'
 import { zh } from '../src/client/locales.ts'
 
@@ -21,7 +21,20 @@ type BashRowProps = Parameters<typeof BashRow>[0]
 // Mirrors the real lookup chain (conversation namespace, then common).
 const t: BashRowProps['t'] = makeTranslate(zh, commonZh)
 
-afterEach(cleanup)
+/** jsdom has no ResizeObserver; StatsLine watches its row for ellipsis truncation through one. */
+class ResizeObserverStub {
+  observe(): void {}
+  unobserve(): void {}
+  disconnect(): void {}
+}
+
+beforeEach(() => { vi.stubGlobal('ResizeObserver', ResizeObserverStub) })
+afterEach(() => {
+  cleanup()
+  vi.unstubAllGlobals()
+  vi.restoreAllMocks()
+  vi.useRealTimers()
+})
 
 const SID = 's1' as SessionId
 
@@ -66,8 +79,11 @@ describe('deriveStats', () => {
     expect(stats.turns).toBe(2)
     expect(stats.steps).toBe(3)
     // Window-scoped by design: the paged window is not an accounting source, so
-    // the fold exposes no token fields at all (billing rides the projection).
-    expect(Object.keys(stats).sort()).toEqual(['llmMs', 'steps', 'toolMs', 'turns'])
+    // the fold exposes no billing fields (billing rides the projection);
+    // decodeTokens is a throughput input, not a billed total.
+    expect(Object.keys(stats).sort()).toEqual(
+      ['decodeMs', 'decodeTokens', 'llmMs', 'steps', 'toolMs', 'ttftMs', 'ttftSteps', 'turns'],
+    )
   })
 
   it('ignores tool results with no call time', () => {
@@ -97,6 +113,23 @@ describe('deriveStats', () => {
     expect(stats.llmMs).toBe(2_500)
     expect(stats.toolMs).toBe(3_000)
   })
+
+  it('sums ttft per recorded step and decode throughput inputs per usage-carrying step', () => {
+    const sampled: AssistantMessageNode = {
+      ...assistant(1, 1, { outputTokens: 40 }),
+      timing: { stepStartTime: 1_000, firstTokenTime: 1_800, completedTime: 4_800 },
+    }
+    const ttftOnly: AssistantMessageNode = {
+      ...assistant(2, 1),
+      timing: { stepStartTime: 5_000, firstTokenTime: 5_400, completedTime: 7_400 },
+    }
+    const stats = deriveStats([sampled, ttftOnly, assistant(3, 2)])
+    expect(stats.ttftMs).toBe(1_200)
+    expect(stats.ttftSteps).toBe(2)
+    // The usage-less step contributes no decode share, keeping the ratio honest.
+    expect(stats.decodeMs).toBe(3_000)
+    expect(stats.decodeTokens).toBe(40)
+  })
 })
 
 describe('formatters', () => {
@@ -142,47 +175,62 @@ describe('StatsLine', () => {
     expect(emptyView.container.textContent).toBe('')
   })
 
-  it('keeps durable token and context groups after the visible step window is empty', () => {
-    const { source } = makeSource()
-    const view = render(<StatsLine {...props(source, {
-      tokenUsage: USAGE,
-      contextPressure: { pressureTokens: 32_000, contextWindow: 128_000 },
-    })} />)
-    expect(view.container.textContent)
-      .toBe('Context 25% of 128K| Cache hit 90%| Input 100 tok · Output 5 tok')
+  it('reveals the full line in a delayed hover tooltip only while the row is clipped', () => {
+    vi.useFakeTimers()
+    // jsdom lays nothing out; fake a row narrower than its content.
+    vi.spyOn(Element.prototype, 'scrollWidth', 'get').mockReturnValue(800)
+    vi.spyOn(Element.prototype, 'clientWidth', 'get').mockReturnValue(400)
+    const { source } = makeSource({ nodes: [assistant(1, 1)] })
+    const view = render(<StatsLine {...props(source)} />)
+    fireEvent.mouseEnter(view.container.firstElementChild!)
+    act(() => { vi.advanceTimersByTime(499) })
+    expect(view.container.querySelector('[role="tooltip"]')).toBeNull()
+    act(() => { vi.advanceTimersByTime(1) })
+    expect(view.container.querySelector('[role="tooltip"]')?.textContent)
+      .toBe('1 turns · 1 steps | Cache hit 90% | Input 100 tok · Output 5 tok')
   })
 
-  it('renders context occupancy only when the projection knows a capacity', () => {
+  it('suppresses the tooltip while the row fits without truncation', () => {
+    vi.useFakeTimers()
     const { source } = makeSource({ nodes: [assistant(1, 1)] })
-    const withCapacity = render(<StatsLine {...props(source, {
+    const view = render(<StatsLine {...props(source)} />)
+    fireEvent.mouseEnter(view.container.firstElementChild!)
+    act(() => { vi.advanceTimersByTime(500) })
+    expect(view.container.querySelector('[role="tooltip"]')).toBeNull()
+  })
+
+  it('renders window latency and throughput beside the wall-time group', () => {
+    const timed: AssistantMessageNode = {
+      ...assistant(1, 1, { outputTokens: 60 }),
+      timing: { stepStartTime: 1_000, firstTokenTime: 1_800, completedTime: 4_800 },
+    }
+    const { source } = makeSource({ nodes: [timed] })
+    const view = render(<StatsLine {...props(source)} />)
+    expect(view.container.textContent).toContain('LLM 3.8s| TTFT avg 0.8s · 20 tok/s')
+  })
+
+  it('keeps durable token groups after the visible step window is empty', () => {
+    const { source } = makeSource()
+    const view = render(<StatsLine {...props(source, {
       tokenUsage: USAGE,
       contextPressure: { pressureTokens: 32_000, contextWindow: 128_000 },
     })} />)
-    expect(withCapacity.container.textContent).toContain('Context 25% of 128K')
-    // Pressure without capacity has no denominator: the group drops out.
-    const noCapacity = render(<StatsLine {...props(source, {
-      tokenUsage: USAGE,
-      contextPressure: { pressureTokens: 32_000 },
-    })} />)
-    expect(noCapacity.container.textContent).not.toContain('Context')
-    // Capacity arrives before usage in the log; no provider sample means there
-    // is no numerator yet, rather than a synthetic 0%.
-    const noPressure = render(<StatsLine {...props(source, {
-      tokenUsage: USAGE,
-      contextPressure: { contextWindow: 128_000 },
-    })} />)
-    expect(noPressure.container.textContent).not.toContain('Context')
+    // Context occupancy lives on the composer's ContextMeter ring, not here.
+    expect(view.container.textContent)
+      .toBe('Cache hit 90%| Input 100 tok · Output 5 tok')
   })
 
-  it('clamps occupancy at 100% when pressure exceeds the recorded capacity', () => {
+  it('computes context occupancy only when both pressure and capacity are known', () => {
+    expect(contextOccupancy({ pressureTokens: 32_000, contextWindow: 128_000 }))
+      .toEqual({ percent: 25, pressureTokens: 32_000, contextWindow: 128_000 })
+    // Pressure without capacity has no denominator; capacity without a provider
+    // sample has no numerator yet, rather than a synthetic 0%.
+    expect(contextOccupancy({ pressureTokens: 32_000 })).toBeNull()
+    expect(contextOccupancy({ contextWindow: 128_000 })).toBeNull()
+    expect(contextOccupancy(undefined)).toBeNull()
     // Capacity and pressure are independent last-wins fields, so a model switch
     // can pair a smaller new window with the previous route's larger prompt.
-    const { source } = makeSource({ nodes: [assistant(1, 1)] })
-    const view = render(<StatsLine {...props(source, {
-      tokenUsage: USAGE,
-      contextPressure: { pressureTokens: 300_000, contextWindow: 128_000 },
-    })} />)
-    expect(view.container.textContent).toContain('Context 100% of 128K')
+    expect(contextOccupancy({ pressureTokens: 300_000, contextWindow: 128_000 })?.percent).toBe(100)
   })
 
   it('drops every token group when no projection is composed', () => {

+ 39 - 0
packages/client/ui-conversation/tests/chat-view.spec.tsx

@@ -540,6 +540,45 @@ describe('ChatView', () => {
     expect(view.getAllByText(/用时 19秒/)).toHaveLength(1)
   })
 
+  it('the settled footer appends first-step ttft and turn decode throughput', () => {
+    const first: AssistantMessageNode = {
+      kind: 'assistant', seq: 2, time: 2_000, turn: 1, step: 1, blocks: [{ kind: 'text', text: 'mid' }],
+      timing: { stepStartTime: 1_000, firstTokenTime: 2_200, completedTime: 5_200 },
+      usage: { outputTokens: 40 },
+    }
+    const second: AssistantMessageNode = {
+      kind: 'assistant', seq: 16, time: 16_000, turn: 1, step: 2, blocks: [{ kind: 'text', text: 'final' }],
+      timing: { stepStartTime: 10_000, firstTokenTime: 10_200, completedTime: 12_200 },
+      usage: { outputTokens: 60 },
+    }
+    const h = makeHarness({
+      nodes: [user(1, 'hi'), first, second],
+      turnTimings: new Map([[1, { startTime: 1_000, endTime: 20_000 }]]),
+      turnEnds: new Map([[1, 20]]),
+    })
+    const view = render(<h.ChatView {...h.props} />)
+    // First-step ttft (1.2s) plus 100 tokens over 5s of decode.
+    expect(view.getAllByText(/用时 19秒/)).toHaveLength(1)
+    expect(view.getAllByText(/首 token 1\.2秒/)).toHaveLength(1)
+    expect(view.getAllByText(/20 tok\/s/)).toHaveLength(1)
+  })
+
+  it('withholds ttft and throughput while the turn is still running', () => {
+    const settled: AssistantMessageNode = {
+      kind: 'assistant', seq: 2, time: 2_000, turn: 1, step: 1, blocks: [{ kind: 'text', text: 'answer' }],
+      timing: { stepStartTime: 1_000, firstTokenTime: 1_500, completedTime: 2_000 },
+      usage: { outputTokens: 10 },
+    }
+    const h = makeHarness({
+      nodes: [user(1, 'hi'), settled],
+      turnTimings: new Map([[1, { startTime: 1_000 }]]),
+      turnEnds: new Map(),
+      running: true,
+    })
+    const view = render(<h.ChatView {...h.props} />)
+    expect(view.queryByText(/首 token|tok\/s/)).toBeNull()
+  })
+
   it('user and assistant message containers scope the hover-revealed time chrome', () => {
     const h = makeHarness({
       nodes: [user(1, 'hi'), assistant(2, 'answer')],

+ 93 - 0
packages/client/ui-conversation/tests/context-meter.spec.tsx

@@ -0,0 +1,93 @@
+// @vitest-environment jsdom
+// ContextMeter (composer trailing control): occupancy ring gating, the
+// click-open breakdown panel, and its close gestures.
+
+import { afterEach, describe, expect, it } from 'vitest'
+import { cleanup, fireEvent, render } from '@testing-library/react'
+import { makeTranslate } from '@deepseek-ai/dsh-client-test-runtime'
+import { zh as commonZh } from '@deepseek-ai/dsh-client-locale/src/locales/zh.ts'
+import { ContextMeter, type ContextMeterProps } from '../src/client/skeleton/ContextMeter.tsx'
+import css from '../src/client/skeleton/ContextMeter.module.css'
+import { zh } from '../src/client/locales.ts'
+
+afterEach(cleanup)
+
+// Mirrors the real lookup chain (conversation namespace, then common).
+const t = makeTranslate(zh, commonZh) as ContextMeterProps['t']
+
+const BREAKDOWN = { systemTokens: 120, toolsTokens: 21_500, messageTokens: 477_000 }
+
+const segmentClass = css.segment
+if (segmentClass === undefined) throw new Error('segment class missing from ContextMeter.module.css')
+
+/** Stub the projection seat: a key-addressed table of whole values. */
+function projections(values: Record<string, unknown>): ContextMeterProps['useProjection'] {
+  return (key: string) => values[key]
+}
+
+function meter(values: Record<string, unknown>) {
+  return render(<ContextMeter useProjection={projections(values)} t={t} />)
+}
+
+describe('ContextMeter', () => {
+  it('renders nothing until both pressure and capacity are known', () => {
+    expect(meter({}).container.textContent).toBe('')
+    expect(meter({ contextPressure: { pressureTokens: 32_000 } }).container.textContent).toBe('')
+    expect(meter({ contextPressure: { contextWindow: 128_000 } }).container.textContent).toBe('')
+  })
+
+  it('shows the occupancy ring and opens the breakdown panel on click', () => {
+    const view = meter({
+      contextPressure: { pressureTokens: 32_000, contextWindow: 128_000 },
+      contextBreakdown: BREAKDOWN,
+    })
+    const trigger = view.getByRole('button', { name: '上下文已用 25%' })
+    expect(view.container.querySelector('[role="dialog"]')).toBeNull()
+    fireEvent.click(trigger)
+    const panel = view.container.querySelector('[role="dialog"]')!
+    expect(panel.textContent).toContain('~32K / 128K')
+    expect(panel.textContent).toContain('25%')
+    expect(panel.textContent).toContain('上下文已用')
+    expect(panel.textContent).toContain('系统提示词~120')
+    expect(panel.textContent).toContain('工具~21.5K')
+    expect(panel.textContent).toContain('对话消息~477K')
+    // The occupancy bar splits into one colored segment per composition row.
+    expect(panel.getElementsByClassName(segmentClass)).toHaveLength(3)
+    // Clicking the trigger again toggles the panel shut.
+    fireEvent.click(trigger)
+    expect(view.container.querySelector('[role="dialog"]')).toBeNull()
+  })
+
+  it('omits the composition rows while the contextBreakdown projection is absent', () => {
+    const view = meter({ contextPressure: { pressureTokens: 32_000, contextWindow: 128_000 } })
+    fireEvent.click(view.getByRole('button', { name: '上下文已用 25%' }))
+    const panel = view.container.querySelector('[role="dialog"]')!
+    expect(panel.textContent).toContain('~32K / 128K')
+    expect(panel.textContent).not.toContain('系统提示词')
+    expect(panel.textContent).not.toContain('对话消息')
+    // Without composition shares, the bar falls back to one plain segment.
+    expect(panel.getElementsByClassName(segmentClass)).toHaveLength(1)
+  })
+
+  it('closes on outside pointerdown and Escape — but not inside clicks', () => {
+    const view = meter({
+      contextPressure: { pressureTokens: 32_000, contextWindow: 128_000 },
+      contextBreakdown: BREAKDOWN,
+    })
+    const trigger = view.getByRole('button', { name: '上下文已用 25%' })
+    const openPanel = () => {
+      fireEvent.click(trigger)
+      return view.container.querySelector('[role="dialog"]')!
+    }
+    // A pointerdown inside the panel keeps it open; outside closes it.
+    const again = openPanel()
+    fireEvent.pointerDown(again)
+    expect(view.container.querySelector('[role="dialog"]')).not.toBeNull()
+    fireEvent.pointerDown(document.body)
+    expect(view.container.querySelector('[role="dialog"]')).toBeNull()
+    // Escape.
+    openPanel()
+    fireEvent.keyDown(document, { key: 'Escape' })
+    expect(view.container.querySelector('[role="dialog"]')).toBeNull()
+  })
+})

+ 13 - 2
packages/client/ui-conversation/tests/gate-branch-tails.spec.tsx

@@ -1,6 +1,6 @@
 // @vitest-environment jsdom
 
-import { afterEach, describe, expect, it, vi } from 'vitest'
+import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
 import { cleanup, render } from '@testing-library/react'
 import { bindSnapshotSelector } from '@deepseek-ai/dsh-client-web-react'
 import { createSnapshotStore } from '@deepseek-ai/dsh-client-runtime/client'
@@ -18,7 +18,18 @@ import { zh } from '../src/client/locales.ts'
 // Mirrors the real lookup chain (conversation namespace, then common).
 const t: AssistantMarkdownProps['t'] = makeTranslate(zh, commonZh)
 
-afterEach(cleanup)
+/** jsdom has no ResizeObserver; StatsLine watches its row for ellipsis truncation through one. */
+class ResizeObserverStub {
+  observe(): void {}
+  unobserve(): void {}
+  disconnect(): void {}
+}
+
+beforeEach(() => { vi.stubGlobal('ResizeObserver', ResizeObserverStub) })
+afterEach(() => {
+  cleanup()
+  vi.unstubAllGlobals()
+})
 
 const SID = 's1' as SessionId
 

+ 154 - 0
packages/client/ui-conversation/tests/turn-metrics.spec.ts

@@ -0,0 +1,154 @@
+// Per-turn latency/throughput fold and the footer figure formatters.
+
+import { describe, expect, it } from 'vitest'
+import type { AssistantMessageNode, ConversationNode, UserMessageNode } from '@deepseek-ai/dsh-client-runtime/client'
+import { assistantStepReading, deriveTurnMetrics } from '../src/client/chat/turn-metrics.ts'
+import { formatLatencySeconds, formatTokensPerSecond } from '../src/client/chat/message-chrome.ts'
+
+interface StepSpec {
+  seq: number
+  turn: number
+  step: number
+  timing?: AssistantMessageNode['timing']
+  usage?: unknown
+}
+
+const assistant = ({ seq, turn, step, timing, usage }: StepSpec): AssistantMessageNode => ({
+  kind: 'assistant', seq, time: seq * 1_000, turn, step, blocks: [{ kind: 'text', text: `t${seq}` }],
+  ...(timing === undefined ? {} : { timing }),
+  ...(usage === undefined ? {} : { usage }),
+})
+
+const user = (seq: number): UserMessageNode => ({
+  kind: 'user', seq, time: seq * 1_000, content: [{ type: 'text', text: 'hi' }] as never, source: null,
+})
+
+describe('assistantStepReading', () => {
+  it('derives ttft, decode time, and output tokens from a fully recorded step', () => {
+    const reading = assistantStepReading(assistant({
+      seq: 2, turn: 1, step: 1,
+      timing: { stepStartTime: 1_000, firstTokenTime: 1_800, completedTime: 6_800 },
+      usage: { outputTokens: 200 },
+    }))
+    expect(reading).toEqual({ ttftMs: 800, decodeMs: 5_000, outputTokens: 200 })
+  })
+
+  it('returns nulls when timing is absent', () => {
+    const reading = assistantStepReading(assistant({ seq: 2, turn: 1, step: 1, usage: { outputTokens: 5 } }))
+    expect(reading).toEqual({ ttftMs: null, decodeMs: null, outputTokens: 5 })
+  })
+
+  it('needs both boundaries for ttft and clamps negative spans to zero', () => {
+    expect(assistantStepReading(assistant({
+      seq: 2, turn: 1, step: 1,
+      timing: { stepStartTime: null, firstTokenTime: 1_800, completedTime: 6_800 },
+    }))).toEqual({ ttftMs: null, decodeMs: 5_000, outputTokens: null })
+    expect(assistantStepReading(assistant({
+      seq: 2, turn: 1, step: 1,
+      timing: { stepStartTime: 1_000, firstTokenTime: null, completedTime: 6_800 },
+    }))).toEqual({ ttftMs: null, decodeMs: null, outputTokens: null })
+    expect(assistantStepReading(assistant({
+      seq: 2, turn: 1, step: 1,
+      timing: { stepStartTime: 2_000, firstTokenTime: 1_500, completedTime: 1_200 },
+    }))).toEqual({ ttftMs: 0, decodeMs: 0, outputTokens: null })
+  })
+
+  it('rejects non-object, missing, and non-finite usage token counts', () => {
+    const timing = { stepStartTime: 1_000, firstTokenTime: 1_500, completedTime: 2_000 }
+    expect(assistantStepReading(assistant({ seq: 2, turn: 1, step: 1, timing, usage: 'weird' })).outputTokens).toBeNull()
+    expect(assistantStepReading(assistant({ seq: 2, turn: 1, step: 1, timing, usage: {} })).outputTokens).toBeNull()
+    const nan = assistant({ seq: 2, turn: 1, step: 1, timing, usage: { outputTokens: Number.NaN } })
+    expect(assistantStepReading(nan).outputTokens).toBeNull()
+    expect(assistantStepReading(assistant({ seq: 2, turn: 1, step: 1, timing, usage: { outputTokens: -3 } })).outputTokens).toBeNull()
+  })
+})
+
+describe('deriveTurnMetrics', () => {
+  it('takes ttft from the lowest step and throughput over all sampled steps', () => {
+    const nodes: ConversationNode[] = [
+      user(1),
+      // Out of step order on purpose: the lowest step owns the ttft slot.
+      assistant({
+        seq: 4, turn: 1, step: 2,
+        timing: { stepStartTime: 10_000, firstTokenTime: 10_200, completedTime: 12_200 },
+        usage: { outputTokens: 60 },
+      }),
+      assistant({
+        seq: 2, turn: 1, step: 1,
+        timing: { stepStartTime: 1_000, firstTokenTime: 2_200, completedTime: 5_200 },
+        usage: { outputTokens: 40 },
+      }),
+    ]
+    // 100 tokens over 5s of decode.
+    expect(deriveTurnMetrics(nodes).get(1)).toEqual({ ttftMs: 1_200, tokensPerSecond: 20 })
+  })
+
+  it('emits ttft without throughput when no step carries usage', () => {
+    const nodes = [assistant({
+      seq: 2, turn: 1, step: 1,
+      timing: { stepStartTime: 1_000, firstTokenTime: 1_900, completedTime: 3_000 },
+    })]
+    expect(deriveTurnMetrics(nodes).get(1)).toEqual({ ttftMs: 900 })
+  })
+
+  it('emits throughput without ttft when only a later step is recorded', () => {
+    const nodes = [
+      assistant({ seq: 2, turn: 1, step: 1 }),
+      assistant({
+        seq: 4, turn: 1, step: 2,
+        timing: { stepStartTime: 10_000, firstTokenTime: 10_500, completedTime: 12_500 },
+        usage: { outputTokens: 30 },
+      }),
+    ]
+    expect(deriveTurnMetrics(nodes).get(1)).toEqual({ tokensPerSecond: 15 })
+  })
+
+  it('omits turns with no readings and zero-decode throughput', () => {
+    const nodes = [
+      assistant({ seq: 2, turn: 1, step: 1 }),
+      assistant({
+        seq: 4, turn: 2, step: 1,
+        timing: { stepStartTime: null, firstTokenTime: 5_000, completedTime: 5_000 },
+        usage: { outputTokens: 10 },
+      }),
+    ]
+    expect(deriveTurnMetrics(nodes).size).toBe(0)
+  })
+
+  it('keeps turns independent and ignores non-assistant nodes', () => {
+    const nodes: ConversationNode[] = [
+      user(1),
+      assistant({
+        seq: 2, turn: 1, step: 1,
+        timing: { stepStartTime: 1_000, firstTokenTime: 1_400, completedTime: 2_400 },
+        usage: { outputTokens: 10 },
+      }),
+      user(3),
+      assistant({
+        seq: 4, turn: 2, step: 1,
+        timing: { stepStartTime: 4_000, firstTokenTime: 4_100, completedTime: 6_100 },
+        usage: { outputTokens: 100 },
+      }),
+    ]
+    const metrics = deriveTurnMetrics(nodes)
+    expect(metrics.get(1)).toEqual({ ttftMs: 400, tokensPerSecond: 10 })
+    expect(metrics.get(2)).toEqual({ ttftMs: 100, tokensPerSecond: 50 })
+  })
+})
+
+describe('footer figure formatters', () => {
+  it('formats latency with one decimal under ten seconds and whole seconds beyond', () => {
+    expect(formatLatencySeconds(840)).toBe('0.8')
+    expect(formatLatencySeconds(1_000)).toBe('1')
+    expect(formatLatencySeconds(9_949)).toBe('9.9')
+    expect(formatLatencySeconds(12_400)).toBe('12')
+    expect(formatLatencySeconds(-5)).toBe('0')
+  })
+
+  it('formats throughput with whole tokens from ten up and one decimal below', () => {
+    expect(formatTokensPerSecond(34.4)).toBe('34')
+    expect(formatTokensPerSecond(9.96)).toBe('10')
+    expect(formatTokensPerSecond(3.14)).toBe('3.1')
+    expect(formatTokensPerSecond(-1)).toBe('0')
+  })
+})

+ 1 - 1
packages/cordis/tool-cordis/src/api-catalog.ts

@@ -1032,7 +1032,7 @@ export const SERVICE_API: readonly ServiceApiEntry[] = [
       },
       {
         signature: 'estimateMessage(message: Message): number',
-        jsDoc: '/**\n * Heuristically price one model-visible message.\n * @param message - message to price without mutation.\n * @returns content and role-framing tokens under the fixed service heuristic.\n */',
+        jsDoc: '/**\n * Heuristically price one model-visible message (instance face of the pure\n * {@link estimateMessage}).\n * @param message - message to price without mutation.\n * @returns content and role-framing tokens under the fixed service heuristic.\n */',
       },
     ],
   },

+ 5 - 46
packages/core/session/src/index.ts

@@ -15,7 +15,7 @@ import type { Message } from '@deepseek-ai/dsh-llm'
 import { SESSION_FORMAT_VERSION, SessionId } from './types.ts'
 import type { CreateSessionOptions, EpochHeader, RequestContext, SessionEvent, SessionEventMap, SessionEventType, SessionHeader, SurfaceIntent, SurfaceEventType } from './types.ts'
 import { snapshotJsonValue } from './json.ts'
-import { SurfaceManager } from './surface.ts'
+import { deriveEventMessage, SurfaceManager } from './surface.ts'
 import type { SessionSurface } from './surface.ts'
 import { foldRequestHeader } from './request-header.ts'
 
@@ -27,7 +27,7 @@ export { interruptedTurnClosers, lastActivityTime, TOOL_NOT_STARTED, TOOL_OUTCOM
 export { decodeStorageRecord, packChunkRuns } from './chunk-rows.ts'
 export type { ChunkRow, StorageRecord } from './chunk-rows.ts'
 export type { SessionSurface, SurfaceFoldReplacement, SurfaceFoldResult } from './surface.ts'
-export { foldSurface, isAppendSurfaceEvent, isReplacementSurfaceEvent, isSurfaceEvent, isSurfaceEligibleType } from './surface.ts'
+export { deriveEventMessage, foldSurface, isAppendSurfaceEvent, isReplacementSurfaceEvent, isSurfaceEvent, isSurfaceEligibleType } from './surface.ts'
 export { canonicalHeader, foldRequestHeader, headerEquals } from './request-header.ts'
 
 /**
@@ -681,54 +681,13 @@ export class Session {
   }
 
   /**
-   * Project a single event into the LLM message it derives to, or null when
-   * it produces none — a non-surface event (chunk, boundary, log-only record)
-   * or an empty-content assistant/message (which exists only to host usage).
-   * The per-node pure function {@link deriveMessages} folds over the surface;
-   * an external reconstructor (or the dev invariant) folds the same function
-   * over a log prefix's surface to rebuild the exact messages any request was
-   * built from (the reconstructability Agent Note). The returned message is
-   * the already frozen message nested in the event wrapper and shared by
-   * delivery, durable history, and model requests.
+   * Instance face of the pure per-node projection rule
+   * {@link deriveEventMessage} (see its contract in `surface.ts`).
    * @param event - the event to project.
    * @returns the derived message, or null when the event produces none.
    */
   deriveEventMessage(event: SessionEvent): Message | null {
-    // Intentionally non-exhaustive: only message-producing events derive
-    // history; turn/step boundaries, chunks, usage, and errors are
-    // trace/replay data.
-
-    switch (event.type) {
-      // Ordinary prompts, injected context, and mid-turn steering project
-      // identically in user role: the event's model-facing content stays
-      // verbatim. Steering's `turn` is log-only. Do NOT
-      // re-add per-type framing (e.g. `<context>`/`<steering>`) here: framing is
-      // caller-owned — a producer bakes it into `content`, as workspace-context
-      // does with `<system-reminder>` — or, if reintroduced, must be driven by
-      // the event `meta` map and a dedicated renderer, keeping this projection a
-      // verbatim pass-through. See the deferred design note in
-      // ../../../../.agents/notes/implemented/simplification/2026-07-20-unwrap-injected-content-envelopes.md
-      case 'user/message': {
-        return event.data
-      }
-      case 'steering/message': {
-        return event.data.message
-      }
-      case 'assistant/message': {
-        // Skip an empty-content assistant/message: it exists only to host a
-        // max-tokens step's usage and must not inject a content-less assistant
-        // turn into the provider transcript.
-        if (event.data.message.content.length === 0) return null
-        return event.data.message
-      }
-      case 'tool/result': {
-        return event.data.message
-      }
-      default:
-        // A non-surface event (boundary, chunk, log-only record) projects to
-        // no message. Merge-extensible union: no assertNever here.
-        return null
-    }
+    return deriveEventMessage(event)
   }
 }
 

+ 51 - 0
packages/core/session/src/surface.ts

@@ -8,6 +8,7 @@
  * @module @deepseek-ai/dsh-session/surface
  */
 
+import type { Message } from '@deepseek-ai/dsh-llm'
 import type { SessionEvent, SurfaceEvent, SurfaceEventType, SurfaceOp } from './types.ts'
 
 /** Runtime counterpart of the message-producing event union. */
@@ -67,6 +68,56 @@ export function isReplacementSurfaceEvent(
   return isSurfaceEvent(event) && event.surfaceOp !== 'append'
 }
 
+/**
+ * Project a single event into the LLM message it derives to, or null when it
+ * produces none — a non-surface event (chunk, boundary, log-only record) or an
+ * empty-content assistant/message (which exists only to host usage). This is
+ * THE per-node projection rule: `Session.deriveMessages` folds it over the
+ * live surface, external reconstructors and pure projections fold the same
+ * function over a log prefix's surface to rebuild the exact messages any
+ * request was built from. The returned message is the already frozen message
+ * nested in the event wrapper and shared by delivery, durable history, and
+ * model requests.
+ * @param event - the event to project.
+ * @returns the derived message, or null when the event produces none.
+ */
+export function deriveEventMessage(event: SessionEvent): Message | null {
+  // Intentionally non-exhaustive: only message-producing events derive
+  // history; turn/step boundaries, chunks, usage, and errors are trace/replay
+  // data.
+  switch (event.type) {
+    // Ordinary prompts, injected context, and mid-turn steering project
+    // identically in user role: the event's model-facing content stays
+    // verbatim. Steering's `turn` is log-only. Do NOT re-add per-type framing
+    // (e.g. `<context>`/`<steering>`) here: framing is caller-owned — a
+    // producer bakes it into `content`, as workspace-context does with
+    // `<system-reminder>` — or, if reintroduced, must be driven by the event
+    // `meta` map and a dedicated renderer, keeping this projection a verbatim
+    // pass-through. See the deferred design note in
+    // ../../../../.agents/notes/implemented/simplification/2026-07-20-unwrap-injected-content-envelopes.md
+    case 'user/message': {
+      return event.data
+    }
+    case 'steering/message': {
+      return event.data.message
+    }
+    case 'assistant/message': {
+      // Skip an empty-content assistant/message: it exists only to host a
+      // max-tokens step's usage and must not inject a content-less assistant
+      // turn into the provider transcript.
+      if (event.data.message.content.length === 0) return null
+      return event.data.message
+    }
+    case 'tool/result': {
+      return event.data.message
+    }
+    default:
+      // A non-surface event (boundary, chunk, log-only record) projects to
+      // no message. Merge-extensible union: no assertNever here.
+      return null
+  }
+}
+
 /** One replacement operation observed while folding a session surface. */
 export interface SurfaceFoldReplacement {
   /** Seq of the event that replaced the prior surface range. */

+ 2 - 2
packages/llm/token-meter/README.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write packages/llm/token-meter/README.md
-README.md: 701893b342f9a93a75bec175634b1054f3d17151
-README.zh.md: a5844e8788422bba669632ed587fb87e1e2a1e58
+README.md: 2b8320221dc314c9a18f93eab3d88a66d22d8342
+README.zh.md: 7c0a433146013e79034d5eb34ebb39e93caf30ff

+ 4 - 2
packages/llm/token-meter/README.md

@@ -23,13 +23,15 @@ Usage accounting sums disjoint input, cache-read, cache-write, and output bucket
 
 ## Session projections
 
-When the composition provides `ctx.sessionProjections`, token-meter registers two units through an optional child fiber.
+When the composition provides `ctx.sessionProjections`, token-meter registers three units through an optional child fiber.
 
 `tokenUsage` carries the complete durable log's `uncachedInputTokens`, `outputTokens`, `cacheReadTokens`, and `cacheWriteTokens`. Usage chunks are counted even when a request later fails; a final assistant-message usage for the same `(turn, step)` replaces that sample instead of double-counting it. Reasoning remains an output subdivision. The single last-sample slot relies on a session-log ordering property: once a later step reports usage, a legal log never reports usage for an earlier step again.
 
 `contextPressure` carries optional `pressureTokens` — the newest provider-reported prompt size, summing uncached input plus cache reads and writes — and optional `contextWindow` from the newest `request/context` record. Pressure stays absent until a provider reports usage; capacity stays absent for a route whose adapter advertises none. Output is excluded, so the numerator holds still while a turn streams and steps forward when the next request reports its usage.
 
-Both units use the standard projection baseline, live frame, higher-seq-wins store, and JSON checkpoint paths. Unloading token-meter removes both keys. A headless or TUI composition without the projection seam keeps the measurement service's existing behavior.
+`contextBreakdown` carries heuristic `systemTokens`, `toolsTokens`, and `messageTokens` — the context's composition rather than its provider-billed size. The envelope figures reprice last-wins on every `request/header`; the message figure folds surface appends and positional replacements, so compaction shrinks it the same way it shrinks the next request. All three figures use the measurement service's fixed heuristic and are estimates: they do not reconcile with the provider-exact `pressureTokens`, and a UI should present them as approximations.
+
+All three units use the standard projection baseline, live frame, higher-seq-wins store, and JSON checkpoint paths. Unloading token-meter removes all three keys. A headless or TUI composition without the projection seam keeps the measurement service's existing behavior.
 
 ### Context occupancy is an approximation, by design
 

+ 4 - 2
packages/llm/token-meter/README.zh.md

@@ -23,13 +23,15 @@ fold 跟踪完整请求标头快照、步骤边界、表层追加与替换、成
 
 ## 会话投影
 
-当组合提供 `ctx.sessionProjections` 时,token-meter 会通过一个可选子 fiber 注册两个单元。
+当组合提供 `ctx.sessionProjections` 时,token-meter 会通过一个可选子 fiber 注册三个单元。
 
 `tokenUsage` 携带完整持久日志中的 `uncachedInputTokens`、`outputTokens`、`cacheReadTokens` 和 `cacheWriteTokens`。即使请求随后失败,用量分片仍会计入;同一 `(turn, step)` 的最终 assistant 消息用量会替换该样本,而不是重复计数。推理仍是输出的一个细分项。只保留单个最新样本,依赖的是会话日志的一条顺序性质:一旦某个更晚的步骤报告了用量,合法日志就绝不会再为更早的步骤报告用量。
 
 `contextPressure` 携带可选的 `pressureTokens`(提供方报告的最新提示词规模,为未缓存输入加缓存读取与写入之和),以及来自最新一条 `request/context` 记录的可选 `contextWindow`。提供方报告用量前压力保持缺失;路由适配器未公布容量时容量也保持缺失。输出不计入其中,因此轮次流式输出期间分子保持不动,等到下一个请求报告用量时才前进。
 
-两个单元都使用标准的投影基线、实时帧、seq 高者胜值仓和 JSON 检查点路径。卸载 token-meter 会移除这两个键。不带投影 seam 的 headless 或 TUI 组合会保留测量服务的既有行为。
+`contextBreakdown` 携带启发式的 `systemTokens`、`toolsTokens` 与 `messageTokens`,描述上下文的组成而非提供方计费规模。envelope 数字在每条 `request/header` 上按后者胜重新计价;消息数字折叠表层追加与位置替换,因此压缩会像缩小下一个请求那样缩小它。三个数字都使用测量服务的固定启发式规则,属于估算值:它们不会与提供方精确的 `pressureTokens` 对账,UI 应以近似值方式呈现。
+
+三个单元都使用标准的投影基线、实时帧、seq 高者胜值仓和 JSON 检查点路径。卸载 token-meter 会移除这三个键。不带投影 seam 的 headless 或 TUI 组合会保留测量服务的既有行为。
 
 ### 上下文占用率是刻意为之的近似值
 

+ 87 - 0
packages/llm/token-meter/src/breakdown-projection.ts

@@ -0,0 +1,87 @@
+/**
+ * Pure fold for the heuristic context-composition projection: system prompt
+ * and tool schemas from the newest request envelope, conversation from the
+ * live surface. Prices with the same shared estimator as the meter service,
+ * so the three figures match `measure()`'s heuristic vocabulary exactly.
+ */
+
+import { z } from 'zod'
+import { canonicalHeader, deriveEventMessage, isSurfaceEvent } from '@deepseek-ai/dsh-session'
+import type { ProjectionDefinition } from '@deepseek-ai/dsh-session-projection'
+import { estimateMessage, estimateSystemTokens, estimateToolsTokens } from './estimate.ts'
+// Import for the `contextBreakdown` SessionProjectionMap key merge.
+import type {} from './projection.ts'
+
+/** One priced surface node (plain JSON for the persisted projection cache). */
+interface BreakdownSurfaceNode {
+  seq: number
+  tokens: number
+}
+
+interface ContextBreakdownState {
+  systemTokens: number
+  toolsTokens: number
+  messageTokens: number
+  surface: BreakdownSurfaceNode[]
+}
+
+const breakdownSchema = z.object({
+  systemTokens: z.number().int().nonnegative(),
+  toolsTokens: z.number().int().nonnegative(),
+  messageTokens: z.number().int().nonnegative(),
+}).strict()
+
+/**
+ * Token-meter's context-composition projection unit.
+ *
+ * Envelope figures are last-wins per `request/header`; the message figure
+ * folds surface appends and positional replacements, so compaction shrinks it
+ * the same way it shrinks the next request. Committed logs are
+ * surface-validated at append time, so an unresolvable replace range here is
+ * log corruption and fails loud rather than skipping the event.
+ */
+export const contextBreakdownProjectionDefinition:
+ProjectionDefinition<'contextBreakdown', ContextBreakdownState> = {
+  key: 'contextBreakdown',
+  schema: breakdownSchema,
+  init: () => ({ systemTokens: 0, toolsTokens: 0, messageTokens: 0, surface: [] }),
+  apply: (state, event) => {
+    if (event.type === 'request/header') {
+      const header = canonicalHeader(event.data.header)
+      const systemTokens = estimateSystemTokens(header)
+      const toolsTokens = estimateToolsTokens(header)
+      if (systemTokens === state.systemTokens && toolsTokens === state.toolsTokens) return state
+      return { ...state, systemTokens, toolsTokens }
+    }
+    if (!isSurfaceEvent(event)) return state
+    const message = deriveEventMessage(event)
+    const tokens = message === null ? 0 : estimateMessage(message)
+    const op = event.surfaceOp
+    if (op === 'append') {
+      return {
+        ...state,
+        messageTokens: state.messageTokens + tokens,
+        surface: [...state.surface, { seq: event.seq, tokens }],
+      }
+    }
+    const startIdx = state.surface.findIndex(node => node.seq === op.start)
+    const endIdx = state.surface.findIndex(node => node.seq === op.end)
+    if (startIdx === -1 || endIdx === -1 || startIdx > endIdx) {
+      throw new Error(
+        `context breakdown: replace at seq ${event.seq} has invalid current range ${op.start}-${op.end}`,
+      )
+    }
+    const removed = state.surface
+      .slice(startIdx, endIdx + 1)
+      .reduce((total, node) => total + node.tokens, 0)
+    const surface = [...state.surface]
+    surface.splice(startIdx, endIdx - startIdx + 1, { seq: event.seq, tokens })
+    return {
+      ...state,
+      messageTokens: state.messageTokens + tokens - removed,
+      surface,
+    }
+  },
+  view: ({ systemTokens, toolsTokens, messageTokens }) => ({ systemTokens, toolsTokens, messageTokens }),
+  stateVersion: 1,
+}

+ 87 - 0
packages/llm/token-meter/src/estimate.ts

@@ -0,0 +1,87 @@
+/**
+ * Fixed-density heuristic token pricing shared by the meter service and the
+ * pure context-breakdown projection, so both surfaces price identical content
+ * to identical numbers.
+ *
+ * @module @deepseek-ai/dsh-token-meter/estimate
+ */
+
+import type { ContentBlock, Message } from '@deepseek-ai/dsh-llm'
+import type { EpochHeader } from '@deepseek-ai/dsh-session'
+
+/** Fixed text-density estimate used until exact tokenization is needed. */
+const CHARS_PER_TOKEN = 4
+
+/** Per-block structural overhead for JSON framing and type tags. */
+const BLOCK_OVERHEAD = 4
+
+/** Role-field framing overhead added to every priced message. */
+export const ROLE_OVERHEAD = 4
+
+/**
+ * Price content blocks recursively under the fixed density heuristic.
+ * @param blocks - content blocks to price without mutation.
+ * @returns heuristic tokens including per-block structural overhead.
+ */
+export function estimateContent(blocks: readonly ContentBlock[]): number {
+  let tokens = 0
+  for (const block of blocks) {
+    switch (block.type) {
+      case 'text':
+      case 'reasoning':
+        tokens += Math.ceil(block.text.length / CHARS_PER_TOKEN) + BLOCK_OVERHEAD
+        break
+      case 'tool-call':
+        tokens += Math.ceil(block.name.length / CHARS_PER_TOKEN)
+          + Math.ceil(block.arguments.length / CHARS_PER_TOKEN)
+          + BLOCK_OVERHEAD
+        break
+      case 'tool-result':
+        tokens += estimateContent(block.content) + BLOCK_OVERHEAD
+        break
+      default:
+        // ContentBlockMap is merge-extensible; unknown blocks retain a
+        // conservative structural JSON price under the fixed heuristic.
+        tokens += BLOCK_OVERHEAD + Math.ceil(JSON.stringify(block).length / CHARS_PER_TOKEN)
+    }
+  }
+  return tokens
+}
+
+/**
+ * Heuristically price one model-visible message.
+ * @param message - message to price without mutation.
+ * @returns content and role-framing tokens under the fixed heuristic.
+ */
+export function estimateMessage(message: Message): number {
+  return estimateContent(message.content) + ROLE_OVERHEAD
+}
+
+/**
+ * Price the system-prompt part of a canonical request envelope.
+ * @param header - canonical envelope, or undefined before any request.
+ * @returns heuristic system-prompt tokens; 0 when absent.
+ */
+export function estimateSystemTokens(header: EpochHeader | undefined): number {
+  if (header?.system === undefined) return 0
+  return Math.ceil(header.system.length / CHARS_PER_TOKEN) + ROLE_OVERHEAD
+}
+
+/**
+ * Price the tool-schema part of a canonical request envelope.
+ * @param header - canonical envelope, or undefined before any request.
+ * @returns heuristic tool-schema tokens; 0 when absent or empty.
+ */
+export function estimateToolsTokens(header: EpochHeader | undefined): number {
+  if (header?.tools === undefined || header.tools.length === 0) return 0
+  return Math.ceil(JSON.stringify(header.tools).length / CHARS_PER_TOKEN) + BLOCK_OVERHEAD
+}
+
+/**
+ * Price the complete non-surface request envelope.
+ * @param header - canonical envelope, or undefined before any request.
+ * @returns heuristic system plus tool tokens.
+ */
+export function estimateHeader(header: EpochHeader | undefined): number {
+  return estimateSystemTokens(header) + estimateToolsTokens(header)
+}

+ 11 - 55
packages/llm/token-meter/src/index.ts

@@ -7,7 +7,7 @@
 import { Context, Service } from 'cordis'
 import z from 'schemastery'
 import { BlockAssembler, deepFreeze } from '@deepseek-ai/dsh-llm'
-import type { ContentBlock, Message, TokenUsage } from '@deepseek-ai/dsh-llm'
+import type { Message, TokenUsage } from '@deepseek-ai/dsh-llm'
 import type { EpochHeader, Session, SessionEvent, SurfaceEvent } from '@deepseek-ai/dsh-session'
 import { canonicalHeader, headerEquals, isSurfaceEvent } from '@deepseek-ai/dsh-session'
 // Type-only: resolves the optional projection registry Context seam.
@@ -18,19 +18,12 @@ import type {
   TokenMeterConfig,
   TokenSurfaceNode,
 } from './types.ts'
+import { contextBreakdownProjectionDefinition } from './breakdown-projection.ts'
 import { contextPressureProjectionDefinition, tokenUsageProjectionDefinition } from './usage-projection.ts'
+import { estimateContent, estimateHeader, estimateMessage, ROLE_OVERHEAD } from './estimate.ts'
 
 export type * from './types.ts'
 
-/** Fixed text-density estimate used until exact tokenization is needed. */
-const CHARS_PER_TOKEN = 4
-
-/** Per-block structural overhead for JSON framing and type tags. */
-const BLOCK_OVERHEAD = 4
-
-/** Role-field framing overhead added to every priced message. */
-const ROLE_OVERHEAD = 4
-
 interface MeasurementAnchor {
   readonly header: EpochHeader | undefined
   readonly surfaceTokens: number
@@ -98,6 +91,7 @@ export class TokenMeterService extends Service {
     ctx.inject(['sessionProjections'], (projectionCtx) => {
       projectionCtx.sessionProjections.register(tokenUsageProjectionDefinition)
       projectionCtx.sessionProjections.register(contextPressureProjectionDefinition)
+      projectionCtx.sessionProjections.register(contextBreakdownProjectionDefinition)
     })
 
     // Readers catch up independently, while eager observation bounds ordinary
@@ -141,7 +135,7 @@ export class TokenMeterService extends Service {
     } else {
       baseline = {
         kind: 'estimated',
-        tokens: this._estimateHeader(header) + state.surfaceTokens,
+        tokens: estimateHeader(header) + state.surfaceTokens,
       }
       surfaceDeltaTokens = 0
     }
@@ -157,12 +151,13 @@ export class TokenMeterService extends Service {
   }
 
   /**
-   * Heuristically price one model-visible message.
+   * Heuristically price one model-visible message (instance face of the pure
+   * {@link estimateMessage}).
    * @param message - message to price without mutation.
    * @returns content and role-framing tokens under the fixed service heuristic.
    */
   estimateMessage(message: Message): number {
-    return this._estimateContent(message.content) + ROLE_OVERHEAD
+    return estimateMessage(message)
   }
 
   /** Catch one session's fold up to the current durable tail. */
@@ -246,7 +241,7 @@ export class TokenMeterService extends Service {
         )
         const anchorSurfaceTokens = stepStart.surfaceTokens + providerAssistantTokens
         const providerTokens = usageTokens(event.data.usage)
-        const estimatedAnchorTokens = this._estimateHeader(nextHeader) + anchorSurfaceTokens
+        const estimatedAnchorTokens = estimateHeader(nextHeader) + anchorSurfaceTokens
         nextAnchor = {
           header: nextHeader,
           surfaceTokens: anchorSurfaceTokens,
@@ -263,7 +258,7 @@ export class TokenMeterService extends Service {
           surfaceTokens: anchorSurfaceTokens,
           baseline: {
             kind: 'estimated',
-            tokens: this._estimateHeader(nextHeader) + anchorSurfaceTokens,
+            tokens: estimateHeader(nextHeader) + anchorSurfaceTokens,
           },
         }
       }
@@ -355,46 +350,7 @@ export class TokenMeterService extends Service {
       assembler.push(sourceEvent.data.chunk)
     }
     const providerContent = assembler.blocks()
-    return providerContent.length === 0 ? 0 : this._estimateContent(providerContent) + ROLE_OVERHEAD
-  }
-
-  /** Price content blocks recursively under the fixed density heuristic. */
-  private _estimateContent(blocks: readonly ContentBlock[]): number {
-    let tokens = 0
-    for (const block of blocks) {
-      switch (block.type) {
-        case 'text':
-        case 'reasoning':
-          tokens += Math.ceil(block.text.length / CHARS_PER_TOKEN) + BLOCK_OVERHEAD
-          break
-        case 'tool-call':
-          tokens += Math.ceil(block.name.length / CHARS_PER_TOKEN)
-            + Math.ceil(block.arguments.length / CHARS_PER_TOKEN)
-            + BLOCK_OVERHEAD
-          break
-        case 'tool-result':
-          tokens += this._estimateContent(block.content) + BLOCK_OVERHEAD
-          break
-        default:
-          // ContentBlockMap is merge-extensible; unknown blocks retain a
-          // conservative structural JSON price under the fixed heuristic.
-          tokens += BLOCK_OVERHEAD + Math.ceil(JSON.stringify(block).length / CHARS_PER_TOKEN)
-      }
-    }
-    return tokens
-  }
-
-  /** Price the canonical non-surface request envelope. */
-  private _estimateHeader(header: EpochHeader | undefined): number {
-    if (header === undefined) return 0
-    let tokens = 0
-    if (header.system !== undefined) {
-      tokens += Math.ceil(header.system.length / CHARS_PER_TOKEN) + ROLE_OVERHEAD
-    }
-    if (header.tools !== undefined && header.tools.length > 0) {
-      tokens += Math.ceil(JSON.stringify(header.tools).length / CHARS_PER_TOKEN) + BLOCK_OVERHEAD
-    }
-    return tokens
+    return providerContent.length === 0 ? 0 : estimateContent(providerContent) + ROLE_OVERHEAD
   }
 }
 

+ 19 - 0
packages/llm/token-meter/src/projection.ts

@@ -40,11 +40,30 @@ export interface ContextPressureProjection {
   contextWindow?: number
 }
 
+/**
+ * Heuristic composition of the next request's context: what the prompt is
+ * made of, not what it costs. All three figures use the meter's fixed
+ * density estimate (they will not sum exactly to the provider-reported
+ * `pressureTokens`, which is billing-grade and one request behind), and the
+ * message figure tracks the live surface, so it moves as content is appended
+ * or compacted while the provider number holds still.
+ */
+export interface ContextBreakdownProjection {
+  /** Heuristic tokens of the newest request envelope's system prompt; 0 before any request. */
+  systemTokens: number
+  /** Heuristic tokens of the newest request envelope's tool schemas; 0 before any request. */
+  toolsTokens: number
+  /** Heuristic tokens of the current model-visible conversation surface. */
+  messageTokens: number
+}
+
 declare module '@deepseek-ai/dsh-session-projection/types' {
   interface SessionProjectionMap {
     /** Provider-reported usage accumulated across the complete durable log. */
     tokenUsage: TokenUsageProjection
     /** Newest request pressure paired with the newest known route capacity. */
     contextPressure: ContextPressureProjection
+    /** Heuristic system/tools/message composition of the next request. */
+    contextBreakdown: ContextBreakdownProjection
   }
 }

+ 1 - 1
packages/llm/token-meter/src/types.ts

@@ -6,7 +6,7 @@
 
 import type { TokenUsage } from '@deepseek-ai/dsh-llm'
 
-export type { ContextPressureProjection, TokenUsageProjection } from './projection.ts'
+export type { ContextBreakdownProjection, ContextPressureProjection, TokenUsageProjection } from './projection.ts'
 
 /** Token-meter plugin configuration; the fixed estimator has no settings. */
 export type TokenMeterConfig = Record<string, never>

+ 195 - 0
packages/llm/token-meter/tests/context-breakdown-projection.spec.ts

@@ -0,0 +1,195 @@
+// contextBreakdown projection: heuristic system/tools/message composition,
+// plus the shared estimator's pricing branches.
+
+import { describe, expect, it } from 'vitest'
+import { Context } from 'cordis'
+import { createMessage, createUserMessage } from '@deepseek-ai/dsh-llm'
+import type { ContentBlock, ToolSchema } from '@deepseek-ai/dsh-llm'
+import SessionStore from '@deepseek-ai/dsh-session'
+import type { Session, SessionEvent } from '@deepseek-ai/dsh-session'
+import SessionProjectionRegistry from '@deepseek-ai/dsh-session-projection'
+import TokenMeterService from '@deepseek-ai/dsh-token-meter'
+import type { ContextBreakdownProjection } from '@deepseek-ai/dsh-token-meter/client'
+import { contextBreakdownProjectionDefinition } from '../src/breakdown-projection.ts'
+import {
+  estimateContent,
+  estimateHeader,
+  estimateMessage,
+  estimateSystemTokens,
+  estimateToolsTokens,
+} from '../src/estimate.ts'
+
+const CONFIG = { provider: 'test', model: 'test-model' }
+
+const TOOLS: ToolSchema[] = [{
+  name: 'bash',
+  description: 'run a command',
+  parameters: { type: 'object', properties: {} },
+}]
+
+async function harness(): Promise<{ ctx: Context; session: Session }> {
+  const ctx = new Context()
+  await ctx.plugin(SessionStore)
+  await ctx.plugin(SessionProjectionRegistry)
+  await ctx.plugin(TokenMeterService)
+  return { ctx, session: ctx.sessions.create() }
+}
+
+const projected = (ctx: Context, session: Session): ContextBreakdownProjection => {
+  const value = ctx.sessionProjections.snapshot(session).values.contextBreakdown
+  if (value === undefined) throw new Error('contextBreakdown projection is not registered')
+  return value
+}
+
+function appendUser(session: Session, text: string): number {
+  return session.append('user/message', createUserMessage({
+    content: [{ type: 'text', text }],
+    source: { kind: 'user' },
+  }), { surfaceOp: 'append' }).seq
+}
+
+describe('contextBreakdown session projection', () => {
+  it('serves zeros for an empty log', async () => {
+    const { ctx, session } = await harness()
+    expect(projected(ctx, session)).toEqual({ systemTokens: 0, toolsTokens: 0, messageTokens: 0 })
+  })
+
+  it('prices the newest envelope last-wins and pushes no change for a restated one', async () => {
+    const { ctx, session } = await harness()
+    session.append('request/header', {
+      header: { config: CONFIG, system: 'You are terse.', tools: TOOLS },
+      reason: 'initial',
+    })
+    expect(projected(ctx, session)).toEqual({
+      systemTokens: estimateSystemTokens({ config: CONFIG, system: 'You are terse.' }),
+      toolsTokens: estimateToolsTokens({ config: CONFIG, tools: TOOLS }),
+      messageTokens: 0,
+    })
+
+    const changed: string[] = []
+    ctx.sessionProjections.onChanged((_session, key) => { changed.push(key) })
+    session.append('request/header', {
+      header: { config: CONFIG, system: 'You are terse.', tools: TOOLS },
+      reason: 'change',
+    })
+    session.append('todo/write', { todos: [] })
+    expect(changed).not.toContain('contextBreakdown')
+
+    // A system-less, tool-less envelope prices back to zero.
+    session.append('request/header', { header: { config: CONFIG }, reason: 'change' })
+    expect(projected(ctx, session)).toEqual({ systemTokens: 0, toolsTokens: 0, messageTokens: 0 })
+  })
+
+  it('sums surface appends and skips an empty-content assistant message', async () => {
+    const { ctx, session } = await harness()
+    appendUser(session, 'abcd')
+    session.append('step/start', { turn: 1, step: 1 })
+    session.append('assistant/message', {
+      turn: 1,
+      step: 1,
+      message: createMessage({
+        role: 'assistant',
+        content: [],
+        source: { kind: 'model', provider: 'mock', model: 'mock' },
+      }),
+      usage: { inputTokens: 9, outputTokens: 0 },
+    }, { surfaceOp: 'append', sourceEventSeqs: [] })
+    session.append('step/end', { turn: 1, step: 1 })
+    // 'abcd' prices to 9 (1 text + 4 block + 4 role); the usage-only assistant
+    // message derives to no transcript entry and adds nothing.
+    expect(projected(ctx, session).messageTokens).toBe(9)
+  })
+
+  it('shrinks the message figure when a replacement compacts the surface', async () => {
+    const { ctx, session } = await harness()
+    const first = appendUser(session, 'before compaction, a longer message')
+    const second = appendUser(session, 'and a second entry')
+    const summary = createUserMessage({
+      content: [{ type: 'text', text: 'summary' }],
+      source: { kind: 'plugin', plugin: 'test' },
+    })
+    session.append('user/message', summary, {
+      surfaceOp: { op: 'replace', start: first, end: second },
+      sourceEventSeqs: [first, second],
+    })
+    expect(projected(ctx, session).messageTokens).toBe(estimateMessage(summary))
+  })
+
+  it('fails loud on a replace range absent from the folded surface', () => {
+    const definition = contextBreakdownProjectionDefinition
+    const replace = (start: number, end: number): SessionEvent => ({
+      type: 'user/message',
+      seq: 9,
+      time: 0,
+      data: createUserMessage({ content: [{ type: 'text', text: 'x' }], source: { kind: 'user' } }),
+      surfaceOp: { op: 'replace', start, end },
+      sourceEventSeqs: [start, end],
+    } as unknown as SessionEvent)
+    const append = (seq: number): SessionEvent => ({
+      type: 'user/message',
+      seq,
+      time: 0,
+      data: createUserMessage({ content: [{ type: 'text', text: 'x' }], source: { kind: 'user' } }),
+      surfaceOp: 'append',
+    } as unknown as SessionEvent)
+    let state = definition.init()
+    state = definition.apply(state, append(1))
+    state = definition.apply(state, append(3))
+    expect(() => definition.apply(state, replace(7, 3))).toThrow('invalid current range')
+    expect(() => definition.apply(state, replace(1, 7))).toThrow('invalid current range')
+    expect(() => definition.apply(state, replace(3, 1))).toThrow('invalid current range')
+  })
+
+  it('restores from a JSON checkpoint and unregisters with the token-meter fiber', async () => {
+    const ctx = new Context()
+    await ctx.plugin(SessionStore)
+    await ctx.plugin(SessionProjectionRegistry)
+    const meterFiber = await ctx.plugin(TokenMeterService)
+    const session = ctx.sessions.create()
+    session.append('request/header', {
+      header: { config: CONFIG, system: 'You are terse.' },
+      reason: 'initial',
+    })
+    appendUser(session, 'abcd')
+    const checkpoint = JSON.parse(JSON.stringify(
+      ctx.sessionProjections.checkpoint(session),
+    )) as ReturnType<typeof ctx.sessionProjections.checkpoint>
+
+    await meterFiber.dispose()
+    expect(ctx.sessionProjections.snapshot(session).values).not.toHaveProperty('contextBreakdown')
+
+    await ctx.plugin(TokenMeterService)
+    expect(ctx.sessionProjections.viewCheckpoint(checkpoint).contextBreakdown).toEqual({
+      systemTokens: estimateSystemTokens({ config: CONFIG, system: 'You are terse.' }),
+      toolsTokens: 0,
+      messageTokens: 9,
+    })
+  })
+})
+
+describe('shared estimator', () => {
+  it('prices every content-block shape under the fixed heuristic', () => {
+    expect(estimateContent([{ type: 'text', text: 'abcd' }])).toBe(5)
+    expect(estimateContent([{ type: 'reasoning', text: 'abcdefgh' }] as ContentBlock[])).toBe(6)
+    expect(estimateContent([{ type: 'tool-call', id: 'c' as never, name: 'bash', arguments: '{"a":1}' }])).toBe(7)
+    expect(estimateContent([{
+      type: 'tool-result', toolCallId: 'c' as never,
+      content: [{ type: 'text', text: 'abcd' }],
+    }])).toBe(9)
+    const unknown = { type: 'mystery', payload: 'abc' } as unknown as ContentBlock
+    expect(estimateContent([unknown])).toBe(4 + Math.ceil(JSON.stringify(unknown).length / 4))
+  })
+
+  it('prices envelope parts independently and absent parts to zero', () => {
+    expect(estimateSystemTokens(undefined)).toBe(0)
+    expect(estimateSystemTokens({ config: CONFIG })).toBe(0)
+    expect(estimateSystemTokens({ config: CONFIG, system: 'abcdefgh' })).toBe(6)
+    expect(estimateToolsTokens(undefined)).toBe(0)
+    expect(estimateToolsTokens({ config: CONFIG, tools: [] })).toBe(0)
+    expect(estimateToolsTokens({ config: CONFIG, tools: TOOLS }))
+      .toBe(Math.ceil(JSON.stringify(TOOLS).length / 4) + 4)
+    expect(estimateHeader(undefined)).toBe(0)
+    expect(estimateHeader({ config: CONFIG, system: 'abcdefgh', tools: TOOLS }))
+      .toBe(6 + Math.ceil(JSON.stringify(TOOLS).length / 4) + 4)
+  })
+})

Nem az összes módosított fájl került megjelenítésre, mert túl sok fájl változott