Răsfoiți Sursa

Merge pull request #4230 from deepseek-harness/feature/model-image-input-setting

feat(web): add model image input settings
Wenlu Wang 3 săptămâni în urmă
părinte
comite
011bbf8eec
20 a modificat fișierele cu 321 adăugiri și 27 ștergeri
  1. 6 0
      .agents/notes/implemented/feature/2026-09-14-model-image-input-settings.i18n.yaml
  2. 27 0
      .agents/notes/implemented/feature/2026-09-14-model-image-input-settings.md
  3. 27 0
      .agents/notes/implemented/feature/2026-09-14-model-image-input-settings.zh.md
  4. 4 4
      apps/web/tests/expected/deepseek-messages-settings/cards.expected.md
  5. 12 1
      apps/web/tests/expected/models-settings/declared-edit.expected.md
  6. 16 4
      apps/web/tests/expected/onboarding-deepseek-config/default-models.expected.md
  7. 6 1
      apps/web/tests/expected/onboarding-deepseek-config/models.expected.md
  8. 33 0
      apps/web/tests/models-settings.e2e.ts
  9. 18 4
      apps/web/tests/onboarding-deepseek-config.e2e.ts
  10. 2 2
      packages/client/ui-settings-models/README.i18n.yaml
  11. 3 1
      packages/client/ui-settings-models/README.md
  12. 3 1
      packages/client/ui-settings-models/README.zh.md
  13. 10 1
      packages/client/ui-settings-models/src/client/DeepSeekModelsEditor.tsx
  14. 61 0
      packages/client/ui-settings-models/src/client/ModelImageInput.tsx
  15. 9 2
      packages/client/ui-settings-models/src/client/ModelListEditor.tsx
  16. 2 2
      packages/client/ui-settings-models/src/client/ModelsSection.module.css
  17. 12 2
      packages/client/ui-settings-models/src/client/locales.ts
  18. 2 1
      packages/client/ui-settings-models/tests/components.client.spec.tsx
  19. 50 0
      packages/client/ui-settings-models/tests/model-image-input.client.spec.tsx
  20. 18 1
      packages/client/ui-settings-models/tests/provider-form.client.spec.tsx

+ 6 - 0
.agents/notes/implemented/feature/2026-09-14-model-image-input-settings.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-09-14-model-image-input-settings.md
+2026-09-14-model-image-input-settings.md: ef328400332aa58c02a450ead0195eff124c5fdf
+2026-09-14-model-image-input-settings.zh.md: 70ec427b138026124cad2ffdb1fc5f60597abe8a

+ 27 - 0
.agents/notes/implemented/feature/2026-09-14-model-image-input-settings.md

@@ -0,0 +1,27 @@
+# Agent Note: Model image-input settings
+
+Status: implemented
+
+English | [中文](2026-09-14-model-image-input-settings.zh.md)
+
+## Problem
+
+Models settings can edit a model id without exposing the input capabilities that determine whether image attachments are accepted. A custom vision model can therefore appear in the picker while retaining a text-only declaration.
+
+## Decision
+
+Each model row exposes image input under Model options. Supported declares text and image; Not supported declares text only; Default removes the model's explicit input field. DeepSeek writes `inputModalities`, whose absent value means text only. Pi-ai writes `input`, whose absent or empty value inherits the installed model catalog or provider default. Opening a row preserves that inheritance without materializing an override.
+
+The shared field replaces one drafted row and preserves unrelated metadata. Selecting text only or default for DeepSeek also removes its image request limits, because the adapter rejects those limits without image input. Saving uses the existing catalog-array settings mutation and adapter validation. Configuration declares an upstream capability; it does not add image processing to a text-only model.
+
+## Alternatives considered
+
+**Keep `input` editable only in the settings document.** The [earlier pi-ai modality decision](../../archived/architecture/2026-08-12-pi-ai-route-default-input-modalities.md) kept this field outside the model-list editor. That leaves users who add custom vision models through the UI unable to enable their image input there. Per-row editing supplies that configuration while the default choice preserves catalog inheritance.
+
+**A two-state switch.** Treating an absent pi-ai declaration as disabled would misrepresent inherited vision support and encourage overwriting catalog defaults. The explicit Default choice preserves the adapter's existing resolution rules.
+
+**Keep image limits when disabling DeepSeek images.** This leaves a configuration that the adapter refuses to save. Clearing the image-specific limits makes the selected text-only state valid while preserving unrelated model fields.
+
+## Consequences
+
+Users can configure image input for DeepSeek and custom pi-ai model rows through the same control. Restoring defaults can change effective capabilities when the installed catalog or provider defaults change. DeepSeek image limits must be configured again after disabling images. Provider routing and the [catalog recovery rules](../bug-fix/2026-09-07-pi-ai-settings-catalog-recovery.md) remain owned by their existing decisions.

+ 27 - 0
.agents/notes/implemented/feature/2026-09-14-model-image-input-settings.zh.md

@@ -0,0 +1,27 @@
+# Agent Note:模型图片输入设置
+
+Status: implemented
+
+[English](2026-09-14-model-image-input-settings.md) | 中文
+
+## 问题
+
+模型设置可以编辑模型 ID,却没有展示决定图片附件是否被接受的输入能力。因此,自定义视觉模型虽然出现在选择器中,仍可能保留仅文本的声明。
+
+## 决策
+
+每个模型行在「模型选项」下提供图片输入设置。「支持」声明文本和图片;「不支持」声明仅文本;「默认」移除模型的显式输入字段。DeepSeek 写入 `inputModalities`,缺省时表示仅文本。Pi-ai 写入 `input`,缺省或空数组时继承已安装模型目录或提供方默认值。打开模型行保留这种继承,不会生成覆盖值。
+
+共享字段替换一个草稿模型行并保留无关元数据。DeepSeek 选择仅文本或默认时,还会移除图片请求限制,因为适配器在没有图片输入时拒绝这些限制。保存使用现有的模型目录数组设置变更和适配器校验。配置声明上游能力,不会为仅文本模型增加图片处理能力。
+
+## 考虑过的替代方案
+
+**仅允许在设置文档中编辑 `input`。** [早期 pi-ai 输入模态决策](../../archived/architecture/2026-08-12-pi-ai-route-default-input-modalities.md)将该字段留在模型列表编辑器之外。这使通过 UI 添加自定义视觉模型的用户无法在同一界面启用图片输入。逐行编辑提供了该配置,而默认选项保留模型目录继承。
+
+**两态开关。** 将缺省的 pi-ai 声明视为禁用,会错误表达继承的视觉能力,并促使用户覆盖模型目录的默认值。显式「默认」选项保留适配器现有的解析规则。
+
+**禁用 DeepSeek 图片时保留图片限制。** 这会留下适配器拒绝保存的配置。清除图片专属限制,使选择的仅文本状态有效,同时保留无关模型字段。
+
+## 影响
+
+用户可以通过相同控件配置 DeepSeek 和自定义 pi-ai 模型行的图片输入。恢复默认值后,有效能力可能随已安装模型目录或提供方默认值变化。禁用图片后,DeepSeek 图片限制需要重新配置。提供方路由和[模型目录恢复规则](../bug-fix/2026-09-07-pi-ai-settings-catalog-recovery.zh.md)仍由现有决策负责。

+ 4 - 4
apps/web/tests/expected/deepseek-messages-settings/cards.expected.md

@@ -43,7 +43,7 @@
           - textbox "显示名称 1":
             - /placeholder: 显示名称
             - text: DeepSeek-V41-Flash
-          - button "容量 1":
+          - button "模型选项 1":
             - img
           - button "删除模型 1":
             - img
@@ -53,7 +53,7 @@
           - textbox "显示名称 2":
             - /placeholder: 显示名称
             - text: DeepSeek-V4-Flash
-          - button "容量 2":
+          - button "模型选项 2":
             - img
           - button "删除模型 2":
             - img
@@ -63,7 +63,7 @@
           - textbox "显示名称 3":
             - /placeholder: 显示名称
             - text: DeepSeek-V4-Pro
-          - button "容量 3":
+          - button "模型选项 3":
             - img
           - button "删除模型 3":
             - img
@@ -73,7 +73,7 @@
           - textbox "显示名称 4":
             - /placeholder: 显示名称
             - text: DeepSeek-V4-Flash-Vision-Exp
-          - button "容量 4":
+          - button "模型选项 4":
             - img
           - button "删除模型 4":
             - img

+ 12 - 1
apps/web/tests/expected/models-settings/declared-edit.expected.md

@@ -58,8 +58,19 @@
             - text: acme-large
           - textbox "显示名称 1":
             - /placeholder: 显示名称
-          - button "容量 1"
+          - button "模型选项 1" [expanded]
           - button "删除模型 1"
+          - text: 上下文窗口
+          - textbox "上下文窗口 1":
+            - /placeholder: 256K
+          - text: 最大输出 token
+          - textbox "最大输出 token 1":
+            - /placeholder: 32K
+          - text: 图片输入
+          - combobox "图片输入 1":
+            - option "使用默认值"
+            - option "支持" [selected]
+            - option "不支持"
           - button "添加模型"
       - button "取消"
       - button "保存"

+ 16 - 4
apps/web/tests/expected/onboarding-deepseek-config/default-models.expected.md

@@ -43,17 +43,29 @@
           - textbox "显示名称 1":
             - /placeholder: 显示名称
             - text: DeepSeek-V41-Flash
-          - button "容量 1":
+          - button "模型选项 1" [expanded]:
             - img
           - button "删除模型 1":
             - img
+          - text: 上下文窗口
+          - textbox "上下文窗口 1":
+            - /placeholder: 1M
+            - text: 1M
+          - text: 最大输出 token 数
+          - textbox "最大输出 token 数 1":
+            - /placeholder: 256K
+          - text: 图片输入
+          - combobox "图片输入 1":
+            - option "默认(仅文本)"
+            - option "支持" [selected]
+            - option "不支持"
           - textbox "模型 ID 2":
             - /placeholder: 模型 ID
             - text: deepseek-v4-flash
           - textbox "显示名称 2":
             - /placeholder: 显示名称
             - text: DeepSeek-V4-Flash
-          - button "容量 2":
+          - button "模型选项 2":
             - img
           - button "删除模型 2":
             - img
@@ -63,7 +75,7 @@
           - textbox "显示名称 3":
             - /placeholder: 显示名称
             - text: DeepSeek-V4-Pro
-          - button "容量 3":
+          - button "模型选项 3":
             - img
           - button "删除模型 3":
             - img
@@ -73,7 +85,7 @@
           - textbox "显示名称 4":
             - /placeholder: 显示名称
             - text: DeepSeek-V4-Flash-Vision-Exp
-          - button "容量 4":
+          - button "模型选项 4":
             - img
           - button "删除模型 4":
             - img

+ 6 - 1
apps/web/tests/expected/onboarding-deepseek-config/models.expected.md

@@ -44,7 +44,7 @@
           - textbox "显示名称 1":
             - /placeholder: 显示名称
             - text: Private Preview
-          - button "容量 1" [expanded]:
+          - button "模型选项 1" [expanded]:
             - img
           - button "删除模型 1":
             - img
@@ -56,6 +56,11 @@
           - textbox "最大输出 token 数 1":
             - /placeholder: 256K
             - text: 64K
+          - text: 图片输入
+          - combobox "图片输入 1":
+            - option "默认(仅文本)"
+            - option "支持" [selected]
+            - option "不支持"
           - button "添加模型":
             - img
             - text: 添加模型

+ 33 - 0
apps/web/tests/models-settings.e2e.ts

@@ -238,12 +238,18 @@ describe('web e2e: Models settings page configures a dormant provider', () => {
     expect(await dialog.getByLabel('推理强度').count()).toBe(0)
     await dialog.getByRole('button', { name: '添加模型' }).click()
     await dialog.getByLabel('模型 ID 1').fill('acme-large')
+    await dialog.getByRole('button', { name: '模型选项 1' }).click()
+    expect(await dialog.getByLabel('图片输入 1').inputValue()).toBe('default')
+    await dialog.getByLabel('图片输入 1').selectOption('enabled')
     await dialog.getByRole('button', { name: '创建提供方', exact: true }).click()
 
     const row = dialog.getByText('Acme Gateway', { exact: true }).first()
     await row.waitFor({ timeout: 10_000 })
     const document = await readFile(join(scaffold.harnessHome, 'settings.yaml'), 'utf8')
     expect(document).toContain('acme-gateway:')
+    await expect(scaffold.ctx.llm.resolveModelInfo('acme-gateway', 'acme-large')).resolves.toMatchObject({
+      inputModalities: ['text', 'image'],
+    })
 
     // The tag follows the adapter's installed catalog: this route is in no
     // catalog, while minimax-cn is — even though both now have profiles.
@@ -269,11 +275,14 @@ describe('web e2e: Models settings page configures a dormant provider', () => {
     expect(await protocol.inputValue()).toBe('openai-completions')
     const name = dialog.getByLabel('显示名称', { exact: true })
     expect(await name.inputValue()).toBe('Acme Gateway')
+    await dialog.getByRole('button', { name: '模型选项 1' }).click()
+    expect(await dialog.getByLabel('图片输入 1').inputValue()).toBe('enabled')
     const snapshot = await captureStableAria(page, '[role="dialog"]', scaffold.workspaceCwd)
     await compareOrRefreshGolden(DECLARED_EDIT_EXPECTED, snapshot, MODE)
 
     await protocol.selectOption('anthropic-messages')
     await name.fill('Acme 网关')
+    await dialog.getByLabel('图片输入 1').selectOption('disabled')
     await dialog.getByRole('button', { name: '保存', exact: true }).click()
     await expect.poll(async () => dialog.getByLabel('API 协议').count(), { timeout: 10_000 }).toBe(0)
     // The adapter re-resolved the route under the new protocol and re-registered
@@ -287,6 +296,30 @@ describe('web e2e: Models settings page configures a dormant provider', () => {
     const document = await readFile(join(scaffold.harnessHome, 'settings.yaml'), 'utf8')
     expect(document).toContain('api: anthropic-messages')
     expect(document).toContain('displayName: Acme 网关')
+    await expect(scaffold.ctx.llm.resolveModelInfo('acme-gateway', 'acme-large')).resolves.toMatchObject({
+      inputModalities: ['text'],
+    })
+    expect(tripwire.pageErrors).toEqual([])
+  }, 60_000)
+
+  it('restores the provider default for image input', async () => {
+    onTestFailed(() => saveFailureShot(page, 'web-e2e-models-image-default'))
+    const dialog = page.getByRole('dialog', { name: '设置' })
+    await dialog.getByRole('button', { name: '编辑 Acme 网关 (acme-gateway)' }).click()
+    await dialog.getByText('自定义设置').click()
+    await dialog.getByRole('button', { name: '模型选项 1' }).click()
+    expect(await dialog.getByLabel('图片输入 1').inputValue()).toBe('disabled')
+    await dialog.getByLabel('图片输入 1').selectOption('default')
+    await dialog.getByRole('button', { name: '保存', exact: true }).click()
+    await dialog.getByLabel('模型 ID 1').waitFor({ state: 'detached', timeout: 10_000 })
+    await expect(scaffold.ctx.llm.resolveModelInfo('acme-gateway', 'acme-large')).resolves.toMatchObject({
+      inputModalities: ['text'],
+    })
+    await dialog.getByRole('button', { name: '编辑 Acme 网关 (acme-gateway)' }).click()
+    await dialog.getByText('自定义设置').click()
+    await dialog.getByRole('button', { name: '模型选项 1' }).click()
+    expect(await dialog.getByLabel('图片输入 1').inputValue()).toBe('default')
+    await dialog.getByRole('button', { name: '取消', exact: true }).click()
     expect(tripwire.pageErrors).toEqual([])
   }, 60_000)
 

+ 18 - 4
apps/web/tests/onboarding-deepseek-config.e2e.ts

@@ -209,19 +209,24 @@ describe.skipIf(MODE === 'record')('web e2e: first-run DeepSeek credential setup
     expect(await settings.getByLabel('模型 ID 3').inputValue()).toBe('deepseek-v4-pro')
     expect(await settings.getByLabel('模型 ID 4').inputValue()).toBe('deepseek-v4-flash-vision-exp')
     expect(await settings.getByRole('button', { name: /删除模型/ }).count()).toBe(4)
+    await settings.getByRole('button', { name: '模型选项 1' }).click()
+    expect(await settings.getByLabel('图片输入 1').inputValue()).toBe('enabled')
     const defaultModels = await captureStableAria(page, '[role="dialog"]', scaffold.workspaceCwd)
     await compareOrRefreshGolden(DEFAULT_MODELS_EXPECTED, defaultModels, MODE)
     await settings.getByLabel('显示名称 1').fill('Configured Flash')
+    await settings.getByLabel('图片输入 1').selectOption('disabled')
     await settings.getByRole('button', { name: '保存', exact: true }).click()
     await settings.getByLabel('模型 ID 1').waitFor({ state: 'detached', timeout: 15_000 })
     const savedDefaults = await readFile(join(scaffold.harnessHome, 'settings.yaml'), 'utf8')
     expect(savedDefaults).toContain('id: deepseek-flash')
     expect(savedDefaults).toContain('inputModalities:')
     expect(savedDefaults).toContain('- text')
-    expect(savedDefaults).toContain('- image')
     expect(savedDefaults).toContain('systemPromptUpdate: in-history')
     await expect(scaffold.ctx.llm.resolveModelInfo('deepseek-official', 'deepseek-flash')).resolves.toMatchObject({
-      name: 'Configured Flash', inputModalities: ['text', 'image'], systemPromptUpdate: 'in-history',
+      name: 'Configured Flash', inputModalities: ['text'], systemPromptUpdate: 'in-history',
+    })
+    await expect(scaffold.ctx.llm.resolveModelInfo('deepseek-official', 'deepseek-v4-flash-vision-exp')).resolves.toMatchObject({
+      inputModalities: ['text', 'image'],
     })
     await deepSeek.locator('xpath=ancestor::li').getByRole('button', { name: '编辑' }).click()
     await settings.getByText('自定义设置').click()
@@ -232,10 +237,11 @@ describe.skipIf(MODE === 'record')('web e2e: first-run DeepSeek credential setup
     const customModelId = settings.getByLabel('模型 ID 1')
     await customModelId.fill('private-preview')
     await settings.getByLabel('显示名称 1').fill('Private Preview')
-    // Capacities live behind the row's own disclosure, as in the pi-ai form.
-    await settings.getByRole('button', { name: '容量 1' }).click()
+    await settings.getByRole('button', { name: '模型选项 1' }).click()
     await settings.getByLabel('上下文窗口 1').fill('131072')
     await settings.getByLabel('最大输出 token 数 1').fill('64K')
+    expect(await settings.getByLabel('图片输入 1').inputValue()).toBe('default')
+    await settings.getByLabel('图片输入 1').selectOption('enabled')
 
     await expect.poll(
       () => settings.getByLabel('API 密钥', { exact: true }).getAttribute('placeholder'),
@@ -252,6 +258,14 @@ describe.skipIf(MODE === 'record')('web e2e: first-run DeepSeek credential setup
     expect(document).toContain('contextWindow: 131072')
     expect(document).toContain('maxTokens: 64000')
     expect(document).not.toContain('id: deepseek-flash')
+    await expect(scaffold.ctx.llm.resolveModelInfo('deepseek-official', 'private-preview')).resolves.toMatchObject({
+      inputModalities: ['text', 'image'],
+    })
+    await deepSeek.locator('xpath=ancestor::li').getByRole('button', { name: '编辑' }).click()
+    await settings.getByText('自定义设置').click()
+    await settings.getByRole('button', { name: '模型选项 1' }).click()
+    expect(await settings.getByLabel('图片输入 1').inputValue()).toBe('enabled')
+    await settings.getByRole('button', { name: '取消', exact: true }).click()
 
     await page.keyboard.press('Escape')
     // A connected Workspace is what puts a live composer — and its model

+ 2 - 2
packages/client/ui-settings-models/README.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write packages/client/ui-settings-models/README.md
-README.md: f1314b1db9ee231b2f67777491ac57aefc5ea85a
-README.zh.md: 17861766c52333e90947a6c643c7df8cab4992b4
+README.md: a26cd22f8785f8f2c948ac74560ff85d80097577
+README.zh.md: 82b9e031b477aa83c59acdfbade38f8282adefbd

+ 3 - 1
packages/client/ui-settings-models/README.md

@@ -35,10 +35,12 @@ The primary field on an editor card is a single **API key** input — the page n
 
 ### Editing a provider
 
-The collapsed 自定义设置 fold carries the curated extras: `baseURL` for both families (the deepseek placeholder shows the public endpoint), each adapter's model catalog, and the **display name** and **API protocol** of a pi-ai route the adapter does not ship. Profile `headers` remain deployment configuration in `settings.yaml` or Cordis config and have no Models-page editor. The Provider ID stays fixed: it is the settings key, the name every other namespace and every logged session references, and the stem of a credential reference the page cannot read back to move. Reasoning effort is deliberately not among the editable fields: it is a per-model capability, so a provider-scoped control could only be set to a value some models reject. Each DeepSeek row edits `id`, optional display `name`, and optional `contextWindow`/`maxTokens`; existing fields outside that curated set survive edits.
+The collapsed 自定义设置 fold carries the curated extras: `baseURL` for both families (the deepseek placeholder shows the public endpoint), each adapter's model catalog, and the **display name** and **API protocol** of a pi-ai route the adapter does not ship. Profile `headers` remain deployment configuration in `settings.yaml` or Cordis config and have no Models-page editor. The Provider ID stays fixed: it is the settings key, the name every other namespace and every logged session references, and the stem of a credential reference the page cannot read back to move. Reasoning effort is deliberately not among the editable fields: it is a per-model capability, so a provider-scoped control could only be set to a value some models reject. Each model row edits `id`, optional display `name`, optional `contextWindow`/`maxTokens`, and image input; unrelated model fields survive edits.
 
 The DeepSeek card edits the shared `llm-deepseek` endpoint, credentials, and model catalog without a protocol selector. When Cordis YAML selects Messages, the public endpoint placeholder is `https://api.deepseek.com/anthropic`. Saving the card preserves protocol configuration.
 
+Expand **Customized settings → Model options → Image input** to declare whether the model supports images. **Supported** writes text and image input; **Not supported** writes text only. DeepSeek's **Default (text only)** removes `inputModalities`; an absent declaration means text only. Pi-ai's **Use default** removes `input` and inherits the installed model catalog or provider default. Selecting **Not supported** or **Default (text only)** for DeepSeek also removes `imagePixelBudget` and `imageMaxBytes`, which its adapter rejects without image input. Declare support only for models that can actually process images.
+
 ### Adding and deleting providers
 
 The add flow is a card carrying the dormant-directory provider select — a bare-mounted `llm-pi-ai` offers its whole installed catalog before any route exists. **Add a custom provider** declares a route pi-ai does not ship; the create card asks for a unique **Provider ID**, an endpoint, a protocol, and at least one uniquely-identified model, because nothing can default those. The endpoint must be a parseable HTTP or HTTPS URL; localhost, IPv4 and IPv6 literals, and custom ports remain valid. A syntax error blocks both discovery and creation at the field, while a request failure remains a separate provider error. **Fetch available models** asks the `llm/discoverModels` Remote about the endpoint the form shows, so adding a provider is one pass instead of save-then-return; the reply opens a searchable picker rather than being written, and nothing is written until **Add selected**. Each selected candidate copies its id, display name, context window, and output-token cap into the editable row when disclosed, while an existing row retains its user-tuned values. Search matches model ids and optional display names without clearing hidden selections. **Select all** adds the visible results, while **Deselect all** clears the entire selection so hidden results cannot be adopted accidentally. A row is deletable only when the user layer alone carries it (removal restores the composition base), and its confirmation dialog names the provider.

+ 3 - 1
packages/client/ui-settings-models/README.zh.md

@@ -35,10 +35,12 @@ kind: "package-reference"
 
 ### 编辑提供方
 
-收起的「自定义设置」折叠区承载精选的额外字段:两个家族都有 `baseURL`(deepseek 的占位符显示公共端点)、各适配器自己的模型目录,以及适配器未提供的 pi-ai 路由的**显示名称**与 **API 协议**。Profile `headers` 仍是 `settings.yaml` 或 Cordis 配置中的部署配置,Models 页面不提供编辑器。Provider ID 保持固定:它是 settings 的键、其他每个 namespace 与每一条已记录会话引用的名字,也是页面读不回、因而搬不走的凭据引用词干。推理等级刻意不在可编辑字段之列:它是按模型的能力,提供方级的控件只可能被设成某些模型会拒绝的值。每个 DeepSeek 行编辑 `id`、可选显示 `name` 与可选 `contextWindow`/`maxTokens`;该精选集之外的现有字段在编辑后仍会保留。
+收起的「自定义设置」折叠区承载精选的额外字段:两个家族都有 `baseURL`(deepseek 的占位符显示公共端点)、各适配器自己的模型目录,以及适配器未提供的 pi-ai 路由的**显示名称**与 **API 协议**。Profile `headers` 仍是 `settings.yaml` 或 Cordis 配置中的部署配置,Models 页面不提供编辑器。Provider ID 保持固定:它是 settings 的键、其他每个 namespace 与每一条已记录会话引用的名字,也是页面读不回、因而搬不走的凭据引用词干。推理等级刻意不在可编辑字段之列:它是按模型的能力,提供方级的控件只可能被设成某些模型会拒绝的值。每个模型行可编辑 `id`、可选显示 `name`、可选 `contextWindow`/`maxTokens` 和图片输入;无关的模型字段在编辑后仍会保留。
 
 `llm-deepseek` 的 DeepSeek 卡片编辑共用的端点、凭据和模型目录,不提供协议选择器。Cordis YAML 选择 Messages 时,官方端点占位符为 `https://api.deepseek.com/anthropic`;保存卡片不会改写协议配置。
 
+展开**自定义设置 → 模型选项 → 图片输入**,声明模型是否支持图片。选择**支持**会写入文本和图片输入;选择**不支持**会写入仅文本。DeepSeek 的**默认(仅文本)**会移除 `inputModalities`;缺省的声明表示仅文本。Pi-ai 的**使用默认值**会移除 `input`,继承已安装模型目录或提供方的默认值。DeepSeek 选择**不支持**或**默认(仅文本)**时,还会移除 `imagePixelBudget` 和 `imageMaxBytes`,因为适配器在没有图片输入时拒绝这些限制。仅为实际能够处理图片的模型声明支持。
+
 ### 新增与删除提供方
 
 「新增」流程是一张承载休眠目录提供方选择框的卡片——裸挂载的 `llm-pi-ai` 在任何路由存在之前就能提供其完整的已安装 catalog。**添加自定义提供方**声明一条 pi-ai 不提供的路由;创建卡片会索要唯一的 **Provider ID**、端点、协议与至少一个可唯一识别的模型,因为没有东西能为它们兜底。端点必须是可解析的 HTTP 或 HTTPS URL;localhost、IPv4 与 IPv6 字面地址以及自定义端口仍然有效。语法错误会在字段处阻止询问与创建,请求失败则继续作为独立的提供方错误显示。**获取可用模型**通过 `llm/discoverModels` Remote 查询表单显示的端点,因此新增提供方一次即可完成,而非先保存再返回;回复打开的是可搜索选择器而非直接写入,只有点击**添加所选**才会写入。每个选中候选会在提供方公布相应信息时,把 id、显示名、上下文窗口与最大输出 token 数复制进可编辑行;已经存在的行保留用户调整过的值。搜索会匹配模型 id 与可选显示名称,且不会清除隐藏项的勾选状态。**全选**会加入可见结果,而**取消全选**会清空全部勾选,以免意外采用隐藏结果。只有用户层单独携带某行时,该行才可删除(删除会恢复组合基线),其确认对话框会指名该提供方。

+ 10 - 1
packages/client/ui-settings-models/src/client/DeepSeekModelsEditor.tsx

@@ -11,6 +11,7 @@ import {
   IconChevronDownOutline14, IconChevronRightOutline14, IconPlusOutline16, IconTrashOutline16,
 } from '@deepseek-ai/dsh-client-ui-primitives'
 import type { en } from './locales.ts'
+import { ModelImageInput } from './ModelImageInput.tsx'
 import styles from './ModelsSection.module.css'
 
 /** One catalog entry kept structurally open so hidden or future fields survive an edit. */
@@ -144,7 +145,7 @@ export interface DeepSeekModelsEditorProps {
 
 /**
  * Render the direct DeepSeek adapter's model catalog: id and display name on
- * each row, capacities behind the row's own disclosure.
+ * each row, capacities and image support behind the row's own disclosure.
  * @param props - effective rows plus the array-level override actions.
  * @returns the catalog editor.
  */
@@ -343,6 +344,14 @@ export function DeepSeekModelsEditor(props: DeepSeekModelsEditorProps): ReactNod
                     <div className={styles['modelAdvanced']}>
                       {capacityField(model, index, 'contextWindow', props.defaultContextWindow)}
                       {capacityField(model, index, 'maxTokens', props.defaultMaxTokens)}
+                      <ModelImageInput
+                        model={model}
+                        field="inputModalities"
+                        position={index + 1}
+                        disabled={props.disabled}
+                        t={props.t}
+                        onChange={(next) => { props.onChange(props.models.map((row, at) => at === index ? next : row)) }}
+                      />
                     </div>
                   )
                   : null}

+ 61 - 0
packages/client/ui-settings-models/src/client/ModelImageInput.tsx

@@ -0,0 +1,61 @@
+/** Image-input declarations shared by the DeepSeek and pi-ai catalog editors. */
+
+import type { ReactNode } from 'react'
+import type { DeepSeekModelDraft } from './DeepSeekModelsEditor.tsx'
+import type { ModelsKey } from './locales.ts'
+import styles from './ModelsSection.module.css'
+
+/** Props of {@link ModelImageInput}. */
+interface ModelImageInputProps {
+  /** Effective model row, including fields outside the curated editor. */
+  model: DeepSeekModelDraft
+  /** Adapter-owned field; pi-ai inherits capabilities when absent or empty. */
+  field: 'inputModalities' | 'input'
+  /** One-based row position for the accessible label. */
+  position: number
+  /** Prevent changes while read-only or saving. */
+  disabled: boolean
+  /** Section copy. */
+  t: (key: ModelsKey) => string
+  /** Replace this row, preserving unrelated configuration. */
+  onChange: (model: DeepSeekModelDraft) => void
+}
+
+/**
+ * Edit image support while retaining an explicit choice to inherit defaults.
+ * @param props - model declaration and row replacement action.
+ * @returns the labeled image-input selector.
+ */
+export function ModelImageInput({ model, field, position, disabled, t, onChange }: ModelImageInputProps): ReactNode {
+  const modalities = model[field]
+  const value = !Array.isArray(modalities) || modalities.length === 0
+    ? 'default'
+    : modalities.includes('image') ? 'enabled' : 'disabled'
+  return (
+    <label className={styles['modelField']}>
+      <span className={styles['modelFieldLabel']}>{t('modelImageInput')}</span>
+      <select
+        className={`${styles['input']} ${styles['selectInput']}`}
+        aria-label={`${t('modelImageInput')} ${String(position)}`}
+        value={value}
+        disabled={disabled}
+        onChange={(event) => {
+          const choice = event.target.value
+          const next = { ...model }
+          if (choice === 'default') Reflect.deleteProperty(next, field)
+          else next[field] = choice === 'enabled' ? ['text', 'image'] : ['text']
+          // DeepSeek rejects image request limits on a text-only model.
+          if (field === 'inputModalities' && choice !== 'enabled') {
+            Reflect.deleteProperty(next, 'imagePixelBudget')
+            Reflect.deleteProperty(next, 'imageMaxBytes')
+          }
+          onChange(next)
+        }}
+      >
+        <option value="default">{t(field === 'inputModalities' ? 'modelImageDefaultText' : 'modelImageDefault')}</option>
+        <option value="enabled">{t('modelImageEnabled')}</option>
+        <option value="disabled">{t('modelImageDisabled')}</option>
+      </select>
+    </label>
+  )
+}

+ 9 - 2
packages/client/ui-settings-models/src/client/ModelListEditor.tsx

@@ -22,6 +22,7 @@ import { formatCapacity, parseCapacity } from './DeepSeekModelsEditor.tsx'
 import type { ModelsOperations } from './operations.ts'
 import type { DeepSeekModelDraft } from './DeepSeekModelsEditor.tsx'
 import type { en } from './locales.ts'
+import { ModelImageInput } from './ModelImageInput.tsx'
 import styles from './ModelsSection.module.css'
 
 /**
@@ -163,8 +164,6 @@ export function ModelListEditor(props: ModelListEditorProps): ReactNode {
   const [candidates, setCandidates] = useState<readonly LlmDiscoveredModel[] | undefined>(undefined)
   const [picked, setPicked] = useState<ReadonlySet<string>>(new Set())
   const [candidateQuery, setCandidateQuery] = useState('')
-  // Rows carry an id and a name; capacities are the exception, so they stay
-  // folded until asked for rather than crowding every row with four inputs.
   const [expanded, setExpanded] = useState<ReadonlySet<number>>(new Set())
   // Capacities are edited as text, so a field's keystrokes are held here rather
   // than re-derived from the parsed count on every change — that would rewrite
@@ -433,6 +432,14 @@ export function ModelListEditor(props: ModelListEditorProps): ReactNode {
                     onChange={(event) => { editCapacity(index, 'maxTokens', event.target.value) }}
                   />
                 </label>
+                <ModelImageInput
+                  model={model}
+                  field="input"
+                  position={index + 1}
+                  disabled={disabled}
+                  t={t}
+                  onChange={(next) => { onChange(models.map((row, at) => at === index ? next : row)) }}
+                />
               </div>
             )
             : null}

+ 2 - 2
packages/client/ui-settings-models/src/client/ModelsSection.module.css

@@ -432,8 +432,8 @@
 }
 
 /* Model list, shared with the pi-ai provider form: one bordered
-   entry per model, id and display name on the row, capacities behind the
-   row's own disclosure. The rules use this stylesheet's token vocabulary —
+   entry per model, id and display name on the row, capacities and image input
+   behind the row's own disclosure. The rules use this stylesheet's token vocabulary —
    `--dsw-alias-border-subtle`, `--dsw-alias-text-tertiary`, and
    `--dsw-alias-text-primary` are undefined here and would resolve to their
    light-mode literals. */

+ 12 - 2
packages/client/ui-settings-models/src/client/locales.ts

@@ -50,7 +50,12 @@ export const en = {
   contextWindowPlaceholder: 'Uses the provider default',
   maxTokens: 'Max output tokens',
   maxTokensPlaceholder: 'Uses the provider default',
-  modelAdvanced: 'Capacities',
+  modelAdvanced: 'Model options',
+  modelImageInput: 'Image input',
+  modelImageDefault: 'Use default',
+  modelImageDefaultText: 'Default (text only)',
+  modelImageEnabled: 'Supported',
+  modelImageDisabled: 'Not supported',
   addModel: 'Add model',
   removeModel: 'Delete model',
   modelsEmpty: 'No models will be shown in the selector. Unlisted IDs can still be sent directly.',
@@ -160,7 +165,12 @@ export const zh: { [Key in keyof typeof en]: string } = {
   contextWindowPlaceholder: '使用提供方默认值',
   maxTokens: '最大输出 token 数',
   maxTokensPlaceholder: '使用提供方默认值',
-  modelAdvanced: '容量',
+  modelAdvanced: '模型选项',
+  modelImageInput: '图片输入',
+  modelImageDefault: '使用默认值',
+  modelImageDefaultText: '默认(仅文本)',
+  modelImageEnabled: '支持',
+  modelImageDisabled: '不支持',
   addModel: '添加模型',
   removeModel: '删除模型',
   modelsEmpty: '模型选择器中将不显示任何模型;目录外 ID 仍可直接发送。',

+ 2 - 1
packages/client/ui-settings-models/tests/components.client.spec.tsx

@@ -648,6 +648,7 @@ describe('ModelsSection', () => {
     fireEvent.change(names[2] as HTMLInputElement, { target: { value: 'Private Preview' } })
     // Only row 3 is open, so its capacity is addressed by its own label.
     fireEvent.change(screen.getByLabelText(`${en.contextWindow} 3`), { target: { value: '131072' } })
+    fireEvent.change(screen.getByLabelText(`${en.modelImageInput} 3`), { target: { value: 'enabled' } })
     fireEvent.click(screen.getByText(en.apply))
 
     await waitFor(() => { expect(mutate).toHaveBeenCalledTimes(1) })
@@ -658,7 +659,7 @@ describe('ModelsSection', () => {
         path: ['models'],
         value: [
           ...DEFAULT_DEEPSEEK_MODELS,
-          { id: 'private-preview', name: 'Private Preview', contextWindow: 131_072 },
+          { id: 'private-preview', name: 'Private Preview', contextWindow: 131_072, inputModalities: ['text', 'image'] },
         ],
       }],
       0,

+ 50 - 0
packages/client/ui-settings-models/tests/model-image-input.client.spec.tsx

@@ -0,0 +1,50 @@
+// @vitest-environment jsdom
+/** Image capability defaults, explicit choices, and hidden model metadata. */
+import { cleanup, fireEvent, render, screen } from '@testing-library/react'
+import { afterEach, describe, expect, it, vi } from 'vitest'
+import { ModelImageInput } from '../src/client/ModelImageInput.tsx'
+import { en } from '../src/client/locales.ts'
+
+afterEach(cleanup)
+
+describe.each(['inputModalities', 'input'] as const)('%s image input', (field) => {
+  it.each([
+    [undefined, 'default'],
+    [[], 'default'],
+    [['text'], 'disabled'],
+    [['text', 'image'], 'enabled'],
+    [['image'], 'enabled'],
+  ] as const)('displays %j without materializing an override', (modalities, selected) => {
+    const onChange = vi.fn()
+    render(<ModelImageInput model={{ id: 'preview', [field]: modalities }} field={field} position={2} disabled={false} t={key => en[key]} onChange={onChange} />)
+    expect(screen.getByRole<HTMLSelectElement>('combobox', { name: `${en.modelImageInput} 2` }).value).toBe(selected)
+    expect(onChange).not.toHaveBeenCalled()
+  })
+
+  it('enables images and keeps unrelated metadata', () => {
+    const onChange = vi.fn()
+    const model = { id: 'preview', contextWindow: 123456, systemPromptUpdate: 'in-history' }
+    render(<ModelImageInput model={model} field={field} position={1} disabled={false} t={key => en[key]} onChange={onChange} />)
+    fireEvent.change(screen.getByRole('combobox'), { target: { value: 'enabled' } })
+    expect(onChange).toHaveBeenCalledWith({ ...model, [field]: ['text', 'image'] })
+    expect(model).not.toHaveProperty(field)
+  })
+
+  it.each(['disabled', 'default'])('selects %s without leaving invalid DeepSeek image limits', (choice) => {
+    const onChange = vi.fn()
+    const model = { id: 'vision', [field]: ['image'], description: 'kept', imagePixelBudget: 'low', imageMaxBytes: 12345 }
+    render(<ModelImageInput model={model} field={field} position={1} disabled={false} t={key => en[key]} onChange={onChange} />)
+    fireEvent.change(screen.getByRole('combobox'), { target: { value: choice } })
+    expect(onChange).toHaveBeenCalledWith({
+      id: 'vision', description: 'kept',
+      ...choice === 'disabled' ? { [field]: ['text'] } : {},
+      ...field === 'input' ? { imagePixelBudget: 'low', imageMaxBytes: 12345 } : {},
+    })
+    expect(model[field]).toEqual(['image'])
+  })
+
+  it('disables the selector while read-only or saving', () => {
+    render(<ModelImageInput model={{ id: 'preview' }} field={field} position={1} disabled t={key => en[key]} onChange={vi.fn()} />)
+    expect(screen.getByRole<HTMLSelectElement>('combobox').disabled).toBe(true)
+  })
+})

+ 18 - 1
packages/client/ui-settings-models/tests/provider-form.client.spec.tsx

@@ -256,6 +256,22 @@ describe('protocolChoices', () => {
 })
 
 describe('model list editing', () => {
+  it('changes image input without rewriting a neighboring model declaration', async () => {
+    const neighbor = { id: 'vision', input: ['image'], name: 'Kept vision model' }
+    const { mutate } = await mountSection({
+      providers: { openai: { models: [{ id: 'preview' }, neighbor] } },
+    })
+    openEditor('openai')
+    expandModel(1)
+    fireEvent.change(screen.getByLabelText(`${en.modelImageInput} 1`), { target: { value: 'enabled' } })
+    fireEvent.click(screen.getByText(en.apply))
+    await waitFor(() => { expect(mutate).toHaveBeenCalled() })
+    expect(firstMutate(mutate).ops).toEqual([{
+      op: 'set', path: ['providers', 'openai', 'models'],
+      value: [{ id: 'preview', input: ['text', 'image'] }, neighbor],
+    }])
+  })
+
   it('adds, edits, and removes rows without storing emptied optional fields', async () => {
     const { mutate } = await mountSection()
     openEditor('openai')
@@ -264,6 +280,7 @@ describe('model list editing', () => {
     fireEvent.change(screen.getByLabelText(`${en.modelId} 1`), { target: { value: 'acme-large' } })
     expandModel(1)
     fireEvent.change(screen.getByLabelText(`${en.modelContextWindow} 1`), { target: { value: '65536' } })
+    fireEvent.change(screen.getByLabelText(`${en.modelImageInput} 1`), { target: { value: 'enabled' } })
     fireEvent.change(screen.getByLabelText(`${en.modelName} 1`), { target: { value: 'Acme' } })
     // Clearing an optional field must drop it rather than store an empty value.
     fireEvent.change(screen.getByLabelText(`${en.modelName} 1`), { target: { value: '' } })
@@ -273,7 +290,7 @@ describe('model list editing', () => {
     expect(firstMutate(mutate)).toMatchObject({
       ns: 'llm-pi-ai',
       expectedRevision: 3,
-      ops: [{ op: 'set', path: ['providers', 'openai', 'models'], value: [{ id: 'acme-large', contextWindow: 65_536 }] }],
+      ops: [{ op: 'set', path: ['providers', 'openai', 'models'], value: [{ id: 'acme-large', contextWindow: 65_536, input: ['text', 'image'] }] }],
     })
   })