|
|
2 meses atrás | |
|---|---|---|
| .. | ||
| src | 2 meses atrás | |
| tests | 2 meses atrás | |
| README.md | 2 meses atrás | |
| package.json | 2 meses atrás | |
| tsconfig.json | 2 meses atrás | |
@deepseek-ai/dsh-llm-retryFunction plugin that retries selected transient model-request failures through the agent/request-error waterfall. It does not wrap ctx.llm.stream(): every adapter call remains one provider attempt, and every retry opens a fresh numbered turn.
The default policy permits two retries for RATE_LIMIT, SERVER, TIMEOUT, and TRANSPORT, using bounded exponential backoff from 500 ms to 10 seconds with 10 percent jitter. Delay bounds must fit Node's supported timer range. A valid providerRetryAfterMs replaces local backoff when it is within the configured cap; an over-cap instruction delegates to the next recovery policy instead.
The recovery listener appends a non-surface llm/retry event after the failed step, waits for the backoff while the failed turn's signal remains live, then calls agent.retry(). The loop closes that failed turn and opens a retry turn over the same durable history. The policy keeps its own retry count across that uninterrupted recovery chain and clears it at terminal agent/idle. Turn cancellation and plugin disposal abort the wait.
The separately published ./invariant companion checks that every retry record appears inside an open turn after its failed step, matches its position in the current retry chain, and carries a positive bounded retry budget and non-negative bounded timer delay. Full jitter may schedule zero milliseconds at its lower boundary.
- name: '@deepseek-ai/dsh-llm-retry'
config:
maxTransientRetries: 2
initialDelayMs: 500
maxDelayMs: 10000
jitterRatio: 0.1
retryableCodes: [RATE_LIMIT, SERVER, TIMEOUT, TRANSPORT]
No retry event, delay, or failure prose is model-visible. The retry turn reconstructs the same explicit provider/model request from durable session history; failed chunks never enter derived messages.
Each retry is a new provider request and may repeat input-token billing. The finite budget caps attempts; llm/retry itself contributes no tokens.
The reconstructed request preserves the prior prefix and is eligible for provider cache reuse under that provider's rules. The non-surface status event does not change cache identity.
ctx.llm.stream() consumers remain single-attempt because a raw stream cannot separate already-emitted chunks durably.llm/retry records completed backoff, not request completion — later step and turn events establish success, exhaustion, or cancellation.