aryswisnu/dsh-eval-regression
dsh-eval-regression is a small, deterministic regression-evaluation plugin for DeepSeek Harness. It registers evaluate_golden_output, a model-callable tool that compares candidate output against required and forbidden fragments. It does not call a model, persist data, or claim semantic correctness — its job is repeatable pass/fail evidence: required fragments catch omissions, forbidden fragments catch known bad claims or unsafe fallbacks, per-case reports make failures reviewable, and deterministic scoring suits CI thresholds. A small CLI runs version-controlled suites in CI.
Install
dsh plugin --profile web add github:aryswisnu/dsh-eval-regressiondsh plugin --profile web add github:aryswisnu/dsh-eval-regression (README 2026-08-31; npm name not published — 404 verified — GitHub install documented). Local dev: git clone ... && cd dsh-eval-regression && npm install && npm run build && dsh plugin --profile web add .. The bundle's cordis.patch.yml registers the tool automatically.
Compatibility
DSH profile; deterministic, model-callable tool; also ships a CLI for CI suites; MIT.
Details
- Repo: aryswisnu/dsh-eval-regression
- Category: Other
- Stars: 2
- Version: repo aryswisnu/dsh-eval-regression (2★)
- Last push: 2026-08-13
- First seen: 2026-08-13
Recent updates
evaluate_golden_output tool; required/forbidden fragment checks; deterministic scoring; per-case reports; CI CLI; no model calls, no persistence.
FAQ
- Does it call a model?
- No — the tool is fully deterministic; it compares candidate output against required and forbidden fragments and reports pass/fail evidence.
- Can I run suites in CI?
- Yes — the plugin ships a small CLI that runs version-controlled golden suites, suitable for release smoke tests and replayed transcripts.
- Why deterministic?
- So regressions are caught with transparent, reviewable, repeatable evidence instead of vibes-based evaluation.
Alternatives
jkrandom-sudo/dsh-ci-doctor · stardustlc666/dsh-flakefinder · dongsheng123132/dsh-benchmark