DSH Eval Harness
BiBoyang/dsh-eval-harness · tools
DSH plug-in evaluation tool: YAML use case driven real agent regression evaluation + baseline comparison PASS/WARN/FAIL access control|Regression eval harness for DeepSeek Harness plugins
★ 14GitHub stars
1Forks
2026-09-03Last updated
TypeScriptLanguage
—License
Key features
- eval_runRun all use cases under cases_dir: headless drives real agent → collect session trace → assert → write report.json/report.md
- evalTeach the model to help users write evaluation use cases (use case format, key points for writing assertions, parsing subset constraints)
- LLM-as-judge (output_judge)expresses semantic expectations that cannot be written regularly such as "explain the reasons instead of just giving the conclusion".
- eval_gateCompare baseline and this report, output access control judgment (OVERALL/EXIT_CODE), strict mode tightens WARN exit code
- eval_judge_validateCalibrate LLM judge on the manual annotation set: report confusion matrix and TPR/TNR (look at them separately, agreement will be deceiving), only when the two indicators meet the standard are calibrated
Install command
dsh plugin --profile web add github:BiBoyang/dsh-eval-harness