Agent evaluation platform for DeepSeek Harness
hccccc01333/dsh-eval · skills
Agent evaluation platform for DeepSeek Harness: benchmark YAML, headless dsh orchestration, trace-based metrics, LLM judge, paired A/B, keyless replay, and cross-harness import.
★ 2GitHub stars
0Forks
2026-08-25Last updated
TypeScriptLanguage
MITLicense
Key features
- Agent Evaluation Platform for deepseek-harness.
- dsh eval run benchmark.yaml — orchestrate one headless dsh subprocess per case × trial
- Trace harvesting from persisted session logs (everything a model sees is reconstructable from the log)
- Automatic metricstask success, tool success, tool-selection accuracy, steps, tokens, latency, cost, retry, invalid tool calls, context usage
- Scripted gradingexpected.tool (tool-selection accuracy) and expected.check (task success)
Install command
dsh plugin --profile web add github:hccccc01333/dsh-eval