Time: ~30 min
Say out loud (or voice-note to yourself)
- Golden task definition + one example
- How suite differs from unit tests
- How suite differs from asset time/token A/B
- What you measure on a tool-calling agent (tools + side effects + refuse)
- What you’d add next for production (LLM judge carefulness, traces, CI gate, larger suite)
Proof into your skill notes
Update your personal skill notes for the product-eval axis:
- Score: after Days 1–5 with lab proof, mark real-but-shaky practice; full confidence only after you’ve run the same harness on a real LLM agent once
- Last proof: path to lab + date + pass rates
- Pain still open: LLM-backed suite + CI
Optional stretch (later week)
- Swap fake agent for one real LLM tool loop
- Add a CI job that runs the suite on the qualifying project
Pack complete when the pack index “Done when” checkboxes are ticked.