Building
Engineering notes
promptslay
Unit tests for your prompts, with a CI gate that fails the build when behavior drifts.
Open source · MIT · TypeScript · started June 2026 · by Adeen Shukla
View sourceThe problem
Prompts are code with no compiler. Editing one word, swapping models or bumping an SDK can silently change behavior in production.
promptslay gives prompts the same safety net code has: versioned test suites, assertions, committed baselines, and a CI gate that turns "the bot feels worse" into a red build with a diff.
How a request flows
- 01SuiteTest cases in TypeScript or YAML, plus JSONL datasets
- 02RunnerParallel cases, retries on transient provider errors
- 03ProviderAnthropic by default, an offline mock, or your own adapter
- 04Gradersexact, contains, regex, JSON Schema, similarity, LLM judge
- 05BaselineClean runs report; drift prints a diff and exits 1
Decisions worth explaining
- Cheap checks first, judges when needed
- Deterministic graders cost nothing. The LLM judge scores a weighted rubric, can drop to a fast, cheap tier, and its spend is tracked per case.
- Baselines explain why something drifted
- Each baseline stores the output plus a hash of the rendered prompt and model, so a report can tell a nondeterministic model apart from a prompt you actually changed.
- Baselines live in git
- They're meant to be committed, so a prompt change and its behavior change show up in the same review.
- Re-running an unchanged suite is free
- Responses are cached on disk by provider, model, rendered prompt and params. Cost and latency are tracked for every case.
- Works offline
- The bundled mock provider is deterministic, so the harness itself can be tested in CI without an API key.
- One interface per provider
- Adding a model provider means implementing a single generate() method and registering it.