Skip to content
Building

Engineering notes

promptslay

Unit tests for your prompts, with a CI gate that fails the build when behavior drifts.

Open source · MIT · TypeScript · started June 2026 · by Adeen Shukla

View source

The problem

Prompts are code with no compiler. Editing one word, swapping models or bumping an SDK can silently change behavior in production.

promptslay gives prompts the same safety net code has: versioned test suites, assertions, committed baselines, and a CI gate that turns "the bot feels worse" into a red build with a diff.

How a request flows

  1. 01SuiteTest cases in TypeScript or YAML, plus JSONL datasets
  2. 02RunnerParallel cases, retries on transient provider errors
  3. 03ProviderAnthropic by default, an offline mock, or your own adapter
  4. 04Gradersexact, contains, regex, JSON Schema, similarity, LLM judge
  5. 05BaselineClean runs report; drift prints a diff and exits 1

Decisions worth explaining

Cheap checks first, judges when needed
Deterministic graders cost nothing. The LLM judge scores a weighted rubric, can drop to a fast, cheap tier, and its spend is tracked per case.
Baselines explain why something drifted
Each baseline stores the output plus a hash of the rendered prompt and model, so a report can tell a nondeterministic model apart from a prompt you actually changed.
Baselines live in git
They're meant to be committed, so a prompt change and its behavior change show up in the same review.
Re-running an unchanged suite is free
Responses are cached on disk by provider, model, rendered prompt and params. Cost and latency are tracked for every case.
Works offline
The bundled mock provider is deterministic, so the harness itself can be tested in CI without an API key.
One interface per provider
Adding a model provider means implementing a single generate() method and registering it.