Citations Fact coverage Abstention Latency SLO Cost

You would not merge code without tests

Prompts get merged without them every week. This harness scores a fixed dataset on citation precision and recall, fact coverage, correct abstention, latency against an SLO, and estimated cost, then fails the build when the numbers drop. Same input, same score, every run.

Open the HTML report Read the source

Scored suite

The bundled dataset holds three cases: a grounded answer, a correct abstention, and one response that cites a source it never retrieved while also breaching the latency SLO. The third one is supposed to fail, and the score says exactly why.