Tensic guides

Guide 11

Run evals and manage prompt versions

Test a project against a dataset of questions and expected answers, compare runs across prompt versions, and restore an earlier prompt when quality drops.

On this page
  1. Create an eval dataset and test cases
  2. Run an evaluation and choose metrics
  3. Read eval results and compare runs
  4. View prompt history and restore an earlier version
  5. Catch and roll back a prompt regression
  6. Run evals automatically and get notified

Evals measure the quality of a project's answers. You build a dataset of test questions (optionally with the answers you expect), run it against the project, and a judge model scores each answer. Every run records the prompt version it ran against, so when you change a system prompt you can see whether quality went up or down, and restore the previous version if it went down.

Create an eval dataset and test cases

A dataset is a collection of test cases evaluated together. A project can have several datasets, for example one per topic or scenario.

  1. Open the project and click Evals in the project header.
  2. Click New Dataset, enter a Name and an optional Description, and click Create.
  3. Click Add First Test Case (later: add further test cases to the dataset).
  4. Enter the Question, the input sent to the project, for example "team.blue is a bad company" for a guard project.
  5. Enter the Expected Answer (optional), for example BLOCK. An expected answer is required for the Correctness metric.
  6. Click Add. Repeat for every case you want to cover.

Make the dataset reflect real use. For a support assistant, add the questions customers actually ask, with the answers you want the assistant to give. A useful dataset contains hundreds of cases specific to your use case. This dataset becomes your baseline: every future prompt change is measured against it.

Run an evaluation and choose metrics

  1. On the Evals page, click the run (play) button next to the dataset.
  2. In Run Evaluation, select the metrics. Each test case is scored per metric:
    • Answer Relevancy: is the answer relevant to the question?
    • Correctness: does the answer match the expected answer? Needs an expected answer on the test cases.
    • Faithfulness: available only for RAG projects.
  3. Check the judge model shown as Judged by … at the top of the dialog. The judge can be the same model as the project, or a different one. A smarter judge gives stricter grading; a cheaper judge is enough for simple checks. Choose based on the use case.
  4. Click Start. The run is a background job and appears under Evaluation runs with the prompt version, for example v2, and a status of DONE when finished.
Run Evaluation dialog with the judge model and the Answer Relevancy and Correctness metrics

Read eval results and compare runs

  1. Click a run under Evaluation runs to open it in Compare. You can compare up to three runs side by side.
  2. Each run shows an overall score per metric, for example Correctness 90%.
  3. Click Show per-case results to see, for each test case, the question, the expected answer, the actual answer and the judge's score with the judge's written reason.
  4. Switch to Timeline to see scores over time, with every run listed by date, prompt version and score.
Eval run with a 90% correctness score and the per-case result for one test case

A score below 100% is not always a failure. In the example, the guard answered BLOCK plus a short reason line. The judge scored it 90% because the extra reason was not in the expected answer, even though the verdict was correct. Read the judge's reason before changing anything.

Evals measure answer quality. They do not report token usage or cost; use the project's usage and budget views for that.

View prompt history and restore an earlier version

Every saved change to a project's prompt creates a new version. The versions are stored by Tensic, not in files.

  1. Open the project's Prompt tab and scroll to Prompt history. Each row shows the time, the change, the version (v1, v2, v3…), the author, the size change, and which version is ACTIVE.
  2. Click restore on the version you want back.
  3. The Restore version dialog shows a diff between the current prompt and the version you selected, with removed lines in red and added lines in green. Review it.
  4. Click Restore vN. Restoring records a new version, and the current prompt stays in the history, so nothing is lost. No separate save is needed.
Restore version dialog showing the diff between two prompt versions

Catch and roll back a prompt regression

Use evals and prompt history together whenever you change a system prompt:

  1. Run the eval on the current prompt to get a baseline score, for example 90% on v2.
  2. Edit the prompt and click Save. It becomes v3.
  3. Run the same eval again, without changing the dataset.
  4. Open Timeline. If the score dropped (in the example, from 90% on v2 to 10% on v3), open the failing test cases and read the judge's reasons to see what the change broke.
  5. Go to Prompt > Prompt history, click restore on v2, check the diff, and click Restore v2.
  6. Run the eval again to confirm the score is back to the baseline.

The version label on each run shows exactly which prompt produced which score.

Run evals automatically and get notified

Evals can also be started through the API, for example from a deployment pipeline after a prompt change. See the API reference (Swagger in the left menu) for the endpoints.

To react to results automatically, subscribe to eval events:

  1. Open the project's Integrations tab and go to Event webhooks.
  2. Enter the endpoint URL (HTTPS only) and, optionally, a Signing secret. The body is signed with HMAC-SHA256 and sent in the X-Tensic-Signature header.
  3. Under Subscriptions, turn on eval_completed. It fires when an eval run finishes, with the score and a per-metric breakdown in the payload.
  4. In your own system, compare the score with your baseline and alert your team, or block a release, when it drops.

Common questions

Do I need an expected answer for every test case?

No. It is optional, but the Correctness metric needs one. Answer Relevancy works without it.

Which judge model should I use?

It depends on the use case. The same model is often fine. Use a stronger model when grading needs nuance, and a cheaper one for simple checks such as ALLOW/BLOCK verdicts.

Why did a correct answer score 90% instead of 100%?

The judge also considers anything extra in the answer. Open the per-case results and read the judge's reason.

Can I see which prompt a run used?

Yes. Every run shows its prompt version (for example v2), in the run list and in Timeline.

Does restoring an old prompt delete the newer versions?

No. Restoring creates a new version, and all previous versions stay in the prompt history.

Do evals show token cost?

No. Evals measure quality only.

Can I run evals on a guard project?

Yes. Evals work on any project. For a guard, use messages as questions and ALLOW or BLOCK as expected answers.