Re-test your agent as it changes.

Models, prompts, tools, permissions, and workflows change. Continuous Agent Validation re-runs relevant security scenarios and regression tests after material changes, and compares the result with your previous assessment.

After an initial assessment · Per material change

An assessment describes the agent you shipped that week.

Every change to a model, prompt, tool set, or permission can reopen an action path that was closed, or open a new one. The evidence from a point-in-time assessment does not carry forward on its own.

A prompt edit changes how the agent treats untrusted content in tool results.
A new tool or MCP server adds an action path that was never in scope.
A permission or approval-threshold change alters what the agent may do without a human.
A model swap changes behaviour under the same adversarial scenarios.

What a validation pass covers

Each pass re-runs the scenarios that apply to your agent against the changed build and reports what moved.

Scenario re-run
The adversarial scenarios from your assessment, executed again against the new build.
Regression comparison
Per-scenario comparison against the previous assessment result: what held, what reopened, what is new.
Confirmed-finding re-test
Previously confirmed findings are re-executed to check whether the remediation still holds.
Capability changes
Tools discovered on the new build are compared with the confirmed capability map so scope changes are visible.

How a validation pass runs

The same pipeline as the initial assessment, run against the changed build and compared with the stored previous result.

01
Trigger
You tell us a material change shipped: model, prompt, tools, permissions, or workflow.
02
Re-discover
The agent’s tools are discovered again and compared with the confirmed capability map.
03
Re-test
Relevant adversarial scenarios and confirmed-finding regression tests run against the new build in the agreed test environment.
04
Compare and report
You receive an updated report with per-scenario regression status against the previous assessment.

Validation is not runtime protection.

Themis Agent Intelligence
Runtime authorization of live consequential actions: each proposed action is evaluated before it executes. Does not depend on an assessment.
Continuous Agent Validation
Adversarial regression testing when the agent changes: scenarios re-run against the new build, results compared with the previous assessment.

They are complementary. Validation tells you what a change did to your security boundaries; Agent Intelligence governs actions while the agent runs.

What you receive

Updated assessment report
Finding Report or Assessment Coverage Report for the new build, with the same evidence structure as the initial assessment.
Regression comparison
Per-scenario status against the previous result.
Updated capability map
The discovered tool surface of the new build and any differences from the confirmed baseline.

Continuous Agent Validation is available after an initial AI Agent Security Assessment. The assessment establishes the confirmed authority baseline and scenario set that validation re-runs.

Need runtime protection instead?

Themis Agent Intelligence evaluates consequential actions while the agent runs and returns ALLOW, ALERT, REQUIRE APPROVAL, or BLOCK before execution. It does not depend on an assessment.

About Themis Agent Intelligence

Why re-test

Same evidence, later date
A validation pass produces the same evidence structure as the assessment, so results are comparable across builds.
Anchored to your baseline
Re-tests run against the authority baseline your team confirmed, not a generic ruleset.
Scoped to what changed
A pass is triggered by a material change and reports against the previous result.

Questions we get asked

Do we need an assessment first?

Yes. The AI Agent Security Assessment establishes the confirmed authority baseline, the capability map, and the scenario set. Validation re-runs against those.

What counts as a material change?

A model swap, a system-prompt change, added or removed tools or MCP servers, a permission or approval-threshold change, or a workflow change that alters what the agent can do.

How is a pass triggered?

You tell us a change shipped and we schedule the pass. Pipeline-triggered runs are not part of the current offering.

How is it priced?

Pricing is agreed per engagement after the initial assessment. We do not publish a subscription price at this stage.

Ship often? Keep the evidence current.

Continuous Agent Validation follows an initial assessment. Start by checking whether your agent is supported.