Know your AI agents are still secure after every change
Humanbound runs scheduled adversarial tests against your agents in production, so the model updates, prompt edits, and new tool integrations that weaken security are caught before your users find them.
Regression detection
Every finding is tracked from the moment it is first detected, so you can see when it was fixed and know right away if it returns in a later release.
An attacker that learns
The engine keeps the strategies that broke through, retires the ones that fail, and builds new ones from how your agent behaves, which makes each cycle harder to pass than the one before.
Posture over time
A series of posture scores shows whether your agents are getting stronger, slipping, or stuck, and gives engineering and leadership the same view of progress.
Every finding has a history
Any change to the model, the system prompt, or the tools can bring back a vulnerability you already fixed, and model providers update their models on their own schedule. Even with no changes, the same input can produce different outputs.
Monitoring follows each finding through every test cycle, so you see when it first appeared, whether the fix held, and when it came back.
Testing over time also enables cross-session leakage detection: the engine plants identifiable tokens in one session and checks whether they appear in later ones. Single-session testing cannot detect this.
Open
A vulnerability that has been detected and is not yet resolved.
Fixed
A previously open finding that has not been reproduced in recent test cycles.
Regressed
A fixed finding that has been detected again, often after a code change, a model update, or configuration drift.
Stale
A finding that has not been triggered in 14 or more consecutive days of testing.
An attacker that knows your agent better with every cycle
Static scanners and traditional red teaming can only find the vulnerability classes they already know to look for. As your agent gains new tools and data sources, its attack surface changes in ways no pre-built test suite can anticipate.
- Learns from every response
- When a strategy fails, the engine works out why and moves to a different angle. When one partly succeeds, it layers further techniques onto the weakness it observed.
- Sharper every cycle
- Strategies that breached your defenses are kept, refined, and used again, and those that keep failing are retired. By the tenth cycle, the engine knows your agent's weak points, response patterns, and edge cases.
- Finds what changed
- When the engine finds a new tool, a data source it can reach, or a permission boundary that shifts under pressure, it forms its own hypotheses about what might be exploitable and tests them without human input.
See where your security is heading
A single posture score tells you where an agent stood on one day. A series of scores shows direction: a rising trend means your agents are hardening, a falling trend means something has regressed, and a flat line despite remediation work means fixes are not reaching production.
With monitoring, posture becomes a time series of score, grade, and change for every cycle, so a dip after a deployment sits right next to the fix that followed it. Engineering, compliance, and leadership each get the evidence they need, and frameworks such as the NIST AI RMF and the EU AI Act expect this kind of post-deployment monitoring.
Testing, monitoring, and protection in one loop
A single test produces a snapshot that starts going out of date as soon as the agent, the model behind it, or the attacks against it change. Monitoring sits between development-time testing and the runtime firewall, and each stage passes what it learns to the next, so your security grows stronger the longer your agent is monitored.
Testing sets the baseline
A one-time test finds vulnerabilities, generates guardrail rules, and trains a firewall classifier for your agent.
Monitoring sharpens the attacker
Each cycle builds on everything earlier cycles learned, adding refined strategies and newly discovered weaknesses.
The firewall absorbs the intelligence
Every cycle adds examples of what attacks and legitimate use look like for your specific agent, which improves classifier accuracy and makes guardrail rules more precise.
Production closes the loop
Firewall verdicts on real user traffic flow back into monitoring, where they guide which strategies to refine and what to test first.
Set it up once and it keeps running
A campaign is a recurring test cycle for one agent. Each cycle produces a posture score, findings with their lifecycle status, and coverage metrics. Tests run on Humanbound's infrastructure, so the CLI does not need to stay running.
Monitoring is a platform feature and requires a Humanbound account.
Schedule
Run tests daily, weekly, or whenever you deploy.
Scope
Choose which threat classes to test, how deep to go, and which orchestrator to use.
Intelligence
Each campaign carries forward its attack strategies, finding history, and coverage data from past cycles.
Alerting
Get webhook notifications when posture changes, new findings appear, or a fixed finding regresses.
What monitoring adds to a single test
A single hb test run gives you a posture score, findings, and guardrail rules for one moment in time. Continuous monitoring runs the same engine on a schedule and adds everything that only becomes possible across many test cycles.
hb test runKeep your agents secure as they change
Turn a one-time test into continuous assurance that gets sharper with every cycle.