Humanbound website

When Opus refused, Hugging Face switched models. Here is how to pick any model for Humanbound Red-Teaming

Hugging Face dropped Opus when its guardrails blocked incident work. hb 2.13 lets you red-team with any OpenAI-compatible model. Five OpenRouter setups compared on empty replies, cost and speed, and why you should pin the model and the host.

9 min read
Humanbound banner: "Bring your own model. Pin the host." A terminal shows hb pointed at DeepSeek V4.1 Flash via an OpenRouter preset, with 0 of 672 empty scorer replies at $0.14 engine cost.

In July, an AI agent working its way out of an OpenAI evaluation sandbox ran a 4.5-day campaign, about two and a half days of it inside Hugging Face's infrastructure. When Hugging Face published its technical timeline on July 27, one paragraph had little to do with the attack itself. It was about the tools the defenders reached for:

"The models we reached for first, Claude Opus and Fable, refused a large part of that work: their safety guardrails treated reverse-engineering an exploit the same as launching one. Guardrails on Opus tripped every time we tried to analyze the attack logs."

So they switched. They ran an open-weights model, GLM-5.2, on their own hardware and "rerouted the entire pipeline through it, with the added benefit of keeping the attacker data on-prem." That was incident investigation, meaning log analysis and decoding the attacker's payloads, not attack generation, so it says nothing directly about red-teaming models.

The questions it raises are still the ones you face when you point an AI red-teamer at your own agent:

  • Will the model do the work?
  • What does it cost?
  • Where does your attack data go?

Humanbound's hb supports a short list of providers directly and as of version 2.13 it also works with any service that speaks the OpenAI API compatible format, including a well renowned service like OpenRouter, which sells access to models from many labs with a single API key. I wrote that change (PRs #166 and #170), so I wanted to see what happens when you use it for real.

One model plays three jobs

When you run a hb test, the selected AI models do three things. It writes the attacks. After every turn it also gives a quick 0 to 10 score for how close the attacker is to its goal, and that score steers the next message. At the end it judges each conversation and decides whether your agent broke and did something it was not supposed to do.

Setup on hb 2.13

You need hb 2.13 or later.

Loading code editor...

HB_MODEL takes the service's own model name, so on OpenRouter it is deepseek/deepseek-v4.1-flash, not an OpenAI name. You can find the model slug on each model's page at openrouter.ai/models.

Screenshot of the DeepSeek V4.1 Flash page on OpenRouter, with the model slug deepseek/deepseek-v4.1-flash highlighted under the title. Summary cards show text and image input with text output, prices from $0.015 input and $1.20 output per million tokens, a 1M-token context window and a release date of September 10, 2026.

You also need something to test. hb arena ships a practice target called hello-world, a shop assistant built to be easy to break. It runs in Docker, so start Docker first. Its own model is set separately, so I pointed it at a small Llama model on OpenRouter too:

Loading code editor...

Use a separate OpenRouter key with a spend limit for this. Anything in the same shell that reads OPENAI_API_KEY without a base URL will send it to OpenAI.

I ran my long tests from the main branch on September 30 and October 2, which already contained both changes. I also started a run on the released 2.13.0 package with the same settings and watched engine calls reach OpenRouter. I stopped that run early, so every number below comes from the main-branch runs.

One model name, 32 endpoints

One model name on OpenRouter can map to many different servers. When I looked, DeepSeek V4.1 Flash was available from 32 endpoints run by 29 companies, charging from $0.015 to $0.60 per million input tokens, some running a compressed version of the model and some not. If you do not choose, OpenRouter picks for you, and it can pick differently from one call to the next.

Screenshot of the Providers tab for DeepSeek V4.1 Flash on OpenRouter, listing the companies that host the same model with their input, output and cache prices, latency, throughput and uptime. Morph is third in the list at $0.021 input and $0.383 output per million tokens, with 1.12s latency, 36 tokens per second and 99.72% uptime. Above it are Open Inference ($0.015 / $1.20) and Relace ($0.02 / $0.60). Wafer, further down, has the lowest latency of the hosts shown at 0.50s, with 103 tokens per second.

Hosts also behave differently from each other. My first GLM run used the plain model name. Of 162 calls, 77 came back with no text at all. The model had spent its whole token allowance thinking and had nothing left to say, and I still paid for those calls. About $0.34 of the $0.50 I spent went to empty answers.

What fixed most of it was a preset: a saved set of routing rules that you refer to like a model name. You can build one in the OpenRouter dashboard, or with a single call that creates the preset under that slug (if the slug already exists, the call adds a new version and makes it the active one):

Loading code editor...

Diagram of how hb routes model calls. Inside the hb engine box, the Attacker, Scorer and Judge all send their calls to OpenRouter (HB_ENDPOINT); a note says one HB_MODEL answers all three jobs. The Attacker also exchanges attacks and replies with the practice target, hb arena hello-world. OpenRouter passes the model name @preset/hb-deepseek to the preset hb-deepseek, which is set to only Morph, no fallbacks and thinking off. The preset routes every call to Morph, the pinned host, and the 31 other endpoints are never called.

Three settings do the work. only pins the host, and allow_fallbacks: false makes sure OpenRouter never routes anywhere else, even when the pinned host is busy. reasoning turns the model's thinking step off, so a short answer limit does not get eaten before the answer starts.

I chose Morph provider because it was the cheapest. To work that out properly I measured what hb actually sends: about 13 input tokens for every output token. Weighted by that mix, Morph came out cheapest at $0.021 in and $0.383 out per million tokens, about a quarter cheaper than Relace in second place. Morph serves an 8-bit (fp8) build, so check that is acceptable before you pin it. Prices on OpenRouter move, so check them the day you run it.

I picked DeepSeek V4.1 Flash by name, not by the ~deepseek/deepseek-flash-latest shortcut. That shortcut is an alias that always points to the newest Flash model. When I sent it a test call, the reply reported V4.1 Flash, served by Makora, since nothing pinned the host. Today the alias points at V4.1 Flash. It will not forever, and a moving alias is fine for a chatbot. For a security test you want to repeat next month, pin the model and the host.

What happened on five setups

Same target and same command each time, 97 conversations per full run, run between September 30 and October 2. The scorer is the 0 to 10 progress score from earlier. hb gives it 50 tokens to answer, and that limit is where the setups differ.

Loading code editor...

The attacker and the judge never returned an empty answer in any of the three full runs. The scorer is where the models differ. GLM did wrap two of its verdicts in a code fence that hb could not read, but nothing came back blank. When a scorer reply is empty, hb falls back to a 5, and the attacker's next prompt tells it that it scored 5 out of 10 and is making some progress. That score is made up.

Sequence diagram of how an empty scorer reply becomes a fake score. The hb attacker sends an attack message to the practice target and gets a reply. It then asks the model, through OpenRouter, to rate the attempt from 0 to 10, sending only the first 200 characters of the reply with a 50-token limit. A model that thinks first uses up all 50 tokens and sends back an empty reply. hb finds no number and falls back to 5, and the next attack prompt says "Your previous attempt scored 5/10". A note says the 5 is not a real score, but it still steers the attacker on the next turn.

OpenAI's own small model lost between two-thirds and four-fifths of its scorer calls, so this is not only a problem for non-OpenAI models. A cheap model with thinking switched off answered every time, at about one-twelfth of the cost of the GLM run.

The cheapest setup was also the slowest. The median attacker call took 9.7 seconds on Morph and 2.7 on Wafer, though that compares two models as well as two hosts. Slow calls are likely a big part of why the DeepSeek run took 36 minutes. If you are waiting on a CI job, you may happily pay more for a faster host.

The rest errored: 8 of the 10 across all runs were the practice target timing out (more on that below), and 2 were GLM verdicts hb could not parse, which hb counts against the grade.

What this does not tell you

I used one practice target and measured only the plumbing: did each model return something usable for each of the three jobs, and what did it cost? I did not measure how many flaws a model finds or whether a judge's verdicts are right, so nothing here ranks models for attack quality. Each setup ran once. None of the empty replies were refusals, but a quick scan found the attacker stepping out of role in a few conversations, which my empty-reply count does not catch.

The 50-token limit on the scorer is the main reason models that think first lose so many scorer replies: every empty reply hit that limit. The scorer also sees only the first 200 characters of each reply. Both are weaknesses in how hb steers its attacker, and I would rather say so here than hide them behind a good-looking row.

The runs also hit the arena's 120 second timeout on a few conversations, because the target is a small model running through a shared service. hb tells you when this happens and leaves those conversations out of the grade: "2 conversation(s) errored and are left out of the posture grade and --fail-on, so this result may look better than it is."

Finally, your attack transcripts leave your machine. They go to OpenRouter and to whichever host you pinned, and on a small provider like Morph that is a company you may know very little about. Hugging Face counted keeping "the attacker data on-prem" as a benefit of running GLM itself. If your agent handles real customer data, read the host's data policy before you pin it, or use Ollama, which keeps everything local.

Things to try next

If you try any of these, I would like to hear what you find.

  • Count the empty replies. hb does not print this today. I wrapped the engine's HTTP call in a short script that logs each call's role, token counts and cost. Do the same before you trust a score from any model you have not checked.
  • Point it at your own agent. The practice target only exists to test the plumbing, and your agent is where model differences will start to matter.
  • Pin two hosts and compare. The same model name on two hosts gave me very different speeds and cost.
  • Try self-hosting. The docs (I wrote that page in #166) say LiteLLM and vLLM work through the same setting. I have not tried either.
  • Look at OpenRouter's data policy options. It documents a data_collection setting for restricting routing to hosts that do not store prompts. I have not tested it.

A full scan on the cheapest setup cost me $0.14 in model fees, or $0.19 on my OpenRouter account once the target's calls are counted. Check the empty replies first, then the bill.

Red-team your own agent for free. Start on the Community plan (Free forever) at app.humanbound.ai, or run it locally with pip install humanbound[engine]. The code is open source on GitHub.