We put Jev at the front door of a security platform. It costs five cents a day.

A cybersecurity client needed to catch misuse at onboarding and weak pentest reports before customers read them. We swapped a domain blocklist and an LLM-as-a-judge for Jev. Here's what we asked it, how we escalate, and what it costs.

Jev in Production: Intent Guardrails and Report Quality Gates

Image: METAHEURISTIC

Alexander Myasoedov

+Alexander Myasoedov Alexander writes about the operational side of shipping production AI - agents, retrieval, evals, and the guardrails that keep them from going sideways.

Someone signs up. Company name looks normal, email domain is real, card goes through. Their first job is a login URL plus a few lines of free text, something like “log in with these credentials, go through checkout and the admin panel, test everything you can.” Fine, except the URL belongs to their direct competitor. Every field in the form validates. And the platform is a few seconds away from running an automated security assessment, with real credentials, against a company that never asked for one.

We shipped the fix for this last week, for one of our clients. I want to write it down while it’s fresh, because the thing that fixed it was cheaper and dumber than I expected: Jev, TypeSafe’s System One model, answering a handful of typed questions at two points in the pipeline.

The client: security assessments driven by free text

The client is a cybersecurity firm. Their customers hand them a website URL and whatever instructions they feel like writing, and the firm runs an assessment. That flow gets long fast. It logs in, walks a bunch of pages, and runs scenarios nobody hardcoded. Some of those scenarios are rulesets the firm maintains. Some are fully agentic, with a model deciding the next step based on what’s on the page.

The other end is just as loose. Findings come back from the firm’s own researchers and from third-party pen testers, and a lot of them are free text too. So you’ve got unstructured input going in, unstructured reports coming out, and real liability in the middle. Jev ended up with three jobs there.

Gate 1
Intent of use
Before a scan: is this client assessing something they're entitled to assess?
Gate 2
Report quality
After a scan: can the customer actually act on this report?
Policy
Escalation
Stop for a human, flag and keep going, or pass. Decided in code.

Why a domain blocklist fails for intent

When we went through a few weeks of signups, the bad patterns were easy to spot by eye. People scanning competitors. People submitting a pile of sites that had nothing to do with them, half of them dead or parked. Government sites. Stuff from the grey corners of the internet. The obvious move is a list: block .gov, pull in a feed of known-bad domains, throttle bulk submissions. We sketched that and threw it away pretty quickly. It catches the boring cases, and the boring cases were never the problem.

The problem is that a URL isn’t good or bad on its own. A competitor’s storefront is a completely legitimate target if you own it, and abuse if you don’t. You can’t put that in a list. You have to look at who the client says they are, what’s on the target site, and what the instructions ask for, and decide whether those three things belong together. That’s a judgement call with a small, closed set of answers, and it has to run on every signup and every new target. Which is pretty much what Jev is for.

Gate 1: a Jev intent guardrail at onboarding

First we strip the data. Profiles and instructions go through PII and PHI redaction before anything leaves the client’s infrastructure: names, emails, phone numbers, anything that smells like health data. Jev doesn’t need to know who a person is to tell whether the request makes sense. It would also be a bit ironic to ship personal data to a third party in order to decide whether to protect someone. What’s left is the company profile, the public text of the target page, and the redacted instructions.

Then it’s one request with several questions. Jev is on OpenRouter as typesafe/jev-1.13 behind POST /api/v1/systemone, so we didn’t need a separate TypeSafe account. All questions see the same state and run in parallel, so we ask everything the policy could possibly need in one go:

import httpx

URL = "https://openrouter.ai/api/v1/systemone"


def check_intent(key: str, profile: dict, target_text: str, instructions: str) -> dict:
    state = {"client": profile, "target_page": target_text, "instructions": instructions}
    questions = {
        "relationship": {
            "type": "choice",
            "instructions": "The most plausible relationship between `client` and the site in `target_page`.",
            "criteria": {
                "owner": "The client operates this site.",
                "engaged": "The client is plausibly engaged to assess it for the owner.",
                "competitor": "The client competes with the site's operator.",
                "unrelated": "No plausible connection between the client and the site.",
            },
        },
        "government": {"type": "noul", "instructions": "`target_page` is a government or public-sector site."},
        "grey": {"type": "noul", "instructions": "`target_page` offers illicit or grey-market goods or services."},
        "dead": {"type": "noul", "instructions": "`target_page` is parked, abandoned or a placeholder."},
        "in_scope": {"type": "noul", "instructions": "`instructions` describe a security assessment of this site."},
    }
    res = httpx.post(
        URL,
        headers={"Authorization": f"Bearer {key}"},
        json={"model": "typesafe/jev-1.13", "state": state, "questions": questions},
        timeout=10,
    )
    res.raise_for_status()
    return res.json()["answers"]

There’s no “is this client malicious?” question in there, and that’s on purpose. Motive is exactly the kind of thing a model ends up guessing at, and you can’t check the guess. Every question that survived is something a person could check by opening the page. The verdict gets assembled from those answers in code, same idea as in the Jev explainer: split the judgement into small literal questions and keep the math somewhere you can review it.

What replaced the blocklist: A question about the relationship between the client and the target. The domain alone never decides anything.

Gate 2: report quality without LLM-as-a-judge

The second gate is on the reports. Some are great. Some are a scanner dump with a summary paragraph pasted on top. Some rate a finding critical and then describe an impact that’s clearly low. Before we got involved, the check was an LLM-as-a-judge: send the report and a rubric to a frontier model, get a verdict back. Two problems. It was expensive to run on every single report, and it wasn’t conclusive. You’d get a paragraph and a number, the number was whatever the paragraph talked itself into, and you couldn’t really set a threshold on that.

With Jev we use a Score against levels written in plain English, plus a couple of Nouls for failure modes we kept running into. A Score gives you a continuous position between the levels and a probability for each one, so “somewhere between thin and actionable” turns into an actual number. The levels below aren’t the client’s real rubric, but they’re close in shape:

questions = {
    "quality": {
        "type": "score",
        "instructions": "How actionable `report` is for the customer who receives it.",
        "criteria": [
            "unusable: no finding a reader could reproduce",
            "thin: findings without repro steps, affected asset or evidence",
            "actionable: every finding has asset, repro steps and evidence",
        ],
    },
    "scanner_dump": {"type": "noul", "instructions": "`report` is mostly raw tool output with no analysis."},
    "severity_mismatch": {"type": "noul", "instructions": "A finding's stated severity does not match the impact it describes."},
}

A call comes back in well under a second and costs a fraction of a cent, so this runs when the report is submitted instead of in some nightly batch. The tester hears about a weak report while they still remember what they tested. One gotcha: long reports have to be split. OpenRouter’s guide lists a 32k-token context for Jev, and TypeSafe’s docs say accuracy drops when the state is full of stuff the question doesn’t need. So we grade each finding separately and roll the results up ourselves.

LLM-as-a-judge

  • A paragraph plus a number it wrote itself
  • Frontier-model price on every report
  • Hard to threshold, hard to reproduce

Jev Score

  • A position on levels you define, with probabilities
  • Priced per input token, output free
  • A float you compare against a constant

Three escalation tiers: stop, flag, pass

Both gates feed one policy with three outcomes. Stop means the flow halts and nothing happens until a human reviews it. Competitor targets and government sites end up there. Flag means we raise a marker so someone takes a look, but the flow keeps going. That’s where borderline reports and half-dead sites go. Pass means nothing unusual, carry on. Most traffic passes, which is what you want. People only get pulled in when there’s actually something to look at.

Which tier a request lands in depends on two things Jev returns: the probabilities, and the confidence on each Choice or Score. High probability of misuse, stop. Middling probability, flag. And if Jev can’t make up its mind, with the probability smeared across options, that’s a flag too. I’d rather have someone glance at it for ten seconds than let a coin flip through. The thresholds here are placeholders, not what runs in production:

def tier(a: dict) -> str:
    rel = a["relationship"]
    misuse = rel["probabilities"]["competitor"] + rel["probabilities"]["unrelated"]
    if a["government"]["noul"] > 0.7 or a["grey"]["noul"] > 0.7 or misuse > 0.8:
        return "stop"
    if misuse > 0.4 or a["dead"]["noul"] > 0.5 or rel["confidence"] < 0.6:
        return "flag"
    return "pass"
Figure - Probability picks the tier; low confidence never passes
P(misuse)
confidence
Drag P(misuse) across the bar and watch the verdict move from pass to flag to stop. Then drop confidence below 0.6 with misuse low: the request can no longer pass, and lands in flag for a human glance.

That’s the whole policy. If the client’s risk team wants to be stricter about bulk submissions, it’s a one-line change to a float, with a test next to it. Nobody has to reword a prompt and then wonder what else the new wording quietly changed.

Low confidence is a tier, not a failure: When the model effectively says I can’t tell, that’s useful. Send it to a human instead of rounding it to yes or no.

What Jev costs in production: five cents a day

Now the fun part. Jev charges $0.042 per million input tokens, and output is free since it doesn’t generate any. With both gates running at this client’s volume, the OpenRouter bill comes to about five cents a day. At that price, five cents is roughly 1.2 million input tokens. And it’s not an estimate: OpenRouter puts a usage.cost field on every response, so we just add them up.

What surprised me more is how that changes which checks you bother to write. With the LLM judge, every gate was a line item someone had to defend, so there were few of them and they ran late in the flow. Now the expensive part of a check is the person it might page, not the model call. We ended up adding questions, because a question that only matters on one branch costs you the tokens in its instructions and nothing else.

Where Jev does not fit

Jev doesn’t explain itself. TypeSafe and OpenRouter both say it up front: no reasoning traces, no free text. For a gate that’s mostly okay, since the questions carry the explanation. A reviewer opening a flagged signup sees that relationship came back as competitor with 0.83 and gets the picture. But if the client ever needs to send a written reason to the person who got stopped, that text has to come from somewhere else.

The bigger thing to keep in mind is that the target page is untrusted. Anyone can put anything on their own website, including text written to push a classifier around. TypeSafe’s own jaggedness notes say Jev doesn’t treat state as hostile by default. So in this setup it’s one signal, never the only thing standing between a scanner and a site it shouldn’t touch. The stop tier always ends with a person, and the hard legal lines are plain deterministic code. It’s also strongest in English, which matters for non-English targets and is part of why the confidence check is in the policy.

Jev guardrails FAQ

What is Jev used for in production?

Typed decisions inside software: classification, routing, risk gating and quality grading. Here it runs two gates for a cybersecurity firm, one checking a new client’s intent of use before a scan and one grading pen test reports as they come in.

Can Jev replace LLM-as-a-judge?

For verdicts with a closed set of answers, yes. A Jev Score returns a position on levels you define plus probabilities, and you threshold that in code. It can’t write a critique or explain a verdict, so keep an LLM around for anything that needs prose.

How do you call Jev through OpenRouter?

Use an OpenRouter API key, the model id typesafe/jev-1.13 (or the ~typesafe/jev-latest alias), and POST https://openrouter.ai/api/v1/systemone with a state and a map of questions. You don’t need a separate TypeSafe account.

How much does Jev cost per day?

Input is $0.042 per million tokens and output is free. For this client, both gates together cost about five cents a day, roughly 1.2 million input tokens.

Is Jev safe to use as a security guardrail?

As one signal, with a human behind the stop tier. Jev doesn’t treat its input as hostile by default, so content on an untrusted website can influence an answer. Pair it with deterministic rules for the hard lines and send low-confidence answers to a person.

If you’re putting typed decisions into an onboarding flow or a review pipeline and want to compare notes, get in touch.

Follow in Google

Make Metaheuristic a preferred source.

One tap and posts like this one surface higher in your Top Stories.

Work with us

Production AI, with guardrails.

Start with a fixed-scope AI Workflow Audit. We map the opportunity and quote a build.

Start a Discovery Sprint →