How do you use Jev for tool calling? You can use it to decide whether to call a tool, which one, and whether the call is risky. Jev doesn’t write free-text arguments for that call, though. It can pick an argument from a fixed list, like a cancellation reason, and the rest you build yourself in code, with predefined steps, and then validate. And you give the decision more outcomes than yes and no, so “not enough information” has somewhere to go. That’s the short version. The rest of this post is how I got there, with the Go code to show it.
Is Jev good for tool calling?
What I wanted to work out is how to tell the use cases apart and apply Jev properly in production. In one of our use cases we used Jev to filter content in a pipeline that was customer-facing and interacted with customers. The data was sanitized and the volume was fairly low, and the job was mostly classifying free-text intents. That part works, and I’ve written about what Jev is and where else we’ve put it already.
Classifying intents is one thing. Tool calling is a much different beast, I think. Let me make up an example. Say a customer writes in with some intent, and you want that intent to end in a tool call. You can classify with Jev whether or not you should call a certain tool. The problem comes when you need to know what parameters to call it with. You can extract a customer ID and some parameters from the message, and say your tool function takes five of them: customer ID, email address, a cancellation reason, a free-text note, and one or two others. In this post I’ll use Jev, System One model and decision model interchangeably. There are plenty of use cases for them. The question is how you actually do the tool-calling part.
How do you know the arguments are correct?
Say we have an intent, and from it we determine that we’d like to call a certain function. How do we know the arguments we extracted are correct? With generation, an LLM can produce them. Give it the free text and it will generate the function call for you, and you can check that it’s a valid signature. Jev doesn’t generate anything. It answers typed questions with probabilities. So with Jev you construct the arguments yourself, with predefined steps. I’d say that’s also the more reliable way to do it: first extract the parameters with rules, or with some specific pieces of logic laid out, and only then call the tool.
That isn’t the only way to split it. Evalgent’s guide to Jev in voice agents makes the opposite split its default: an LLM proposes the tool call and its arguments, and Jev validates it before your code executes, confirms or rejects. Their Jev-first variant matches what follows here, where Jev picks the tool and the arguments that come from a fixed list, and only the free-form fields need something else.
In Go that’s plain code. Here’s a toy version for a cancel request:
var (
customerIDRe = regexp.MustCompile(`\b\d{5,10}\b`)
emailRe = regexp.MustCompile(`[^\s@]+@[^\s@]+\.[a-z]{2,}`)
)
type CancelArgs struct {
CustomerID string
Email string
Reason string // one of a closed set, picked by Jev
Note string // the customer's own words, verbatim, for whoever reviews it
}
func extractCancel(msg, reason string) (CancelArgs, error) {
args := CancelArgs{
CustomerID: customerIDRe.FindString(msg),
Email: emailRe.FindString(msg),
Reason: reason,
Note: msg,
}
if args.CustomerID == "" {
return args, errors.New("customer_id missing")
}
if args.Reason == "" || args.Reason == "unclear" {
return args, errors.New("reason unclear")
}
return args, nil
}
Reason is the one argument that needs judgement, so it isn’t free text. It’s a Choice
over a fixed list, answered in the same Jev request as everything else, and code only
checks that it landed on something usable. Note is never interpreted at all. It goes to
the human who approves the cancellation.
None of that touches the model. That’s the point of doing it this way: the part that has to be exactly right is a regex and a check you can unit test. But it still leaves the question of how you verify the call as a whole is correct before anything runs.
Which tool, and is it risky?
One option is to make the tool itself a choice. If you have multiple tools, which one should you call? Then you end up with more questions. If I choose this tool, is it the correct choice? Are the arguments correct? Is it a risky action? The risky action can be a boolean on each of these calls.
Take a tool like cancel_subscription. It’s a risky action and it needs some human
supervision. It might take five parameters to execute. How do you know all of them are
correct without supervision? Those are the things you have to address with proper
instrumentation in the harness. A fast classifier doesn’t remove the harness. It still
requires one.
With Jev, the tool choice, the risk flag and the cancellation reason are questions in the
same request. A Choice takes up to 255 options. If you have more tools than that, the
Evalgent guide suggests two stages: score each candidate, then run a Choice over the top
few plus a none option. I’m
reusing the Question, Choice and Noul helpers from
the Jev explainer, and pointing the client at OpenRouter:
func decide(ctx context.Context, c *jev.Client, msg string) (*jev.Response, error) {
return c.SystemOne(ctx,
map[string]any{"message": msg},
map[string]jev.Question{
"tool": jev.Choice("Which tool the request in `message` needs", map[string]string{
"cancel_subscription": "The customer wants to stop their subscription.",
"update_email": "The customer wants to change the email on their account.",
"none": "The request needs none of these tools.",
}),
"risky": jev.Noul("Acting on `message` changes or removes something the customer cannot easily undo"),
"reason": jev.Choice("Why the customer in `message` wants to cancel", map[string]string{
"too_expensive": "Price or billing is the stated reason.",
"not_using": "They no longer use the product.",
"switching": "They are moving to another product.",
"unclear": "No reason is stated or it can't be told apart.",
}),
})
}
Do you need LangChain for Jev?
There are plenty of examples of how to use Jev through LangChain. In my experience there is no fucking reason to use LangChain at all. Our application was in Go, and we didn’t care about the LangChain implementation either way. It’s a very simple API call to Jev. Constructing the correct types and answers inside the codebase is a small SDK for the decision, and we even called it the decision SDK. That was much more transparent than pulling in a grenade of abstraction from LangChain.
We used OpenRouter as the router for Jev, and it was fine for us. We didn’t need anything else. The SDK stayed very slim and very traceable, and we liked the cost and latency budget for it because it’s low. TypeSafe publishes 70ms to 500ms end-to-end and says those evals were generally run from laptops on the West Coast, where the service is based, so measure from your own region before you plan around it. The rate limits on direct access currently read 100K tokens and 80 requests per second, and TypeSafe warns they change without notice. The whole transport is this:
c := &jev.Client{
URL: "https://openrouter.ai/api/v1/systemone",
Model: "typesafe/jev-1.13",
Key: os.Getenv("OPENROUTER_API_KEY"),
HTTP: &http.Client{Timeout: time.Second},
}
Decision SDK A client struct, a handful of typed questions and a switch statement. Slim, traceable, and every line between the customer’s message and the tool call is yours to read.
One Jev call or several?
So tool calling with Jev isn’t trivial. It’s still possible, and it’s fast, but you might end up calling Jev several times to cover all the criteria. I’m not saying multiple calls are necessary. You can do it in one shot, in one API call. Calling it several times is still common and acceptable, though.
You can squash those calls into one request to Jev: is this tool needed, is this the right tool to call, is it risky. You construct the schemas beforehand, and that probably optimizes your pipeline cycle. Every question sees the same state and runs in parallel, so the extra questions barely cost anything.
Then you have a set of if/else conditions on the decision: whether to call the tool, whether there’s some other action to take, whether you need to document it, whether you need to do something else. That logic is Go, next to the argument checks:
The one-second timeout comes from the same Evalgent guide. If Jev times out, hits a rate limit or isn’t confident, nothing executes. You go back to the customer:
res, err := decide(ctx, c, msg)
if err != nil {
return AskCustomer // timeout or 429: don't act on a missing decision
}
type Route int
const (
Call Route = iota
NeedsApproval
AskCustomer
NotApplicable
)
func route(a map[string]jev.Answer, argsErr error) Route {
switch {
case a["tool"].Choice == "none":
return NotApplicable
case argsErr != nil, a["tool"].Confidence < 0.6:
return AskCustomer
case a["risky"].Noul > 0.5:
return NeedsApproval
}
return Call
}
Why yes or no is not enough
The big takeaway for me is separating the decision of whether to call a tool from the argument generation. On top of that, we sometimes needed to add a flag, some extra options, so the decision isn’t binary, approve or reject. Sometimes you need “not enough information”, or “not applicable”, or some other choice that makes the decision better.
In practice it looks like this. You classify the intent into categories, you classify
the risky action, and you add another field that says there’s not enough information.
It works as a proxy for the question you actually care about: is this risky, or is it
safe? It definitely helped improve the quality of the pipeline itself. In the request
it’s one more Choice:
"info": jev.Choice("Whether `message` contains what the requested action needs", map[string]string{
"enough": "The customer gave everything needed to act.",
"not_enough": "Something needed to act is missing or ambiguous.",
"not_applicable": "No action applies to this message.",
}),
and two more cases at the top of route:
case a["info"].Choice == "not_applicable":
return NotApplicable
case a["info"].Choice == "not_enough":
return AskCustomer
Two details from the Evalgent guide carry over directly. Jev
always returns one of your options and has no built-in “I don’t know”, so the escape
option only exists if you write it. And Jev reads the option descriptions, never the
keys, so not_enough needs a description as concrete as the real options or it won’t
get picked. They also suggest tracking how often the escape options come back. If that
rate climbs, your transcripts or your option set have drifted.
Third option When the model can’t tell, give it an option that says so. Otherwise the uncertainty gets forced into a yes or a no, and you can’t tell those apart from the real ones.
Should you log Jev inputs and decisions?
Another thing we found very useful is collecting the logs, either at the OpenRouter level or in your own code: the inputs and the classifications. That helps if you end up going down that road later. Maybe not in the next six months, but in 18 months you might want to retrain a model and recalibrate the schemas for a better fit. Either way it’s a valuable piece of data, and you should collect it one way or another.
When does collecting it start to pay off? I don’t know. If you collect enough evidence, it would be nice to calibrate and fine-tune the pipeline. Maybe you need to improve the quality, change the fields, or change the schema of the decision. The decision schema is something that emerges over time. And we can’t say what models we’ll have access to in 18 months, or what they’ll be able to do. You might end up with a model you can run with far fewer parameters, cheaper than Jev is today, at a fraction of this cost and this latency. That would be nice to calibrate against, and maybe it gives you a better solution.
Throwing this data away isn’t wise. Just collect it. One JSON line per decision is enough:
type DecisionLog struct {
At time.Time `json:"at"`
Message string `json:"message"`
Answers map[string]jev.Answer `json:"answers"`
Route Route `json:"route"`
Latency time.Duration `json:"latency_ns"`
Outcome string `json:"outcome"` // tool result, approved, rejected, customer reply
}
func logDecision(w io.Writer, d DecisionLog) error {
return json.NewEncoder(w).Encode(d)
}
The latency field tells you whether you’re still inside your budget, and the outcome is what turns a logged decision into a labelled one. Evalgent’s list of per-decision fields goes further, with a call ID, turn number and a hash of the state, and per-tool rates for wrong tool, missed call and false blocks.
Retention Keep the inputs alongside the answers. The answers alone can’t be replayed against a different schema or a different model.
Production logs as unit tests in waiting
Besides that, the logs help you calibrate, they help with maintenance, and they’re useful to have in production anyway. In a way they’re an emerging unit test suite for your production. They aren’t shaped as unit tests yet. But you have your cases, a dataset of real cases, and you can reuse it for just about anything. Once someone has reviewed a few of them, they drop straight into a table-driven test:
func TestRoute(t *testing.T) {
cases := loadReviewed(t, "testdata/decisions.jsonl")
for _, c := range cases {
t.Run(c.Message, func(t *testing.T) {
res, err := decide(t.Context(), client(t), c.Message)
if err != nil {
t.Fatal(err)
}
_, argsErr := extractCancel(c.Message, res.Answers["reason"].Choice)
if got := route(res.Answers, argsErr); got != c.Route {
t.Errorf("route = %v, want %v", got, c.Route)
}
})
}
}
So how do you use Jev for tool calling?
You use Jev for the decision and code for the rest. One request asks whether a tool is needed, which one, whether it’s risky and whether there’s enough information to act. Your own code extracts the arguments with rules and validates them, and a plain if/else policy decides whether the call runs, waits for a human, goes back to the customer, or doesn’t apply. A small decision SDK over OpenRouter is all the infrastructure that takes. Log every input and decision from day one, because that data is what you’ll recalibrate on, and maybe what you’ll move to a cheaper model with later.



