Jev for SEO: our content audit is ten questions and a threshold.

We ran TypeSafe's Jev over every Markdown page on a Hugo blog and asked it the kind of thing a Google quality rater would ask. Is this thin? Generic? Written by a model? Each answer is a probability, the flagged pages go to a rewrite agent, and the whole scan costs less than a coffee.

Jev Use Case for SEO: Auditing Blog Content Quality

Image: METAHEURISTIC

Alexander Myasoedov

+Alexander Myasoedov Alexander writes about the operational side of shipping production AI - agents, retrieval, evals, and the guardrails that keep them from going sideways.

Open Search Console on a blog that’s been around for a couple of years and you’ll see the same thing we did. Lots of pages indexed. Some impressions. Barely any clicks. Some posts were written by hand, some with a model helping, a few by people who don’t work here anymore. Which ones are hurting the site? No idea. To really know, somebody would have to read every post with Google’s rater guidelines open next to it, and nobody was ever going to do that. So the blog just kept getting bigger.

We didn’t read it. We had Jev read it. I didn’t want a chat model writing me a review of every page, I’ve done that before and ended up with a folder of essays nobody opened. What I wanted was ten small questions per page, each with a probability I could sort by. The tool ended up as one Go file, around 300 lines. Most of the time went into the questions, not the code.

Why word counts and Lighthouse miss thin content

The mechanical side of SEO is a solved problem. Titles, meta descriptions, canonicals, broken links, Core Web Vitals, any crawler will list those in a few minutes, and you don’t need a model for any of it. What’s left is the stuff Google’s helpful content guidance keeps going on about. Was this written for people or for the search engine? Has the author actually done the thing? Does the page say anything the top ten results don’t? People usually fall back on word count here, and it barely tells you anything. I’ve seen 3,000-word listicles with less in them than a 700-word field report.

Here’s why Jev works for this. Every one of those questions has a short, fixed list of answers. Thin or not thin. Keep, refresh, rewrite, prune. I don’t need an explanation, I need a number to sort on and a threshold we can fight about in review. And it has to run on every page, so it has to be cheap enough that rerunning it is a non-event.

Seven labels, each a yes/no question

Most of the tool is seven Nouls. That’s Jev’s yes/no question type. Each one targets one way a page can go wrong, and you give it two short descriptions, one for what true looks like and one for false. I kept them separate on purpose. Plenty of pages are generic without being salesy, or sound AI-written and still have real depth. Fold all of that into one “SEO quality” score and you lose the part an editor actually needs to know.

var labels = map[string]question{
    "thin": noul("Is this page too thin to satisfy the search intent implied by its title?",
        "Covers the topic superficially; a searcher would need to look elsewhere.",
        "Covers the topic with enough depth that a searcher is satisfied."),
    "no_value": noul("Does the page lack information gain, adding nothing beyond common knowledge?",
        "Restates what top search results already say; no new data, method, or insight.",
        "Adds original data, methods, examples, or insight a reader would not find elsewhere."),
    // generic, ai_written, keyword_stuffing, outdated, salesy
}

The full list is thin, generic, no_value, ai_written, keyword_stuffing, outdated and salesy. An editor could check any of them just by reading the page. We followed the same rule in the intent guardrails post: don’t ask about motive or ranking, only about what’s actually on the page. If a label’s probability hits the threshold (0.6 unless you change it), it’s an issue. Pages under 600 words get a too_short issue straight away, and Jev isn’t involved in that one.

Figure - Seven independent yes/no labels, one threshold (toy pages)
page
threshold
Switch pages and watch each Noul land on its own: a listicle can be generic without being salesy. Drag the threshold and only the bars that cross it become issues, each carrying its problem and goal text into the agent task on the right.

Small thing, big payoff: Write the true criterion as the problem and the false one as the goal. When a label fires, you already have the brief for fixing it.

Two scores and a verdict

The labels tell you what’s wrong with a page. They don’t say how bad it is overall, or what to do about it, so three more questions go into the same request. quality is a Score on five levels, “Useless” up to “Best in class”, and it compares the page to the best result on the web, not to nothing. google_seo is another five-level Score, from “Lowest” to “Highest”, worded as the rating a Google Search Quality Rater would give under the helpful content and E-E-A-T guidelines. Then action is a Choice between four options, each defined in one line:

"action": {"choice", "What should an SEO editor do with this page?", map[string]string{
    "keep":    "Strong page; leave as is.",
    "refresh": "Solid base; update facts, add specifics, tighten.",
    "rewrite": "Low quality; needs a substantial rewrite.",
    "prune":   "No salvageable value; merge into another page or delete.",
}},

Scores don’t come back as buckets. You get a position between the levels, so the report says something like google_seo 1.4/4, and that makes sorting a lot easier. Pages are ordered by how many issues they have, then by google_seo from lowest up, then by quality. So whatever’s at the top of the screen is the worst stuff. One caveat I want to be upfront about: google_seo is a model’s take on Google’s published guidelines. It has no idea how Google actually ranks anything. We use it to sort, nothing more.

Chat model as SEO reviewer

  • A critique per page, different every run
  • Priced per output token
  • You read the prose to decide what to fix

Jev with typed questions

  • Ten numbers per page, same shape every run
  • Input tokens only, output free
  • You sort, threshold and hand the list to an agent

The harness: scan, ask in parallel, sort

The rest is boring code. It walks content/, skips _index.md, drafts and the translated copies, pulls the front matter out with a regex, and strips Hugo shortcodes so Jev gets prose and not template calls. Each page turns into a state: URL, title, meta description, word count, Markdown body. All ten questions go in a single request per page. Jev answers every question against the same state in parallel, and adding one more question hardly changes the response time, so there was no reason to split them up.

Jev has a 32k-token context, so we chop bodies at 60,000 characters. For English prose that leaves a lot of room. Eight requests run at once behind a semaphore. A 429 or 5xx gets retried with exponential backoff, up to four attempts. Then there are a few flags: -match for one section, -n and -offset to page through a big site, -threshold and -min-words if you want to be stricter or looser. Every response includes a usage block with tokens and cost, so the summary at the end shows what you actually paid.

Which isn’t much. Jev costs $0.042 per million input tokens and output is free. A 2,000-word post plus the questions comes to a few thousand tokens, so roughly two hundredths of a cent a page. The biggest page the tool will send, 60,000 characters, still costs less than a tenth of a cent. That’s cheap enough that we don’t save the audit for once a quarter. We can just run it again whenever content/ changes.

JSON for the rewrite agent

The readable report goes to stderr. Pass -json and stdout gets a task for a coding agent instead. It lists the flagged pages with path, scores and verdict, plus the issues. Each issue has the label, its probability, and the problem and goal text taken straight from the question definitions. Above all that there’s one paragraph of instructions:

const agentInstructions = "Each page below was flagged for SEO content-quality issues. For each page, edit the markdown file at `path` " +
    "so every listed issue moves from `problem` to `goal`. Keep front matter keys, URLs and shortcodes intact. " +
    "Do not invent statistics, sources, or product claims; if a fix needs facts you do not have, leave a TODO comment instead. " +
    "Pages with action=prune should be reported back, not rewritten."

Almost all of it is about keeping the agent from doing damage. Tell a model to fix no_value and it’ll go looking for a statistic, and a made-up number hurts you more on E-E-A-T than having no number at all. So it leaves a TODO and a person fills it in. URLs and front matter don’t change, because moving a page that ranks just to fix its copy is a good way to lose the ranking. And prune verdicts come back to a human. Deleting a page means deciding where it redirects, and that’s not a call I’d hand to an editing agent.

Check the rewrite: After the agent is done, scan the same paths again. Whatever was flagged before should now be under the threshold. If it isn’t, the rewrite didn’t fix what it was asked to.

Where it does not fit

Yes, we’re asking a model whether text sounds like a model wrote it. I know. We kept ai_written anyway, but we read it as a style signal and nothing more. It fires on stock phrases, hedging and padded listicles, which are worth fixing no matter who wrote them. It will also miss clean AI prose and occasionally flag a stiff human writer, so nothing gets rejected on that label alone. outdated has a similar catch. There’s no date in the state, so Jev can only tell you a page looks stale, not that it is.

Jev also won’t tell you why. generic at 0.87 says where to look, not which paragraph is the problem, so the agent still reads the whole page. And all Jev sees is the Markdown. No rendered page, no internal links, no backlinks, no search data. So this doesn’t replace Search Console, it sits next to it. Search Console tells you which pages are underperforming. Jev tells you what’s wrong with the writing on them.

Jev for SEO FAQ

What is the Jev use case for SEO?

Content-quality triage at scale. Jev reads each page and answers typed questions about it, so you can sort a whole blog by likely weakness and hand the worst pages to an editor or a rewrite agent, instead of reading every post by hand.

Can Jev check SEO content quality?

It can answer typed questions about a page’s content: whether it’s thin, generic, AI-sounding, keyword-stuffed, outdated or salesy, plus scores for overall usefulness and a helpful-content style rating. It doesn’t crawl, render, or see search data, so use it next to a crawler and Search Console.

How much does a Jev SEO audit cost?

Jev charges $0.042 per million input tokens and nothing for output. A typical blog post with ten questions is a few thousand input tokens, about two hundredths of a cent per page, and the response’s usage.cost field gives the exact figure.

Does the google_seo score predict rankings?

No. It’s Jev’s reading of how a human quality rater would score the page under Google’s published guidelines. It’s useful for sorting pages by likely weakness, not as a prediction of where a page will rank.

How do the flagged pages get fixed?

The tool’s -json output is a task for a coding agent: each flagged page with its issues as problem and goal pairs, plus instructions to keep URLs and front matter intact, never invent facts, and send prune verdicts back to a human.

If you take one thing from this, spend your time on the questions. Ask things an editor could check by reading the page, write the false side as the goal, and the numbers you sort by double as the brief for the fix. If you want to compare notes on typed decisions in a content pipeline, get in touch.

Follow in Google

Make Metaheuristic a preferred source.

One tap and posts like this one surface higher in your Top Stories.

Work with us

Production AI, with guardrails.

Start with a fixed-scope AI Workflow Audit. We map the opportunity and quote a build.

Start a Discovery Sprint →