Open Search Console on a blog that’s been around for a couple of years and you’ll see the same thing we did. Lots of pages indexed. Some impressions. Barely any clicks. Some posts were written by hand, some with a model helping, a few by people who don’t work here anymore. Which ones are hurting the site? No idea. To really know, somebody would have to read every post with Google’s rater guidelines open next to it, and nobody was ever going to do that. So the blog just kept getting bigger.
We didn’t read it. We had Jev read it. I didn’t want a chat model writing me a review of every page, I’ve done that before and ended up with a folder of essays nobody opened. What I wanted was ten small questions per page, each with a probability I could sort by. The tool ended up as one Go file, around 300 lines. Most of the time went into the questions, not the code.
Why word counts and Lighthouse miss thin content
The mechanical side of SEO is a solved problem. Titles, meta descriptions, canonicals, broken links, Core Web Vitals, any crawler will list those in a few minutes, and you don’t need a model for any of it. What’s left is the stuff Google’s helpful content guidance keeps going on about. Was this written for people or for the search engine? Has the author actually done the thing? Does the page say anything the top ten results don’t? People usually fall back on word count here, and it barely tells you anything. I’ve seen 3,000-word listicles with less in them than a 700-word field report.
Here’s why Jev works for this. Every one of those questions has a short, fixed list of answers. Thin or not thin. Keep, refresh, rewrite, prune. I don’t need an explanation, I need a number to sort on and a threshold we can fight about in review. And it has to run on every page, so it has to be cheap enough that rerunning it is a non-event.
Seven labels, each a yes/no question
Most of the tool is seven Nouls. That’s Jev’s yes/no question type. Each one targets one
way a page can go wrong, and you give it two short descriptions, one for what true
looks like and one for false. I kept them separate on purpose. Plenty of pages are
generic without being salesy, or sound AI-written and still have real depth. Fold all of
that into one “SEO quality” score and you lose the part an editor actually needs to know.
var labels = map[string]question{
"thin": noul("Is this page too thin to satisfy the search intent implied by its title?",
"Covers the topic superficially; a searcher would need to look elsewhere.",
"Covers the topic with enough depth that a searcher is satisfied."),
"no_value": noul("Does the page lack information gain, adding nothing beyond common knowledge?",
"Restates what top search results already say; no new data, method, or insight.",
"Adds original data, methods, examples, or insight a reader would not find elsewhere."),
// generic, ai_written, keyword_stuffing, outdated, salesy
}
The full list is thin, generic, no_value, ai_written, keyword_stuffing,
outdated and salesy. An editor could check any of them just by reading the page.
We followed the same rule in the intent guardrails
post: don’t ask about motive or ranking, only about what’s actually on the page. If a
label’s probability hits the threshold (0.6 unless you change it), it’s an issue. Pages
under 600 words get a too_short issue straight away, and Jev isn’t involved in that one.
Small thing, big payoff: Write the true criterion as the problem and the false one as
the goal. When a label fires, you already have the brief for fixing it.
Two scores and a verdict
The labels tell you what’s wrong with a page. They don’t say how bad it is overall, or
what to do about it, so three more questions go into the same request. quality is a
Score on five levels, “Useless” up to “Best in class”, and it compares the page to the
best result on the web, not to nothing. google_seo is another five-level Score, from
“Lowest” to “Highest”, worded as the rating a Google Search Quality Rater would give
under the helpful content and E-E-A-T guidelines. Then action is a Choice between four
options, each defined in one line:
"action": {"choice", "What should an SEO editor do with this page?", map[string]string{
"keep": "Strong page; leave as is.",
"refresh": "Solid base; update facts, add specifics, tighten.",
"rewrite": "Low quality; needs a substantial rewrite.",
"prune": "No salvageable value; merge into another page or delete.",
}},
Scores don’t come back as buckets. You get a position between the levels, so the report
says something like google_seo 1.4/4, and that makes sorting a lot easier. Pages are
ordered by how many issues they have, then by google_seo from lowest up, then by
quality. So whatever’s at the top of the screen is the worst stuff. One caveat I want
to be upfront about: google_seo is a model’s take on Google’s published guidelines.
It has no idea how Google actually ranks anything. We use it to sort, nothing more.
Chat model as SEO reviewer
- A critique per page, different every run
- Priced per output token
- You read the prose to decide what to fix
Jev with typed questions
- Ten numbers per page, same shape every run
- Input tokens only, output free
- You sort, threshold and hand the list to an agent
The harness: scan, ask in parallel, sort
The rest is boring code. It walks content/, skips _index.md, drafts and the
translated copies, pulls the front matter out with a regex, and strips Hugo shortcodes so
Jev gets prose and not template calls. Each page turns into a state: URL, title, meta
description, word count, Markdown body. All ten questions go in a single request per
page. Jev answers every question against the same state in parallel, and adding one more
question hardly changes the response time, so there was no reason to split them up.
Jev has a 32k-token context, so we chop bodies at 60,000 characters. For English prose
that leaves a lot of room. Eight requests run at once behind a semaphore. A 429 or 5xx
gets retried with exponential backoff, up to four attempts. Then there are a few flags:
-match for one section, -n and -offset to page through a big site, -threshold and
-min-words if you want to be stricter or looser. Every response includes a usage
block with tokens and cost, so the summary at the end shows what you actually paid.
Which isn’t much. Jev costs $0.042 per million input tokens and output is free. A
2,000-word post plus the questions comes to a few thousand tokens, so roughly two
hundredths of a cent a page. The biggest page the tool will send, 60,000 characters,
still costs less than a tenth of a cent. That’s cheap enough that we don’t save the
audit for once a quarter. We can just run it again whenever content/ changes.
JSON for the rewrite agent
The readable report goes to stderr. Pass -json and stdout gets a task for a coding
agent instead. It lists the flagged pages with path, scores and verdict, plus the
issues. Each issue has the label, its probability, and the problem and goal text taken
straight from the question definitions. Above all that there’s one paragraph of
instructions:
const agentInstructions = "Each page below was flagged for SEO content-quality issues. For each page, edit the markdown file at `path` " +
"so every listed issue moves from `problem` to `goal`. Keep front matter keys, URLs and shortcodes intact. " +
"Do not invent statistics, sources, or product claims; if a fix needs facts you do not have, leave a TODO comment instead. " +
"Pages with action=prune should be reported back, not rewritten."
Almost all of it is about keeping the agent from doing damage. Tell a model to fix
no_value and it’ll go looking for a statistic, and a made-up number hurts you more on
E-E-A-T than having no number at all. So it leaves a TODO and a person fills it in. URLs
and front matter don’t change, because moving a page that ranks just to fix its copy is a
good way to lose the ranking. And prune verdicts come back to a human. Deleting a page
means deciding where it redirects, and that’s not a call I’d hand to an editing agent.
Check the rewrite: After the agent is done, scan the same paths again. Whatever was flagged before should now be under the threshold. If it isn’t, the rewrite didn’t fix what it was asked to.
Where it does not fit
Yes, we’re asking a model whether text sounds like a model wrote it. I know. We kept
ai_written anyway, but we read it as a style signal and nothing more. It fires on stock
phrases, hedging and padded listicles, which are worth fixing no matter who wrote them.
It will also miss clean AI prose and occasionally flag a stiff human writer, so nothing
gets rejected on that label alone. outdated has a similar catch. There’s no date in the
state, so Jev can only tell you a page looks stale, not that it is.
Jev also won’t tell you why. generic at 0.87 says where to look, not which paragraph is
the problem, so the agent still reads the whole page. And all Jev sees is the Markdown. No
rendered page, no internal links, no backlinks, no search data. So this doesn’t replace
Search Console, it sits next to it. Search Console tells you which pages are
underperforming. Jev tells you what’s wrong with the writing on them.
Jev for SEO FAQ
What is the Jev use case for SEO?
Content-quality triage at scale. Jev reads each page and answers typed questions about it, so you can sort a whole blog by likely weakness and hand the worst pages to an editor or a rewrite agent, instead of reading every post by hand.
Can Jev check SEO content quality?
It can answer typed questions about a page’s content: whether it’s thin, generic, AI-sounding, keyword-stuffed, outdated or salesy, plus scores for overall usefulness and a helpful-content style rating. It doesn’t crawl, render, or see search data, so use it next to a crawler and Search Console.
How much does a Jev SEO audit cost?
Jev charges $0.042 per million input tokens and nothing for output. A typical blog post
with ten questions is a few thousand input tokens, about two hundredths of a cent per
page, and the response’s usage.cost field gives the exact figure.
Does the google_seo score predict rankings?
No. It’s Jev’s reading of how a human quality rater would score the page under Google’s published guidelines. It’s useful for sorting pages by likely weakness, not as a prediction of where a page will rank.
How do the flagged pages get fixed?
The tool’s -json output is a task for a coding agent: each flagged page with its
issues as problem and goal pairs, plus instructions to keep URLs and front matter
intact, never invent facts, and send prune verdicts back to a human.
If you take one thing from this, spend your time on the questions. Ask things an editor
could check by reading the page, write the false side as the goal, and the numbers
you sort by double as the brief for the fix. If you want to compare notes on typed
decisions in a content pipeline, get in touch.



