Skip to main content
25% AI slop, checked by JevBook a free call

Articles · How we build · September 30, 2026 · 8 min read

How we used Jev to weigh every decision on our website

We asked an AI judgment model to play our customers before we changed anything: the headline, the demos, the free tools, the outreach email. What moved, what didn’t, and where the method misled us.

MAE-AI Solutions is a small company, and our website gets a few dozen real visits a month. That is far too few to A/B test anything. So for two days we tried something else: before we changed a page, we asked Jev, a judgment model from TypeSafe, to read it as specific customers would, and we only kept changes that scored better against a fair comparison. We logged every question, score and decision, including the ones where we overrode it.

None of the numbers below are conversions. They are Jev’s estimates of how a described person would react, from 0 to 1 (or 0 to 3 on a scale). They are a way to decide faster and more honestly, not a substitute for real customers.

How it worked

  • Give Jev the real thing to read: the page text as a browser renders it, and for the demos, real transcripts of people talking to them.
  • Ask as specific people: a dentist who owns two practices, an HVAC owner who misses calls after hours, an office manager at a small law firm. Two per industry, six industries.
  • Ask narrow, typed questions: would you book a call after this (a probability), how much does this raise your trust (a score), which of these options (a choice).
  • Change one thing, then ask the same people the same question about the old and new versions side by side.
  • For any rewrite, ask one more question: does the new version drop a limit or promise the old one stated? Restore whatever it names.

What moved

  • The homepage headline. An independent AI-writing check flagged “Every customer answered, even when you’re busy” as 74% generic. Jev agreed it was generic (0.78) and rated it barely honest (0.08), because nobody can promise every customer is answered. “Answer the calls you miss, without hiring.” kept the same pull (0.63 against 0.64) and scored 0.67 honest.
  • Six industry pages rewritten from contract language into plain English: the chance of booking rose 9 to 14 points on each. The first draft quietly dropped some limits (up to 0.66 on Jev’s “lost a limit” question); after restoring them, every page scored 0.33 or lower.
  • Five service demo pages: the receptionist demo went from 0.32 to 0.49 and the help desk from 0.30 to 0.45, mostly by explaining in owner’s words what the demo does and doesn’t do, and showing the price.
  • Hard-to-read sentences on our sales pages: Jev scored every one; those below 1.5 of 3 fell from 45 to 18.
  • Our first outreach email: four drafts guessed that a business’s calls go to voicemail. Jev rated that guess 0.37 honest. Asking instead (“I couldn’t tell what a patient hears at 6”) and pointing to our own AI-answered phone line raised the estimated reply rate from 0.26 to 0.42 and cut the “feels like spam” score from 0.58 to 0.33.
  • Our free resume tool: Jev judged, skill by skill, whether each resume really showed what a job posting asked for, and we scored the tool against that. It agreed 43% of the time; after fixes, 89% on those resumes and 65% on five new ones it had never seen (up from 37%).

What didn’t, and where the method misled us

  • Our first demo survey showed a big jump in booking. It was an artifact: we had cut the old pages’ text short, so Jev never saw the section we changed. Measured fairly, booking didn’t move; only one question (“what does it cost and how long does setup take”) did, from 0.39 to 1.32. Always compare against an honest control.
  • The same trap on the resume tool: 89% on the resumes we tuned it on looked great; 65% on new ones is the real number.
  • A credit forecast built from raw usage said our voice agents would run out of credits in two days. Broken down by product, the agents used about 2,000 a day and 94% of the month had gone on one-off image and video generation. The model was fine; the question was wrong.
  • Jev judges what it is shown. It sometimes rated a harmless phrase as a risk because it had no context, and it doubted things we could build in an evening. We logged those overrides rather than hiding them.
  • Some things only a person can supply. “No proof” was the biggest problem on 20 of 33 pages, and no rewrite fixes that: it needs real customers saying real things.

Jev inside the product, too

The same model runs in our website-update demo. You type a change the way you would text your office manager (“we are closed Thanksgiving and the day after”), code finds every time and date in what you wrote, and Jev only chooses among them, in roughly 0.15 to 0.7 seconds per decision. Testing it the same way caught a real bug: “we stopped doing weekend showings” was being read as closing the office on Saturdays.

Try the website-update demo →

Would we do it again?

Yes, with the caveats above. It made us faster, and more honest, which we didn’t expect: several of our best changes were Jev pointing out that a line promised more than we can deliver. It is not a replacement for talking to customers. It is a better way to decide what to show them first.

See the site it shaped →Try the free tools →

Book a free call: 30 minutes with Aaron

Prefer to talk now? Call (804) 391-0234.