🔍 Read the full analysis: How To Use Jev For Better AI Decision-Making: 24 Approaches on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
In a Sept. 29 article, Thorsten Meyer describes 24 ways to use Jev for small, repeated classification and routing decisions. He says three applications are live in his publishing operation, 12 are strong fits, seven need measurement and two are poor fits; those figures and performance results are his own reports.
Thorsten Meyer published a guide on Sept. 29 outlining 24 possible uses for Jev, a tool he describes as returning typed answers to decision questions for software to act on. Meyer says three uses are live in his publishing operation, while 12 meet his four-part fit test, seven need measurement and two are poor fits. The account is based on his own experience and measurements; the supplied material does not include independent verification.
Meyer says Jev accepts a text or JSON state alongside typed questions and returns answers such as a yes probability, a choice with probabilities, or a score on ordered levels. His description emphasizes confidence values: code can take action on clear answers and send uncertain cases for another review. Jev itself does not write or summarize content, he says; the surrounding software determines what to do with each answer.
The three live examples are a story-to-site relevance gate, a language check and a fallback topic classifier. Meyer reports that a scan of 78,889 articles cost $2.01 and found 1,576 non-English items, of which 1,553 were fixed. He also reports about 10,000 relevance pairings judged over three days, with 22% clearly on-topic, and 89% agreement with a frontier LLM for the classifier, rising to 97%–99% when Jev’s confidence was at least 0.8. These are figures from Meyer’s operation, not independently audited results.
The guide also lists publishing applications such as detecting thin sources, identifying disclosures, checking headline quality and moderating comments. For each, Meyer gives a question and an action rule. Meyer says a canary test of the deduplication use case found zero duplicates and recommends against adding the check without evidence of a problem.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
A Test Before Adding More Checks
Meyer recommends considering Jev for high-volume, narrow decisions where errors are inexpensive or uncertain cases can be escalated. He also says a current heuristic should be visibly failing and that the failure should be measured. The guide presents the 24 ideas as candidates for evaluation; it does not establish that each application will improve a workflow.
For teams handling repetitive checks, confidence-based routing could allow more cases to be reviewed without sending every one to a larger model or a person. Meyer’s results do not establish performance in other organizations or tasks. The reported agreement with a frontier LLM compares Jev with that model and does not show that either answer was correct in every case.
How Meyer Proposes Testing Jev
Meyer proposes replaying 300 to 500 past decisions, comparing results overall and across confidence bands, then reviewing 20 disagreements to determine which system was right. He says teams should wire in Jev only where the high-confidence band reaches 95%. His suggested rollout uses a separate feature flag, initially off, followed by a 5%–10% canary.
The process reflects the guide’s four conditions: high volume, a narrow question, low-cost errors or a route for uncertain cases, and a measured failure in the current heuristic. Meyer says existing keyword rules should remain in place if they work. His live examples focus on publishing, while the full list is described as covering commerce, software, business operations and home uses; the supplied source excerpt ends partway through the commerce section.
““Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.””
— Thorsten Meyer
Results Beyond Meyer’s Publishing Fleet
The supplied material does not provide independent validation of Jev’s cost, speed or accuracy figures, nor does it specify the full evaluation method behind the reported agreement rates. It also does not establish whether the same results would hold for other publishers, languages or decision types. The source excerpt is incomplete, so details for many of the 24 proposed applications are not available here.
Measure Before Wider Deployment
Meyer recommends testing proposed applications against real past decisions, reviewing disagreements and measuring performance by confidence level. He advises a small canary rollout only after high-confidence results meet his stated threshold. The source does not announce a product release or a scheduled milestone; whether the other proposed applications will be built or adopted remains open.
Key Questions
What does Jev do?
Meyer describes Jev as a tool that answers typed decision questions about text or JSON input, returning probabilities, classifications or scores for software to use.
How many Jev uses does Meyer say are live?
He reports three live uses in his publishing operation: a relevance gate, a language check and a fallback topic classifier.
What does Meyer recommend before deploying Jev?
He recommends replaying 300 to 500 past decisions, reviewing disagreements and measuring performance by confidence band before a limited canary rollout.
Are the reported accuracy and cost figures independently verified?
The supplied source presents them as Meyer’s measurements from his own operation. It does not include independent verification or enough methodological detail to assess how they would generalize.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
