🔍 Read the full analysis: Jev For AI Teams: 24 Decision-Model Use Cases on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Thorsten Meyer has mapped 24 potential uses for Jev, a decision-model tool that returns typed answers for software to act on. He says three uses are already running in his publishing operation, 12 are strong fits, seven need measurement and two are poor fits; the results are his own and have not been independently verified.
Thorsten Meyer has outlined 24 potential uses for Jev, a tool for making small, high-volume decisions that software can act on, and says three are already running in his publishing operation. His assessment classifies 12 more as strong fits, seven as needing measurement and two as poor fits. It also sets out criteria for evaluating whether routine judgments are suitable for automation.
Meyer says Jev receives text or JSON state alongside typed questions and returns answers that code can use directly, rather than writing or summarizing prose. Its answer types include a yes-or-no probability, a choice with probabilities and confidence, or a score on ordered levels. He estimates one call takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens.
Three uses are described as live in his publishing workflow: checking whether a story fits a site, detecting whether an article is in English, and classifying a headline when a primary language model fails. Meyer reports that a scan of 78,889 articles cost $2.01; it identified 1,576 non-English articles, of which 1,553 were fixed. He also reports about 10,000 story-to-site judgments over three days, with 22% clearly on topic, and 89% agreement with a frontier model for the classifier fallback.
The proposed publishing uses include detecting thin source material, identifying product mismatches in roundups, checking disclosures, assessing headlines and moderating comments. Meyer classifies disclosure checks and comment moderation as strong fits, while marking several other proposals as requiring measurement. He labels same-event deduplication a poor fit, citing a canary test that found no duplicates. The source material also says 88% of the news items he processes start from a bare headline, which he cites as a reason to evaluate a thin-source detector.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
A Test Before Automating Decisions
The proposal addresses how teams might evaluate models for routine classification, filtering or routing. Meyer says a low-cost call is useful only when a problem has been measured and the cost of a wrong answer is limited. For tasks where errors could matter, he recommends acting on high-confidence answers and routing uncertain cases to a person or a more capable system.
Meyer recommends replaying 300 to 500 past decisions, comparing results by confidence band and reviewing 20 disagreements. He says teams should put the system into production only where the high-confidence band reaches 95%, then use a feature flag and a small canary rollout. These are thresholds he proposes; the source does not present them as independently validated standards.
How Jev Fits Into a Workflow
Meyer describes Jev as a decision component, not a content generator. A calling system supplies the material and questions; its own code decides what happens next. For example, a publishing system could keep its existing path for uncertain relevance judgments while dropping only stories that Jev rates as clearly low fit with sufficient confidence.
His four-part fit test requires high volume, a narrow question, errors that are cheap or can be routed for review, and evidence that an existing heuristic fails. The final condition is meant to prevent teams from adding a model where a simple keyword rule already works. Meyer recommends testing against real historical decisions before enabling an automated action.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer
Performance Beyond Meyer’s Tests
The reported costs, timings, accuracy figures and production results come from Meyer’s account. The source material does not provide independent testing, detailed evaluation data or enough information to reproduce the measurements. It is also unclear how performance varies across different topics, languages, datasets or operating conditions.
Seven use cases are marked for measurement because a failing current heuristic has not been established. The source excerpt gives detailed examples mainly from publishing and begins a section on commerce and customer operations, but does not include the full descriptions of all 24 uses. The two poor-fit cases are identified in the overview as a category; the excerpt specifically explains the duplicate-detection example.
Measure Before Wider Rollout
Meyer recommends that teams replay historical decisions, inspect disagreements and compare accuracy across confidence bands. If a use case meets the stated threshold, he recommends enabling it behind a flag, starting with 5% to 10% of units and expanding after observing results. The article does not announce a product launch or give a schedule for additional Jev deployments.
For the seven measurement-first ideas, the next step is to determine whether an existing rule or matcher is failing often enough to justify a change. The article presents these ideas as candidates for evaluation; it does not report measured error rates for them.
Key Questions
What is Jev?
Meyer describes Jev as a tool that takes text or JSON and typed questions, then returns structured answers such as probabilities, choices or scores for software to use.
How many uses does Meyer say are ready?
He says three are already live and 12 more are strong fits, for 15 that are ready to build or already running. Seven need measurement first, and two are classified as poor fits.
What makes a task a strong fit?
Meyer’s test calls for high volume, a narrow question, low-cost or reviewable errors, and measured evidence that the existing heuristic fails.
Are the reported results independently verified?
The supplied source gives Meyer’s own measurements and examples. It does not include independent verification or enough underlying data to reproduce them.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
