🔍 Read the full analysis: The Three-Part AI Workflow I’m Using This September on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Thorsten Meyer describes a three-part AI workflow as of Sept. 29, 2026: Claude Opus 5.5 builds, GPT-6.1 Sol handles detailed investigation and review, and lower-cost alternatives cover specific tasks. His cited benchmark figures suggest substantial cost differences among models with relatively close scores, but they do not establish which model will work best for other users’ workloads.
Thorsten Meyer said Sept. 29 that his current AI workflow uses Claude Opus 5.5 for building and newly released GPT-6.1 Sol for detailed investigation and review, with other models assigned narrower tasks. The account matters to developers weighing model costs, but its benchmark comparisons do not establish which model is best for any particular project.
Meyer’s workflow has three main parts: Opus 5.5 as the primary builder, GPT-6.1 Sol as a second model for investigating specific files or diffs and reviewing work, and a set of alternatives for defined jobs. He uses Opus at high effort for features, APIs, multi-file work and refactors, and at xhigh for harder work such as architecture, migrations and trust boundaries. He assigns documents and scoped subtasks to Sonnet 5.5, while Luna handles classification, extraction and routing.
The figures Meyer cites come from Artificial Analysis Intelligence Index v4.3.x, which he describes as a general-capability index rather than a measure of performance on a reader’s own workload. In that index, Opus 5.5 scores 54 at high effort and costs $1.82 per task; at xhigh, it scores 56 and costs $3.46. GPT-6.1 Sol scores 50 at high effort for $0.32 per task and 51 at xhigh for $0.39. These are the source’s reported task costs; results and costs may differ on other work.
Meyer says Sol’s advantage for review is its lower reported task cost, which makes it practical for him to use routinely as another model examines Opus’s output. He says he turns to Astra or Fable for a second opinion if Sol and Opus disagree. His guidance is to shadow-test models before switching and to treat passing tests as evidence rather than automatic approval to ship.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
A Lower-Cost Seat for Code Review
Meyer’s example shows how developers can divide AI work by task and cost rather than choosing one model for every stage. A lower-cost model may be easier to run routinely as a reviewer, while a higher-cost model remains assigned to work where its results justify the expense. That approach could affect a team’s model budget, but the cited benchmark alone cannot show whether the review catches defects in a real codebase.
He also draws attention to effort settings. On the index figures he cites, raising Opus from medium to max increases its per-task cost from $1.34 to $5.98 while its score rises from 51 to 58. Meyer therefore uses high or xhigh for development and says max is rarely worthwhile for his needs. The trade-off is specific to the reported benchmark and his workflow; teams would need to compare quality, latency and cost on their own tasks.
Benchmarks Behind the Model Split
Meyer’s report frames model selection as a cost-per-task decision. Its table lists six models with scores within about 21 index points, while their reported costs range from $0.07 to $7.63 per task. The comparison includes Opus 5.5, Sonnet 5.5, Fable 5.1, GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna. Meyer says Opus 5.5 has the highest score in that group and Sol and Luna offer lower-cost options for work they can handle.
The report says GPT-6.1 Sol launched Sept. 29 at the same listed token prices as GPT-6 Sol: $2 per million input tokens and $10 per million output tokens. Artificial Analysis had listed medium, high and xhigh settings in the figures Meyer cites. Sol’s high and xhigh results came with reported times to first token of 57 and 69 seconds, respectively, a limitation for interactive use. The comparison is a dated snapshot; the source notes that index results and model settings can change.
““The practical reading: Sol is not the model I ask to build. It is the model I can afford to run on everything.””
— Thorsten Meyer
How the Scores Translate to Projects
The source reports benchmark scores and task costs, but does not provide enough information here to establish how representative those figures are for different software teams, project types or production usage. It does not report results from an independent evaluation of Meyer’s workflow, such as how often Sol catches issues or how the model split affects overall development time.
Meyer also notes that one index point may fall within measurement noise, that low and max settings for GPT-6.1 Sol were not yet listed, and that its higher-effort settings have long times to first token. It remains unclear how the models compare on a reader’s own code, how their prices or benchmark entries may change, and whether a separate model review catches errors when both models share the same flawed requirements.
Test the Split on Real Work
Meyer advises readers to shadow-test before switching models. A team following that approach could compare outputs on representative tasks and record quality, latency, model cost and the human time required to review results. Those checks would show whether the proposed split fits its work better than a single-model setup.
As GPT-6.1 Sol’s index entries and model offerings develop, the reported snapshot may change. Meyer’s account does not give a date for a follow-up or a scheduled benchmark update, so the next meaningful evidence will depend on further published evaluations or results from teams testing the models on their own workloads.
Key Questions
What is the three-part workflow?
Opus 5.5 builds, GPT-6.1 Sol investigates details and reviews work, and alternatives such as Sonnet 5.5 or Luna handle narrower tasks.
What benchmark supports Meyer’s comparisons?
He cites the Artificial Analysis Intelligence Index v4.3.x and says it measures general capability, not performance on a reader’s specific workload.
Why does Meyer use GPT-6.1 Sol for review?
In his cited figures, Sol costs $0.32 to $0.39 per task at high or xhigh effort. Meyer says that price makes routine review practical for him; it does not prove the model will catch issues in every codebase.
What limitation does the report give for Sol?
Meyer reports that Sol at high and xhigh effort took 57 to 69 seconds to produce a first token on the index, which may make those settings less suitable for interactive work.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
