How To Check Ironclad’s Fine Print On Training Agents With OpenAI
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How To Check Ironclad’s Fine Print On Training Agents With OpenAI on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI’s October 6 post describes training a frontier model in hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task criteria, while estimated completion times were simulations; the results do not establish customer-ready performance or real-world time savings.

OpenAI published results on October 6 from a project training its GPT-6 Astra model in hosted copies of Ironclad’s contract-management software, reporting an average 55% of evaluation criteria met across 11 tasks. The work matters because it tests agents on specialized business workflows, but OpenAI says its completion-time figures are simulated and the results do not establish that the tasks can be safely handed to agents without human review.

The tasks were selected by Ironclad staff and OpenAI employees who use the product. They covered legal, commercial and procurement work, including creating nondisclosure agreements, configuring procurement approval processes, and updating a reusable contract clause to reflect a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes to complete each task.

OpenAI scored each task against a rubric of 8 to 50 criteria, depending on its complexity. It says Ironclad supplied hosted product environments for model practice and that training tasks were generated from publicly filed contracts in the SEC’s EDGAR database, with personal information filtered out. OpenAI says it did not use its customer data, its internal contracts or non-public Ironclad customer data.

OpenAI reported that GPT-6 Astra met an average 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol in a high-compute setting. The estimated time per attempt was 19.2 minutes for Astra and 37.0 minutes for Sol. An internal model used in Astra’s development reached 63.7%, while Astra met about 94% of criteria on one showcase task. These are rubric results, not percentages of tasks completed successfully.

At a glance
reportWhen: Published October 6; OpenAI describes t…
The developmentOpenAI published results from a training project with Ironclad, using contract-management workflows to test whether models can perform multi-step work in specialized business software.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Workflow Scores Matter

The results offer a concrete look at a broader AI industry effort: teaching agents to carry out multi-step work inside professional software, rather than only answer questions or generate text. If models become reliable in these environments, they could help with tasks that currently require people to navigate product screens and apply business rules.

But the reported average is not evidence that an agent can safely execute a contract workflow on its own. Approval requirements can be consequential: a procurement process might require Finance approval above a spending limit, a Security review for certain requests, and Legal review of nonstandard terms. Missing even one rule could undermine the workflow. OpenAI’s criteria score does not, by itself, show which requirements were missed or whether each attempt produced an acceptable result.

The project also has implications for software vendors. Working with an AI developer may help expose difficult product workflows and improve an agent’s ability to use them. At the same time, as agents handle more interaction, vendors may need to demonstrate that their products’ value lies in business rules, data, audit records and controls, not only in the screens users operate.

Amazon

contract management software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Test Was Set Up

OpenAI framed the project around training models to understand business rules, complete multi-step workflows in specialized software and check the result against the original requirements. The Ironclad work provides a bounded research test of that aim: a set of 11 selected tasks, a hosted product environment and criteria tailored to each task. It does not amount to a broad trial across Ironclad customers or a measurement of performance across all contract work.

OpenAI also invited a small number of other software companies to partner on tasks that current agents cannot reliably complete. The company said prospective partners should bring a specific failing example, subject-matter experts, a secure testing environment and data suitable for research. This makes the Ironclad project both a model evaluation and an example of how software companies could contribute specialized workflows to AI training.

The source post cautioned that human oversight remains necessary when agents may lose track of a business rule. Ironclad CTO Sunita Verma emphasized that agents must preserve the controls teams depend on. The project therefore presents progress on a difficult technical problem while acknowledging that reliability and oversight remain unresolved.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Evaluation Does Not Show

OpenAI’s published results do not establish real-world customer time savings: the time figures are simulated estimates, cover the 11 research tasks and are not measurements from customers using Ironclad. The criteria averages also do not identify, in the information provided here, which specific requirements Astra missed across the tasks or how often it completed every requirement to an acceptable standard.

It is also unclear how performance would hold up across a wider range of contracts, organizations, jurisdictions and unusual cases, or how the system would behave when business rules conflict or information is incomplete. OpenAI says no non-public Ironclad customer data was used, but the reported results do not answer questions about deployment safeguards, error rates in production or the level of review a customer would need.

The source post names GPT-6 Astra and GPT-5.6 Sol and provides comparative figures, but the material supplied does not give enough detail to independently assess the evaluation design or reproduce its results. The figures should be treated as company-reported research findings, not independent evidence of deployment readiness.

Amazon

AI document review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Questions for Vendors and Buyers

OpenAI says it is seeking a small number of software-company partners to bring difficult tasks, domain experts, secure environments and research-suitable data. Its next steps and a timetable for any further partner results were not specified in the source material.

Organizations considering agents in contract or procurement systems can ask vendors for more than an average score: which criteria were missed, how performance was tested, what data was used, what approvals remain under human control, and how actions are recorded for audit. Buyers should also ask whether reported speed figures come from measured customer use or simulations. Until those details and broader performance evidence are available, the Ironclad test is best read as a research milestone, not proof that agents can independently run high-consequence workflows.

Amazon

contract analysis AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI test with Ironclad?

OpenAI tested models on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software. Tasks included setting up nondisclosure agreements and procurement approval processes.

Does the 55% result mean GPT-6 Astra completed 55% of tasks?

No. OpenAI reported an average 55% of evaluation criteria met. That is not the share of tasks completed, and it does not show that each workflow was safe or fully correct.

Did Astra save customers time?

The reported attempt times are simulated estimates based on assumed processing and generation speeds. OpenAI did not report measured time savings from customers using Ironclad.

What data did OpenAI say it used?

OpenAI said it created synthetic training tasks from public contracts filed in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.

Can businesses use agents for contract workflows now?

The reported test does not establish that agents are ready to handle contract workflows without review. OpenAI’s own account says human oversight remains important, and the results do not provide production error rates or a deployment timetable.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Ultimate Guide To AI-Powered Web And Mobile Development With Grok

xAI has announced Grok Build, a new cross-platform tool for web and mobile development, but key details like features, pricing, and rollout remain unclear.

Razer Surges In Global Coverage

Search interest and media coverage of Razer have surged significantly, with 15 mentions in recent reports, signaling increased public and industry attention.

Oracle Surges In Global Coverage

Oracle’s media mentions have increased significantly, with 39 reports in the latest window, marking a 12-fold rise from baseline, amid rising interest in its recent activities.

Will Kai And Speed Beat The Minecraft Challenge By August 13?

Kai and Speed are attempting to beat a Minecraft challenge set for August 13, with betting markets indicating high confidence. The outcome remains uncertain.