🔍 Read the full analysis: How To Check Ironclad’s Fine Print On Training Agents With OpenAI on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI’s October 6 post describes training a frontier model in hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task criteria, while estimated completion times were simulations; the results do not establish customer-ready performance or real-world time savings.
OpenAI published results on October 6 from a project training its GPT-6 Astra model in hosted copies of Ironclad’s contract-management software, reporting an average 55% of evaluation criteria met across 11 tasks. The work matters because it tests agents on specialized business workflows, but OpenAI says its completion-time figures are simulated and the results do not establish that the tasks can be safely handed to agents without human review.
The tasks were selected by Ironclad staff and OpenAI employees who use the product. They covered legal, commercial and procurement work, including creating nondisclosure agreements, configuring procurement approval processes, and updating a reusable contract clause to reflect a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes to complete each task.
OpenAI scored each task against a rubric of 8 to 50 criteria, depending on its complexity. It says Ironclad supplied hosted product environments for model practice and that training tasks were generated from publicly filed contracts in the SEC’s EDGAR database, with personal information filtered out. OpenAI says it did not use its customer data, its internal contracts or non-public Ironclad customer data.
OpenAI reported that GPT-6 Astra met an average 55.0% of criteria, compared with 41.6% for GPT-5.6 Sol in a high-compute setting. The estimated time per attempt was 19.2 minutes for Astra and 37.0 minutes for Sol. An internal model used in Astra’s development reached 63.7%, while Astra met about 94% of criteria on one showcase task. These are rubric results, not percentages of tasks completed successfully.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Contract Workflow Scores Matter
The results offer a concrete look at a broader AI industry effort: teaching agents to carry out multi-step work inside professional software, rather than only answer questions or generate text. If models become reliable in these environments, they could help with tasks that currently require people to navigate product screens and apply business rules.
But the reported average is not evidence that an agent can safely execute a contract workflow on its own. Approval requirements can be consequential: a procurement process might require Finance approval above a spending limit, a Security review for certain requests, and Legal review of nonstandard terms. Missing even one rule could undermine the workflow. OpenAI’s criteria score does not, by itself, show which requirements were missed or whether each attempt produced an acceptable result.
The project also has implications for software vendors. Working with an AI developer may help expose difficult product workflows and improve an agent’s ability to use them. At the same time, as agents handle more interaction, vendors may need to demonstrate that their products’ value lies in business rules, data, audit records and controls, not only in the screens users operate.
contract management software tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Test Was Set Up
OpenAI framed the project around training models to understand business rules, complete multi-step workflows in specialized software and check the result against the original requirements. The Ironclad work provides a bounded research test of that aim: a set of 11 selected tasks, a hosted product environment and criteria tailored to each task. It does not amount to a broad trial across Ironclad customers or a measurement of performance across all contract work.
OpenAI also invited a small number of other software companies to partner on tasks that current agents cannot reliably complete. The company said prospective partners should bring a specific failing example, subject-matter experts, a secure testing environment and data suitable for research. This makes the Ironclad project both a model evaluation and an example of how software companies could contribute specialized workflows to AI training.
The source post cautioned that human oversight remains necessary when agents may lose track of a business rule. Ironclad CTO Sunita Verma emphasized that agents must preserve the controls teams depend on. The project therefore presents progress on a difficult technical problem while acknowledging that reliability and oversight remain unresolved.
As an affiliate, we earn on qualifying purchases.
What the Evaluation Does Not Show
OpenAI’s published results do not establish real-world customer time savings: the time figures are simulated estimates, cover the 11 research tasks and are not measurements from customers using Ironclad. The criteria averages also do not identify, in the information provided here, which specific requirements Astra missed across the tasks or how often it completed every requirement to an acceptable standard.
It is also unclear how performance would hold up across a wider range of contracts, organizations, jurisdictions and unusual cases, or how the system would behave when business rules conflict or information is incomplete. OpenAI says no non-public Ironclad customer data was used, but the reported results do not answer questions about deployment safeguards, error rates in production or the level of review a customer would need.
The source post names GPT-6 Astra and GPT-5.6 Sol and provides comparative figures, but the material supplied does not give enough detail to independently assess the evaluation design or reproduce its results. The figures should be treated as company-reported research findings, not independent evidence of deployment readiness.
As an affiliate, we earn on qualifying purchases.
Questions for Vendors and Buyers
OpenAI says it is seeking a small number of software-company partners to bring difficult tasks, domain experts, secure environments and research-suitable data. Its next steps and a timetable for any further partner results were not specified in the source material.
Organizations considering agents in contract or procurement systems can ask vendors for more than an average score: which criteria were missed, how performance was tested, what data was used, what approvals remain under human control, and how actions are recorded for audit. Buyers should also ask whether reported speed figures come from measured customer use or simulations. Until those details and broader performance evidence are available, the Ironclad test is best read as a research milestone, not proof that agents can independently run high-consequence workflows.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI test with Ironclad?
OpenAI tested models on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software. Tasks included setting up nondisclosure agreements and procurement approval processes.
Does the 55% result mean GPT-6 Astra completed 55% of tasks?
No. OpenAI reported an average 55% of evaluation criteria met. That is not the share of tasks completed, and it does not show that each workflow was safe or fully correct.
Did Astra save customers time?
The reported attempt times are simulated estimates based on assumed processing and generation speeds. OpenAI did not report measured time savings from customers using Ironclad.
What data did OpenAI say it used?
OpenAI said it created synthetic training tasks from public contracts filed in the SEC’s EDGAR database and filtered out personal information. It said it did not use OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
Can businesses use agents for contract workflows now?
The reported test does not establish that agents are ready to handle contract workflows without review. OpenAI’s own account says human oversight remains important, and the results do not provide production error rates or a deployment timetable.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
