Skip to content
Chokmah
Scenario: not a client engagement

First-pass contract clause review in a procurement team

Abstract stack of clause blocks with select ones lifted and outlined, brand blue to violet on white

This is a methodology scenario, not a client engagement. Chokmah has not delivered this engagement for a named client. The workflow, constraints and sequence describe how we would run it. Any figures shown are illustrative targets, not measured results.

First-pass contract clause review locates playbook-relevant clauses and flags deviations for a reviewer, before legal. It suits agentic assistance because it is comparison against a known standard, but the stakes make recall the priority. This page describes how Chokmah would build it as triage, never as approval.

  • The system triages and cites the clause text; a human, and legal, decide.
  • Make the playbook explicit first; implicit clause positions are out of scope.
  • Recall is weighted above precision: a missed clause costs a commitment.
  • The false-negative rate on material clauses is the primary metric and a release gate.
  • Unlocatable clauses are surfaced as gaps, never suppressed to look complete.

The situation

A procurement team reviews inbound supplier contracts and redlines against a company playbook: acceptable liability caps, payment terms, indemnity language, data-protection and termination clauses. Standard agreements are checked against known-good positions; the difficult ones go to legal.

The first pass is where the time goes. Someone reads each contract, locates the clauses that matter, compares each against the playbook position, and flags the deviations. Much of this is pattern location and comparison against a known standard (work an agentic system can assist), but the stakes make the boundary sharp. A missed non-standard indemnity or a mis-read liability cap is not a productivity issue; it is a commitment the company did not intend to make.

The correct framing is a first-pass triage that surfaces clauses and deviations for a human, never an approval. This scenario is a composite of contract first-pass review as it appears in procurement functions. It describes no specific client and reports no achieved result.

Constraints

  • A contract clause is a binding commitment. A missed or misread clause is a legal and financial exposure, not a defect to be tuned away later.
  • The playbook is not fully written down. Some acceptable positions live in the heads of a few reviewers and in the history of past negotiations.
  • Contracts are unstructured and inconsistent. The same clause appears under different headings, in different order, and in varied language across suppliers.
  • Legal review is the control point and cannot be bypassed. The system operates before legal, never instead of it.
  • A confident wrong reading is worse than an obvious gap. A flag the reviewer trusts and should not is the dangerous output.

How we would run it

  1. 1
    Week 1

    Extract the playbook into explicit, testable positions

    Before any extraction model, make the standard explicit. Work with the reviewers to write down the acceptable position for each clause type that matters (liability, indemnity, payment terms, data protection, termination) including the positions that currently live only in reviewers' heads. If the playbook cannot be made explicit for a clause type, that clause type is out of scope for automation, because there is nothing stable to compare against.

  2. 2
    Week 1

    Shadow a reviewer through several real first-pass reviews

    Watch how a reviewer locates clauses in an unfamiliar contract, decides whether a deviation matters, and chooses what to escalate to legal. The judgement about which deviations are routine and which are dealbreakers is the specification for the triage, and it is exactly the part that cannot be inferred from the contracts alone.

  3. 3
    Week 2

    Baseline first-pass time, escalation quality and miss rate

    Record time to complete a first pass, how often legal receives a clean escalation versus a contract that still needs re-reading, and (using a sample of past contracts with known issues) the current miss rate. The miss rate is the metric that matters and the hardest to get, so it is agreed carefully in writing. A build that speeds up the pass but raises misses is a failure regardless of the time saved.

  4. 4
    Weeks 3-4

    Build a clause locator and deviation flagger that cites the text

    The system locates each playbook-relevant clause in the contract, quotes the exact text, states the playbook position, and flags whether it deviates and how. It does not approve, sign off, or hide anything. Every flag points at the specific contract text so the reviewer verifies against the source rather than trusting a summary. Clauses the system cannot confidently locate are surfaced as gaps, not silently dropped, because a missing flag is the failure mode with legal consequences.

  5. 5
    Weeks 4-5

    Evaluate recall on known-issue contracts, not fluency

    Build a versioned set of past contracts with the clauses and deviations a reviewer identified, and score recall: did the system surface every clause that mattered and flag every real deviation. Recall is weighted above precision here: a false flag costs a reviewer a few seconds, a missed clause costs a commitment. The false-negative rate on material clauses is the primary metric and a release gate.

  6. 6
    Week 6

    Hand over with the escalation boundary written into the process

    The procurement team keeps the repository, the evaluation set and a written statement of what the system flags, what it cannot, and the fact that legal review remains the control point. Because the workflow sits in front of a legal control, the handover makes explicit that the tool triages and a human decides, so no later change can quietly promote a flag into an approval.

What we would not automate

Approving, signing off or accepting any clause

The system triages; it never approves. A contract clause is a binding commitment and the decision to accept a deviation is a legal judgement that stays with a person and, for anything non-standard, with legal. Allowing the tool to clear a clause would convert a review aid into an unaccountable approver, and we would refuse to build that regardless of how good the flagging became.

Clause types whose playbook position cannot be made explicit

If reviewers cannot agree and write down the acceptable position for a clause type, there is no stable standard to compare against, and an automated flag would be guessing. Those clause types stay entirely manual until the standard exists. Automating against an implicit, contested position produces confident flags that are sometimes wrong in both directions.

Suppressing low-confidence clauses to make the review look complete

The tempting failure is to drop clauses the system cannot confidently locate so the output looks clean. That hides exactly the clauses most likely to be non-standard. The correct behaviour is to surface every uncertain or unlocatable clause as an explicit gap for the reviewer, because a missing flag is the output with legal consequences.

Illustrative targets

These are targets used to frame the engagement, not measured results from a delivered client project.

Illustrative target: False-negative rate on material clauses: The primary metric and a release gate

Measured on a versioned set of past contracts with reviewer-identified clauses and deviations. A missed material clause is a commitment the company did not intend, so recall on material clauses is the number the build is held to, weighted above precision because a false flag is cheap and a miss is not.

Illustrative target: Clause types covered at first release: Explicit-playbook clause types only

Scope is set to clause types whose acceptable position the reviewers were able to write down and agree in week one. Clause types that remain implicit or contested stay manual, so coverage is defined by the playbook's explicitness, not by the contract volume.

Illustrative target: First-pass review time: Reduction vs baseline, after recall

Baselined before the build and remeasured after with the same definition, but reported only alongside the false-negative rate. A faster pass that misses more clauses is a regression, so the time figure is never presented on its own.

Frequently asked questions

It can do a first pass: locating the clauses that matter, quoting the text, stating the playbook position and flagging deviations for a reviewer. It should not approve or accept anything, because a contract clause is a binding commitment and that judgement stays with a person, with anything non-standard going to legal. The build is triage in front of the existing legal control, not a replacement for it.

Because the two errors have very different costs. A false flag makes a reviewer spend a few seconds dismissing it; a missed clause lets a non-standard liability or indemnity through as a real commitment. So the evaluation weights recall on material clauses above precision, and the false-negative rate on a set of known-issue contracts is the primary metric and a release gate.

The playbook has to be made explicit. The acceptable position for each clause type must be written down and agreed, including the positions that currently live only in experienced reviewers' heads. Clause types where that cannot be done stay manual, because there is no stable standard for a system to compare against and an automated flag would simply be guessing.

It is a Workflow Sprint: four to six weeks, built with three to five of your own procurement and engineering people, who keep the code and the evaluation set. The first weeks make the playbook explicit, shadow reviewers and baseline the miss rate; the build and evaluation follow; the handover writes the triage-not-approval boundary into the process.

No. This is a methodology scenario, not a case study. Chokmah has not delivered this for a named client. The workflow is a composite of how first-pass contract review appears in procurement functions, and every figure on the page is an illustrative target measured against a baseline, not an achieved result.

Have a workflow like this?

We name the workflow before we start. Book a free AI Reality Check and build one live.