A partner sends over a 40-page master services agreement at 4pm on a Friday and wants a redline before the client call Monday morning. An associate spends three hours doing what is mostly pattern matching, flagging the same indemnity carve-out and the same unilateral termination clause they flagged last month on a nearly identical agreement. The billable value of those three hours is low and the risk of missing something at hour two and a half is real.
This is the gap ai contract review is supposed to close, and in narrow ways it does. The problem is that vendor demos show the narrow ways and stay quiet about everything else. What follows is a practical account of where automated review holds up, where it fails, and why the attorney still owns the outcome.
What it genuinely does well
Extraction is the strongest use case. Given a stack of executed agreements, current tools reliably pull out governing law, term length, renewal mechanics, notice periods, payment terms, assignment restrictions, and the presence or absence of named clause types. This is structured data capture from unstructured text, and it works because contracts are formulaic documents written by people who copy from templates.
The second strong case is comparison against a known standard. If your firm has a playbook, a preferred NDA form, or a set of client-specific fallback positions, software can compare an incoming draft against it and mark deviations. It will tell you that the limitation of liability is uncapped where your standard caps at fees paid in the preceding twelve months. It will tell you the confidentiality term is perpetual where you normally accept five years. This is difference detection, and machines are good at it.
The third case is triage across volume. In a diligence review with 300 vendor agreements, ai contract review can sort by change of control language and flag the 40 that need human eyes first. It does not need to be right about all 300. It needs to be right enough about which 40 matter, and a partner can spot check the rest.
Finally, it is useful for consistency checking within a single document. Defined terms used before they are defined, cross-references to sections that were deleted in an earlier round, a party name that changed on page 12 but not page 31. These are the errors that survive three rounds of human proofreading because human attention degrades over long documents and machine attention does not.
Where it fails, plainly
The failure modes are not random. They cluster.
- Novel or bespoke drafting. The model has seen thousands of standard indemnity clauses. It has not seen the one your opposing counsel wrote at 11pm to solve a specific commercial problem. Non-standard language is where extraction accuracy drops and where the tool is most likely to categorise something incorrectly while reporting high confidence.
- Absence. Software is much better at finding what is present than flagging what is missing. A contract with no limitation of liability clause at all may simply return nothing on that field. A clean report is not the same as a clean contract.
- Interaction between clauses. The termination provision might be fine. The transition services obligation might be fine. The combination, where you can terminate but remain obligated to provide services for eighteen months at a rate set by the other party, is the actual risk. Reading clauses in isolation misses this, and most tools read in isolation.
- Commercial context. Nothing in the document tells the software that this counterparty is your client's largest customer, that the relationship is already strained, or that the client has decided to accept a bad indemnity because the revenue justifies it. Risk is not a property of text.
- Scanned and poorly formatted documents. Optical character recognition on a faxed copy of a 1990s lease with handwritten margin notes produces garbage input, and garbage input produces confident garbage output.
- Confident errors. This is the most dangerous failure mode. When the tool is wrong, it is usually wrong in a way that reads as authoritative. It will state that the agreement contains a mutual non-solicit when it contains a one-way non-solicit, and it will not hedge.
A worked example
Consider a mid-size firm doing commercial real estate work. They receive between 15 and 25 commercial leases a month for review on behalf of tenant clients. Historically an associate reads each one against a checklist and produces a memo.
A workable automated layer looks like this. The lease arrives by email and lands in the matter folder in Filevine or Clio. An intake step extracts the standard data points: premises description, base rent, escalation formula, term, renewal options and their notice deadlines, operating expense structure, assignment and subletting rights, holdover provisions, and the guaranty if there is one. Those fields populate a structured summary and, critically, the renewal and notice dates are written into the matter calendar so they exist somewhere other than a paragraph in a memo.
A second step compares the lease against the firm's tenant-side playbook and produces a deviation list ranked by how far the clause sits from the firm's preferred position. The associate now opens the file with a populated data summary, a calendared set of deadlines, and a list of twelve deviations to assess rather than a blank page and 60 pages of text.
What the associate still does, entirely: decide which of the twelve deviations matter for this client, this building, this market. Read the exclusive use clause against what the client actually intends to sell. Notice that the operating expense definition permits capital expenditure pass-through in a way the deviation list flagged as minor but which, on a fifteen-year term in a building with an ageing HVAC system, is the largest financial exposure in the document. Call the client. Write the advice.
The automation compressed the mechanical portion. It did not touch the judgment.
Why attorney review stays mandatory
The obvious reason is professional responsibility. Model Rule 1.1 and its state analogues make competence non-delegable. There is no version of ai contract review where an unreviewed machine output goes to a client as advice. Supervising lawyers remain responsible for work product regardless of what produced the first draft, and several state bars have now issued guidance making this explicit for generative tools.
The less obvious reason is that the failure modes described above are exactly the ones that a checklist-following associate would also miss, which means the tool does not compensate for the weakest human review. It compensates for the most tedious part of a strong review. A firm that uses automation to skip the senior read is not saving time, it is deferring a problem.
There is also a workflow reason. Contract review is rarely the deliverable. The deliverable is advice, a negotiation position, a redline that reflects a strategy someone decided on. Software can produce a summary. It cannot decide that on this deal you concede the indemnity to hold the line on the liability cap, because that decision depends on facts that live in the client relationship rather than the document.
Where to start
Pick one document type you handle repeatedly. NDAs, standard vendor agreements, commercial leases, whatever your volume actually is. Write down the eight to twelve data points you extract from that document type every single time, and the deviations from standard that you always check for. If you cannot write that list, you do not have a playbook yet, and building the playbook is the real work. The software is the easy part.
Then run automated extraction on 20 recently closed matters where you already know the answers, and measure how often it is wrong and how it is wrong. You will learn more from that exercise than from any demo. Alphovia builds this kind of workflow automation for small and mid-size firms, but the diagnostic step is worth doing before you talk to anyone, including us. If the tool cannot reliably handle documents you already understand, it will not handle the ones you do not.