Clause

Production-grade agentic contract intelligence — a measured, evidence-driven pipeline over CUAD, with honest negative-result reporting and a research-backed architecture pivot.

$0 / local-first CUAD · 77 test contracts 47 automated tests Phase 3 research: concluded

Where the project landed

ConfigurationMacro F1Notes
Local LLM, unconstrained (v4)0.117zero-shot, all 41 categories in one pass
Local LLM + category batching (v5-G)0.262best LLM configuration measured
Local LLM + batching + BM25 retrieval (v5-H)0.257retrieval did not help further at full scale
Classical: TF-IDF + Logistic Regression0.453current production floor

RED for further standalone-LLM extraction iteration  →  YELLOW conditional go on the resulting hybrid design (ADR-026).

The production architecture pivots to a classical-primary, LLM-secondary hybrid: the TF-IDF+LogReg model decides whether a clause exists; the LLM only explains and cites clauses the classical model already flagged — a job a follow-up validation probe measured at 100% citation-verbatim accuracy and 93.8% explanation accuracy (with a stated upper-bound caveat: that probe used a frontier model, not the deployed local one).

What makes this worth reading

Full documentation