Production-grade agentic contract intelligence — a measured, evidence-driven pipeline over CUAD, with honest negative-result reporting and a research-backed architecture pivot.
| Configuration | Macro F1 | Notes |
|---|---|---|
| Local LLM, unconstrained (v4) | 0.117 | zero-shot, all 41 categories in one pass |
| Local LLM + category batching (v5-G) | 0.262 | best LLM configuration measured |
| Local LLM + batching + BM25 retrieval (v5-H) | 0.257 | retrieval did not help further at full scale |
| Classical: TF-IDF + Logistic Regression | 0.453 | current production floor |
RED for further standalone-LLM extraction iteration → YELLOW conditional go on the resulting hybrid design (ADR-026).
The production architecture pivots to a classical-primary, LLM-secondary hybrid: the TF-IDF+LogReg model decides whether a clause exists; the LLM only explains and cites clauses the classical model already flagged — a job a follow-up validation probe measured at 100% citation-verbatim accuracy and 93.8% explanation accuracy (with a stated upper-bound caveat: that probe used a frontier model, not the deployed local one).
Full source, tests, and complete commit history.
Direct answers to the objections these results invite.
Literature review, experiments, and the ADR-026 pivot.
The pre-build validation harness and its results table.