[{"data":1,"prerenderedAt":51},["ShallowReactive",2],{"blog-tag-ai-governance":3},[4,24,37],{"id":5,"slug":6,"body":7,"html":8,"title":9,"description":10,"category":11,"tags":12,"author":17,"date":18,"year":19,"month":20,"quarter":21,"status":22,"featured":23},"2026\u002F09\u002Findustry-applications\u002Fproject-knowledge-plane-contextkeep","project-knowledge-plane-contextkeep","\nA document controller finds the tab at 4:40 p.m. A coordinator has pasted six pages of the client’s executed contract — retention, liquidated damages, the confidentiality schedule — into a personal ChatGPT account to “check the wording.” The answer looks clean. There is no project boundary, no revision stamp, and no record of what left the building. IT’s draft policy arrives the next morning: ban consumer AI for client files. By Friday, people are still pasting — from home laptops, from personal phones, from the same hunger that made the ban feel urgent.\n\nOr the other version of the same failure. A PM asks an assistant which fire-rating detail applies to Level 3. The model answers from a sheet that was superseded two weeks ago. The RFI goes out citing Rev B. Shop drawings move. Field work starts. The document controller finds the mismatch when Rev D was already current. Nobody can show what the model retrieved, because the session lived in a personal account that was never part of the job.\n\nThat is not an AI capability gap. It is a missing project knowledge plane: tenanted spaces, mandatory citations, permissioned skills and an auditable trail — before anyone treats chat as production practice.\n\n## Banning paste does not stop the hunt\n\nProject Directors and Innovation leads already know the demand. People want answers from the job — drawings, contracts, specs, RFIs, submittals — not from generic training data. They want reusable skills that travel across jobs: contract readers, RFI drafters, rate look-ups, scheduling helpers. They want something that feels like an operating system for that work, not one more chat window.\n\nWhat they do not have is a way to run that practice on live projects without three unacceptable outcomes: client PDFs leaking into personal accounts, agents that write into Procore or email without a named human, and a trail that evaporates when counsel or the owner asks what touched the job.\n\nIT bans push usage underground. Personal Claude and ChatGPT sessions become the unofficial knowledge layer. Custom GPTs and laptop skills accumulate on individual machines. Excel “company memory” of rates and lessons never links back to project provenance. The firm still pays for Procore, Aconex or SharePoint as the document store — and still cannot prove what an assistant saw or did.\n\nThe competing status quo is not “no AI.” It is shadow AI with no tenancy, no citations and no approval gates.\n\n## Wrong revision is not a soft error\n\nKnowledge failures on jobs were expensive before generative tools. Wrong drawing revision cited. Outdated rate used. Lesson learned never found. Models amplify the risk because the answer looks authoritative while the source is invisible.\n\nSupersede has to be structural. When a drawing or spec revision is superseded, default retrieval must prefer current. Historical revisions stay available for deliberate history queries — with that fact disclosed in the citation — so yesterday’s sheet cannot silently answer today’s question. Document controllers already own revision discipline in the CDE; the knowledge plane has to honour the same map, not invent a second filing tree that drifts.\n\nUncited answers that export into RFIs, emails or commercial packs are how field and commercial errors get dressed as confidence. If an answer cannot name document identity, revision, page or chunk locus and retrieval time, it should be marked ungrounded and blocked from export. Citation is not etiquette. It is the difference between assist and liability.\n\n## Prove what touched the job\n\nOwner contracts and confidentiality clauses make personal uploads structurally unacceptable for many firms. Data residency and retention are not policy PDFs — they are product obligations. When a dispute or owner audit arrives, the firm needs an append-only record: who asked, which space, which skill version, which model route, which tools were called, which citations were used, who approved any side effect, and a hash of what went out.\n\nThat is the question Innovation and IT both care about, even when they use different words: can we show what AI touched on this live project, under whose authority?\n\nAn agent that drafts an RFI from cited sources is useful. An agent that files it, emails the client or mutates a schedule without a named approver is a commercial and legal liability. Side effects outside the knowledge plane — create, send, write, mutate — belong in an approval queue with the proposed payload and citations visible until an authorised human confirms. Deny-by-default tool grants. No privilege escalation at runtime. Fail closed and audit the attempt.\n\nSafety-critical means and methods stay human-owned. The plane drafts and retrieves. It does not certify how to build.\n\n## The unit is the knowledge space, not the chat thread\n\nThe unit of tenancy is the knowledge space — project space and company space. Documents, retrievals, skill runs and agent actions are scoped to a space. No silent cross-space retrieval. Client A drawings do not appear in Client B answers. Company memory — historical rates, lessons, standard procedures — does not leak into another client’s space without an explicit, audited promotion path with named approval and optional redaction of client identifiers.\n\nSkills are first-class, versioned artefacts: declared inputs, tool allow-lists, model policy, space scopes. Ad-hoc prompts may help draft a skill; they cannot permanently elevate privileges. Digital champions author and publish versions; document controllers own ingest quality and supersede maps; security owns residency, retention and legal hold.\n\nThe plane stays model-agnostic. Skills declare an allowed model class or pin; operators re-route providers when quality or cost shifts. Claude-to-elsewhere churn must not destroy the library. Locking the firm to one vendor’s proprietary skill format as the sole runtime contradicts how buyers already behave.\n\nContextkeep is not the system of record for drawings, contracts or RFIs. It syncs with Procore, Aconex, SharePoint and peers, records external object identifiers, and keeps indexed derivatives and citations so truth can be reconciled upstream. Firms will not rip out the CDE. Adoption starts by mirroring a live project’s document tree into a space — not by promising another mega-platform replacement.\n\nWhat people actually open: a **space home** for the live job; a **document library** with supersede maps and ingest fitness flags; **Ask with citations** where every answer carries document, revision and page or chunk — plus a **citation proof viewer** when counsel asks how you knew; a **skills gallery and studio** for versioned estimating, contracts and scheduling skills with tool allow-lists and model routes that can change without rewriting the skill; **skill run detail** showing retrieval, tools and citations for one run; an **action approvals** queue for anything that would write to Procore, email or schedule; **company memory** promotion with redaction; **connectors** health; and org admin for residency, permissions and audit export. Ungrounded answers stay in the console — they do not export into an RFI or a bid.\n\n## What “better” looks like on the ground\n\nValue shows up in measures Project Directors and Innovation leads already argue about:\n\n- **Hours hunting docs** — time PMs and coordinators spend searching instead of acting, once answers come from the space with citations.\n- **Share of answers with complete citations** — grounded runs versus ungrounded assists that never leave the console.\n- **Wrong-revision rework** — RFIs and submittals rooted in superseded sheets, driven toward near zero when export is blocked without citations and supersede is enforced.\n- **Governed usage versus shadow AI** — skill runs and approved actions on tenanted spaces versus personal consumer accounts of client PDFs.\n- **Side effects through the gate** — count of agent writes that passed approval versus anything that would have fired unsupervised.\n- **Time-to-first useful skill run** on a new project, and dispute-ready audit export time when counsel asks.\n\nThose are operational outcomes. They do not require a foundation-model training story on customer documents. Training on client corpora is a later, planned question — not the first cut.\n\n## What this is not\n\nIt is not a Procore feature-parity pitch. The CDE stays the system of record. The knowledge plane is the governed practice sitting on top of how people already try to use AI — with citations, permissions and audit.\n\nIt is not Quantspan. Rate libraries and past-bid memory may live here for cited retrieval into estimating; the estimating worksheet and bid package export stay Quantspan’s lane.\n\nIt is not Planvector. Drawing PDFs and metadata are stored and cited here; sheet geometry and take-off-ready vectorization stay Planvector’s job.\n\nIt is not Crewspan. Cited answers and draft payloads can feed the execution cockpit; the PM’s daily home for RFIs, look-aheads and field issues is Crewspan.\n\nIt is not Awardbind. Clause and exhibit citation feeds commercial instruments; award recommendations and the commercial spine stay Awardbind.\n\nIt is not a “build a knowledge base with Claude” tutorial. Consumer chat tools remain outside and unsupported as a store of client documents. The product is tenanted retrieval and permissioned skill runs — not another prompt library on a laptop.\n\n## First cut on one live space\n\nStart narrow. Pick one live project where personal paste is already the pain, and where document control can stand behind the ingest:\n\n1. Stand up one project knowledge space that mirrors the job’s CDE folders — Procore, Aconex or SharePoint — with ACLs, residency and revision supersede enforced.\n2. Publish a small pack of governed skills (typically a handful, not a marketplace): declared tool allow-lists, version pins, deny-by-default scopes. Skills may draft; they may not act outside the plane without approval.\n3. Require citations on anything exported — RFI language, email paste, clause packs, rate seeds. Ungrounded answers stay in the console; they do not leave.\n4. Put agent side effects in an approval queue with payload and citations visible. Measure approval latency and the share of side effects that never bypass the gate.\n5. Leave estimating worksheets in Quantspan, sheet geometry in Planvector, day-to-day coordination in Crewspan and commercial instruments in Awardbind. Measure hunting hours, citation rate, wrong-revision incidents and shadow-AI displacement on the Contextkeep slice alone.\n\nThat is what [Contextkeep](https:\u002F\u002Fxzero.media\u002Fatlas\u002Fapps\u002Fcontextkeep) is built to be: Atlas’s project knowledge plane and governed skills OS — the operating layer between consumer chat tools and the systems of record contractors already run. Mid-market GCs and specialty trades already experimenting with Claude Skills and “chat with the job folder” are the natural wedge: enough AI hunger to hurt, enough confidentiality pressure that bans alone will not hold.\n\nScope the cut in a [Solution Definition Sprint](\u002Fservices\u002Fsolution-definition-sprint): which project, which document classes, which skills, which approvers, which residency and retention rules, which audit export path.\n\nSee [AEC and built environment](\u002Findustries\u002Faec-built-environment), explore [Contextkeep on the Atlas](https:\u002F\u002Fxzero.media\u002Fatlas\u002Fapps\u002Fcontextkeep), or [bring us the paste problem IT cannot ban away](\u002Fcontact).\n","\u003Cp>A document controller finds the tab at 4:40 p.m. A coordinator has pasted six pages of the client’s executed contract — retention, liquidated damages, the confidentiality schedule — into a personal ChatGPT account to “check the wording.” The answer looks clean. There is no project boundary, no revision stamp, and no record of what left the building. IT’s draft policy arrives the next morning: ban consumer AI for client files. By Friday, people are still pasting — from home laptops, from personal phones, from the same hunger that made the ban feel urgent.\u003C\u002Fp>\n\u003Cp>Or the other version of the same failure. A PM asks an assistant which fire-rating detail applies to Level 3. The model answers from a sheet that was superseded two weeks ago. The RFI goes out citing Rev B. Shop drawings move. Field work starts. The document controller finds the mismatch when Rev D was already current. Nobody can show what the model retrieved, because the session lived in a personal account that was never part of the job.\u003C\u002Fp>\n\u003Cp>That is not an AI capability gap. It is a missing project knowledge plane: tenanted spaces, mandatory citations, permissioned skills and an auditable trail — before anyone treats chat as production practice.\u003C\u002Fp>\n\u003Ch2>Banning paste does not stop the hunt\u003C\u002Fh2>\n\u003Cp>Project Directors and Innovation leads already know the demand. People want answers from the job — drawings, contracts, specs, RFIs, submittals — not from generic training data. They want reusable skills that travel across jobs: contract readers, RFI drafters, rate look-ups, scheduling helpers. They want something that feels like an operating system for that work, not one more chat window.\u003C\u002Fp>\n\u003Cp>What they do not have is a way to run that practice on live projects without three unacceptable outcomes: client PDFs leaking into personal accounts, agents that write into Procore or email without a named human, and a trail that evaporates when counsel or the owner asks what touched the job.\u003C\u002Fp>\n\u003Cp>IT bans push usage underground. Personal Claude and ChatGPT sessions become the unofficial knowledge layer. Custom GPTs and laptop skills accumulate on individual machines. Excel “company memory” of rates and lessons never links back to project provenance. The firm still pays for Procore, Aconex or SharePoint as the document store — and still cannot prove what an assistant saw or did.\u003C\u002Fp>\n\u003Cp>The competing status quo is not “no AI.” It is shadow AI with no tenancy, no citations and no approval gates.\u003C\u002Fp>\n\u003Ch2>Wrong revision is not a soft error\u003C\u002Fh2>\n\u003Cp>Knowledge failures on jobs were expensive before generative tools. Wrong drawing revision cited. Outdated rate used. Lesson learned never found. Models amplify the risk because the answer looks authoritative while the source is invisible.\u003C\u002Fp>\n\u003Cp>Supersede has to be structural. When a drawing or spec revision is superseded, default retrieval must prefer current. Historical revisions stay available for deliberate history queries — with that fact disclosed in the citation — so yesterday’s sheet cannot silently answer today’s question. Document controllers already own revision discipline in the CDE; the knowledge plane has to honour the same map, not invent a second filing tree that drifts.\u003C\u002Fp>\n\u003Cp>Uncited answers that export into RFIs, emails or commercial packs are how field and commercial errors get dressed as confidence. If an answer cannot name document identity, revision, page or chunk locus and retrieval time, it should be marked ungrounded and blocked from export. Citation is not etiquette. It is the difference between assist and liability.\u003C\u002Fp>\n\u003Ch2>Prove what touched the job\u003C\u002Fh2>\n\u003Cp>Owner contracts and confidentiality clauses make personal uploads structurally unacceptable for many firms. Data residency and retention are not policy PDFs — they are product obligations. When a dispute or owner audit arrives, the firm needs an append-only record: who asked, which space, which skill version, which model route, which tools were called, which citations were used, who approved any side effect, and a hash of what went out.\u003C\u002Fp>\n\u003Cp>That is the question Innovation and IT both care about, even when they use different words: can we show what AI touched on this live project, under whose authority?\u003C\u002Fp>\n\u003Cp>An agent that drafts an RFI from cited sources is useful. An agent that files it, emails the client or mutates a schedule without a named approver is a commercial and legal liability. Side effects outside the knowledge plane — create, send, write, mutate — belong in an approval queue with the proposed payload and citations visible until an authorised human confirms. Deny-by-default tool grants. No privilege escalation at runtime. Fail closed and audit the attempt.\u003C\u002Fp>\n\u003Cp>Safety-critical means and methods stay human-owned. The plane drafts and retrieves. It does not certify how to build.\u003C\u002Fp>\n\u003Ch2>The unit is the knowledge space, not the chat thread\u003C\u002Fh2>\n\u003Cp>The unit of tenancy is the knowledge space — project space and company space. Documents, retrievals, skill runs and agent actions are scoped to a space. No silent cross-space retrieval. Client A drawings do not appear in Client B answers. Company memory — historical rates, lessons, standard procedures — does not leak into another client’s space without an explicit, audited promotion path with named approval and optional redaction of client identifiers.\u003C\u002Fp>\n\u003Cp>Skills are first-class, versioned artefacts: declared inputs, tool allow-lists, model policy, space scopes. Ad-hoc prompts may help draft a skill; they cannot permanently elevate privileges. Digital champions author and publish versions; document controllers own ingest quality and supersede maps; security owns residency, retention and legal hold.\u003C\u002Fp>\n\u003Cp>The plane stays model-agnostic. Skills declare an allowed model class or pin; operators re-route providers when quality or cost shifts. Claude-to-elsewhere churn must not destroy the library. Locking the firm to one vendor’s proprietary skill format as the sole runtime contradicts how buyers already behave.\u003C\u002Fp>\n\u003Cp>Contextkeep is not the system of record for drawings, contracts or RFIs. It syncs with Procore, Aconex, SharePoint and peers, records external object identifiers, and keeps indexed derivatives and citations so truth can be reconciled upstream. Firms will not rip out the CDE. Adoption starts by mirroring a live project’s document tree into a space — not by promising another mega-platform replacement.\u003C\u002Fp>\n\u003Cp>What people actually open: a \u003Cstrong>space home\u003C\u002Fstrong> for the live job; a \u003Cstrong>document library\u003C\u002Fstrong> with supersede maps and ingest fitness flags; \u003Cstrong>Ask with citations\u003C\u002Fstrong> where every answer carries document, revision and page or chunk — plus a \u003Cstrong>citation proof viewer\u003C\u002Fstrong> when counsel asks how you knew; a \u003Cstrong>skills gallery and studio\u003C\u002Fstrong> for versioned estimating, contracts and scheduling skills with tool allow-lists and model routes that can change without rewriting the skill; \u003Cstrong>skill run detail\u003C\u002Fstrong> showing retrieval, tools and citations for one run; an \u003Cstrong>action approvals\u003C\u002Fstrong> queue for anything that would write to Procore, email or schedule; \u003Cstrong>company memory\u003C\u002Fstrong> promotion with redaction; \u003Cstrong>connectors\u003C\u002Fstrong> health; and org admin for residency, permissions and audit export. Ungrounded answers stay in the console — they do not export into an RFI or a bid.\u003C\u002Fp>\n\u003Ch2>What “better” looks like on the ground\u003C\u002Fh2>\n\u003Cp>Value shows up in measures Project Directors and Innovation leads already argue about:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Hours hunting docs\u003C\u002Fstrong> — time PMs and coordinators spend searching instead of acting, once answers come from the space with citations.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Share of answers with complete citations\u003C\u002Fstrong> — grounded runs versus ungrounded assists that never leave the console.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Wrong-revision rework\u003C\u002Fstrong> — RFIs and submittals rooted in superseded sheets, driven toward near zero when export is blocked without citations and supersede is enforced.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Governed usage versus shadow AI\u003C\u002Fstrong> — skill runs and approved actions on tenanted spaces versus personal consumer accounts of client PDFs.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Side effects through the gate\u003C\u002Fstrong> — count of agent writes that passed approval versus anything that would have fired unsupervised.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Time-to-first useful skill run\u003C\u002Fstrong> on a new project, and dispute-ready audit export time when counsel asks.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Those are operational outcomes. They do not require a foundation-model training story on customer documents. Training on client corpora is a later, planned question — not the first cut.\u003C\u002Fp>\n\u003Ch2>What this is not\u003C\u002Fh2>\n\u003Cp>It is not a Procore feature-parity pitch. The CDE stays the system of record. The knowledge plane is the governed practice sitting on top of how people already try to use AI — with citations, permissions and audit.\u003C\u002Fp>\n\u003Cp>It is not Quantspan. Rate libraries and past-bid memory may live here for cited retrieval into estimating; the estimating worksheet and bid package export stay Quantspan’s lane.\u003C\u002Fp>\n\u003Cp>It is not Planvector. Drawing PDFs and metadata are stored and cited here; sheet geometry and take-off-ready vectorization stay Planvector’s job.\u003C\u002Fp>\n\u003Cp>It is not Crewspan. Cited answers and draft payloads can feed the execution cockpit; the PM’s daily home for RFIs, look-aheads and field issues is Crewspan.\u003C\u002Fp>\n\u003Cp>It is not Awardbind. Clause and exhibit citation feeds commercial instruments; award recommendations and the commercial spine stay Awardbind.\u003C\u002Fp>\n\u003Cp>It is not a “build a knowledge base with Claude” tutorial. Consumer chat tools remain outside and unsupported as a store of client documents. The product is tenanted retrieval and permissioned skill runs — not another prompt library on a laptop.\u003C\u002Fp>\n\u003Ch2>First cut on one live space\u003C\u002Fh2>\n\u003Cp>Start narrow. Pick one live project where personal paste is already the pain, and where document control can stand behind the ingest:\u003C\u002Fp>\n\u003Col>\n\u003Cli>Stand up one project knowledge space that mirrors the job’s CDE folders — Procore, Aconex or SharePoint — with ACLs, residency and revision supersede enforced.\u003C\u002Fli>\n\u003Cli>Publish a small pack of governed skills (typically a handful, not a marketplace): declared tool allow-lists, version pins, deny-by-default scopes. Skills may draft; they may not act outside the plane without approval.\u003C\u002Fli>\n\u003Cli>Require citations on anything exported — RFI language, email paste, clause packs, rate seeds. Ungrounded answers stay in the console; they do not leave.\u003C\u002Fli>\n\u003Cli>Put agent side effects in an approval queue with payload and citations visible. Measure approval latency and the share of side effects that never bypass the gate.\u003C\u002Fli>\n\u003Cli>Leave estimating worksheets in Quantspan, sheet geometry in Planvector, day-to-day coordination in Crewspan and commercial instruments in Awardbind. Measure hunting hours, citation rate, wrong-revision incidents and shadow-AI displacement on the Contextkeep slice alone.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Cp>That is what \u003Ca href=\"https:\u002F\u002Fxzero.media\u002Fatlas\u002Fapps\u002Fcontextkeep\">Contextkeep\u003C\u002Fa> is built to be: Atlas’s project knowledge plane and governed skills OS — the operating layer between consumer chat tools and the systems of record contractors already run. Mid-market GCs and specialty trades already experimenting with Claude Skills and “chat with the job folder” are the natural wedge: enough AI hunger to hurt, enough confidentiality pressure that bans alone will not hold.\u003C\u002Fp>\n\u003Cp>Scope the cut in a \u003Ca href=\"\u002Fservices\u002Fsolution-definition-sprint\">Solution Definition Sprint\u003C\u002Fa>: which project, which document classes, which skills, which approvers, which residency and retention rules, which audit export path.\u003C\u002Fp>\n\u003Cp>See \u003Ca href=\"\u002Findustries\u002Faec-built-environment\">AEC and built environment\u003C\u002Fa>, explore \u003Ca href=\"https:\u002F\u002Fxzero.media\u002Fatlas\u002Fapps\u002Fcontextkeep\">Contextkeep on the Atlas\u003C\u002Fa>, or \u003Ca href=\"\u002Fcontact\">bring us the paste problem IT cannot ban away\u003C\u002Fa>.\u003C\u002Fp>\n","A project knowledge plane: cited answers and governed skills, not personal AI paste","Contextkeep turns project drawings, contracts and specs into cited retrieval and permissioned skill runs with an audit trail on every action.","industry-applications",[13,14,15,16],"aec","knowledge-retrieval","ai-governance","evidence","xzero-media-editorial","2026-09-26T00:00:00.000Z",2026,9,3,"published",false,{"id":25,"slug":26,"body":27,"html":28,"title":29,"description":30,"category":11,"tags":31,"author":17,"date":35,"year":19,"month":36,"quarter":21,"status":22,"featured":23},"2026\u002F07\u002Findustry-applications\u002Fai-model-governance-as-an-application","ai-model-governance-as-an-application","\nMost enterprises now have an AI policy. Far fewer have an AI governance **system**. The policy says every model must be inventoried, evaluated, approved and monitored. In practice, the inventory is a spreadsheet, the evaluations are in notebooks, approvals happen in email and monitoring depends on whoever built the model.\n\nThat works for five models. It fails at fifty, and it fails immediately when an auditor or supervisor asks for evidence.\n\n## The workflow behind “AI governance”\n\nThe **AI governance** family in the Atlas treats governance as an operational workflow with a system of record:\n\n1. **Register.** Every model and AI use case gets an owner, a purpose, a risk tier, its data sources and where it is deployed. That includes vendor models, LLM features and internal models.\n2. **Evaluate.** Structured evaluations against defined criteria: accuracy, robustness, bias and fairness, and for LLM features, groundedness and safety. Results are stored as evidence, not screenshots.\n3. **Approve.** Deployment requests route through the right reviewers, such as model risk, security, the business owner and compliance, based on the risk tier. Every decision is recorded.\n4. **Monitor.** Production behaviour is tracked against thresholds. Drift and incidents raise cases with owners.\n5. **Evidence.** Packs for internal audit, the board or supervisors are generated from the record.\n\n## Where AI helps inside the governance application\n\nIt sounds recursive, but it's useful:\n\n- **Summarization** of model documentation and evaluation results for reviewers\n- **Classification** of new use cases into risk tiers, as a suggestion for a human to confirm\n- **Evaluation assistance**, generating test cases and red-team prompts for LLM features\n- **Drafting** evidence-pack narratives from structured records\n\nEvery one of these is a draft for a human. The approval decision is never automated.\n\n## Who uses it\n\n- **Head of AI and the AI platform team:** keep the portfolio visible and deployable.\n- **Model risk managers:** run reviews with consistent criteria.\n- **Risk and compliance officers:** answer supervisors and auditors from one record.\n- **CIO, CDO and CDAO:** see where AI is used, by whom, and at what risk.\n\n## Integrations that matter\n\nModel registries and ML platforms, CI\u002FCD pipelines (so deployment approval is a real gate rather than a formality), the identity provider for reviewer roles, ticketing, and data catalogues for lineage.\n\n## Controls designed in\n\n- Segregation between model owner and approver\n- An immutable decision history\n- Required evidence before approval can proceed\n- Periodic re-review based on risk tier and staleness\n- Role-based access to sensitive evaluation data\n\n## Why it belongs in financial services first\n\nBanks and insurers already run model risk management for credit and pricing models. Generative AI has multiplied the number of “models” and blurred their edges. A governance application extends existing discipline to the new portfolio instead of creating a parallel process.\n\nThe same foundation applies across enterprise operations, government and any organization preparing for AI-specific regulation.\n\n## Starting point\n\nThe fastest start is to take one line of business's AI inventory and move it into the application, with the approval workflow switched on for new deployments only. The [Solution Definition Sprint](\u002Fservices\u002Fsolution-definition-sprint) scopes the delta: your risk tiers, reviewers, evaluation criteria and integrations.\n\nSee the [financial services](\u002Findustries\u002Ffinancial-services) page, search the [Atlas](https:\u002F\u002Fxzero.media\u002Fatlas), or [bring us your AI inventory](\u002Fcontact).\n","\u003Cp>Most enterprises now have an AI policy. Far fewer have an AI governance \u003Cstrong>system\u003C\u002Fstrong>. The policy says every model must be inventoried, evaluated, approved and monitored. In practice, the inventory is a spreadsheet, the evaluations are in notebooks, approvals happen in email and monitoring depends on whoever built the model.\u003C\u002Fp>\n\u003Cp>That works for five models. It fails at fifty, and it fails immediately when an auditor or supervisor asks for evidence.\u003C\u002Fp>\n\u003Ch2>The workflow behind “AI governance”\u003C\u002Fh2>\n\u003Cp>The \u003Cstrong>AI governance\u003C\u002Fstrong> family in the Atlas treats governance as an operational workflow with a system of record:\u003C\u002Fp>\n\u003Col>\n\u003Cli>\u003Cstrong>Register.\u003C\u002Fstrong> Every model and AI use case gets an owner, a purpose, a risk tier, its data sources and where it is deployed. That includes vendor models, LLM features and internal models.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Evaluate.\u003C\u002Fstrong> Structured evaluations against defined criteria: accuracy, robustness, bias and fairness, and for LLM features, groundedness and safety. Results are stored as evidence, not screenshots.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Approve.\u003C\u002Fstrong> Deployment requests route through the right reviewers, such as model risk, security, the business owner and compliance, based on the risk tier. Every decision is recorded.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Monitor.\u003C\u002Fstrong> Production behaviour is tracked against thresholds. Drift and incidents raise cases with owners.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Evidence.\u003C\u002Fstrong> Packs for internal audit, the board or supervisors are generated from the record.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Ch2>Where AI helps inside the governance application\u003C\u002Fh2>\n\u003Cp>It sounds recursive, but it&#39;s useful:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Summarization\u003C\u002Fstrong> of model documentation and evaluation results for reviewers\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Classification\u003C\u002Fstrong> of new use cases into risk tiers, as a suggestion for a human to confirm\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Evaluation assistance\u003C\u002Fstrong>, generating test cases and red-team prompts for LLM features\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Drafting\u003C\u002Fstrong> evidence-pack narratives from structured records\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Every one of these is a draft for a human. The approval decision is never automated.\u003C\u002Fp>\n\u003Ch2>Who uses it\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Head of AI and the AI platform team:\u003C\u002Fstrong> keep the portfolio visible and deployable.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Model risk managers:\u003C\u002Fstrong> run reviews with consistent criteria.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Risk and compliance officers:\u003C\u002Fstrong> answer supervisors and auditors from one record.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>CIO, CDO and CDAO:\u003C\u002Fstrong> see where AI is used, by whom, and at what risk.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Integrations that matter\u003C\u002Fh2>\n\u003Cp>Model registries and ML platforms, CI\u002FCD pipelines (so deployment approval is a real gate rather than a formality), the identity provider for reviewer roles, ticketing, and data catalogues for lineage.\u003C\u002Fp>\n\u003Ch2>Controls designed in\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>Segregation between model owner and approver\u003C\u002Fli>\n\u003Cli>An immutable decision history\u003C\u002Fli>\n\u003Cli>Required evidence before approval can proceed\u003C\u002Fli>\n\u003Cli>Periodic re-review based on risk tier and staleness\u003C\u002Fli>\n\u003Cli>Role-based access to sensitive evaluation data\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Why it belongs in financial services first\u003C\u002Fh2>\n\u003Cp>Banks and insurers already run model risk management for credit and pricing models. Generative AI has multiplied the number of “models” and blurred their edges. A governance application extends existing discipline to the new portfolio instead of creating a parallel process.\u003C\u002Fp>\n\u003Cp>The same foundation applies across enterprise operations, government and any organization preparing for AI-specific regulation.\u003C\u002Fp>\n\u003Ch2>Starting point\u003C\u002Fh2>\n\u003Cp>The fastest start is to take one line of business&#39;s AI inventory and move it into the application, with the approval workflow switched on for new deployments only. The \u003Ca href=\"\u002Fservices\u002Fsolution-definition-sprint\">Solution Definition Sprint\u003C\u002Fa> scopes the delta: your risk tiers, reviewers, evaluation criteria and integrations.\u003C\u002Fp>\n\u003Cp>See the \u003Ca href=\"\u002Findustries\u002Ffinancial-services\">financial services\u003C\u002Fa> page, search the \u003Ca href=\"https:\u002F\u002Fxzero.media\u002Fatlas\">Atlas\u003C\u002Fa>, or \u003Ca href=\"\u002Fcontact\">bring us your AI inventory\u003C\u002Fa>.\u003C\u002Fp>\n","AI model governance should be an application, not a policy document","Model inventory, evaluation, deployment approval and monitoring as one governed workflow, so AI governance produces evidence instead of meetings.",[15,32,33,16,34],"financial-services","evaluation","risk","2026-07-02T00:00:00.000Z",7,{"id":38,"slug":39,"body":40,"html":41,"title":42,"description":43,"category":44,"tags":45,"author":17,"date":48,"year":19,"month":49,"quarter":50,"status":22,"featured":23},"2026\u002F06\u002Fai-in-production\u002Fevaluation-and-guardrails-before-production","evaluation-and-guardrails-before-production","\nTraditional software has a comforting property: the same input produces the same output, so a passing test suite means something. AI features don't behave that way. The same prompt can produce different answers, a model upgrade can change behaviour silently, and a content change can make a previously correct answer wrong.\n\nSo AI features need their own form of testing, **evaluation**, and it has to be a delivery gate, not a one-off exercise before a demo.\n\n## Four layers of evaluation\n\n**1. Task quality.** Does the feature do its job? For extraction, field-level accuracy against labelled documents. For classification, precision and recall per class. For summarization, coverage of required facts. For retrieval, whether the right sources come back.\n\n**2. Groundedness.** For anything generated from sources, is every claim supported by the retrieved material, and are citations correct? An ungrounded answer is a defect even when it happens to be true.\n\n**3. Safety and policy.** Does the feature refuse what it should: out-of-scope questions, requests for data the user can't access, instructions hidden in documents (prompt injection)? Does it avoid prohibited content and claims?\n\n**4. Regression.** Every change to prompts, models, retrieval settings or content is re-evaluated against the same test sets, and the results are compared with the last accepted baseline.\n\n## Building the test sets\n\nGood test sets come from the workflow, not from the vendor:\n\n- real questions and documents from the pilot, anonymized where needed\n- edge cases the business owner worries about\n- known-hard cases collected from production feedback\n- adversarial cases: injection attempts, ambiguous requests, missing data\n\nEvery case has an expected outcome defined by a person who owns the domain.\n\n## Guardrails in the application\n\nEvaluation tells you how the feature behaves. Guardrails constrain it in production:\n\n- **Grounding rules:** answer only from retrieved, authorized sources, or say you don't know.\n- **Output validation:** structured outputs checked against schemas and business rules before use.\n- **Allow-lists:** an AI can only reference entities that exist. It can't invent a product, a customer or a case number.\n- **Human checkpoints:** consequential outputs are drafts until a person accepts them.\n- **Untrusted-input handling:** document and user content is treated as data, never as instructions.\n- **Fallbacks:** if the model is unavailable or uncertain, the workflow continues deterministically.\n\n## Monitoring after launch\n\nIn production, keep measuring: acceptance and edit rates on AI drafts, user flags, drift in evaluation scores on a sampled stream, and latency and cost. Those signals feed the next round of test cases.\n\n## How the factory handles it\n\nIn our architecture, evaluation sits alongside the automated test suite. Every application foundation that includes AI features ships with an evaluation harness, and a Production Sprint doesn't close until the agreed evaluation thresholds are met. It's one of the [quality gates](\u002Fservices\u002Fai-production-sprint) we use to decide whether something is done.\n\nRelated: [from AI pilot to production application](\u002Fblog\u002Ffrom-ai-pilot-to-production-application) and [AI model governance as an application](\u002Fblog\u002Fai-model-governance-as-an-application).\n\nHave a pilot that's never been evaluated properly? [Bring it to us](\u002Fcontact).\n","\u003Cp>Traditional software has a comforting property: the same input produces the same output, so a passing test suite means something. AI features don&#39;t behave that way. The same prompt can produce different answers, a model upgrade can change behaviour silently, and a content change can make a previously correct answer wrong.\u003C\u002Fp>\n\u003Cp>So AI features need their own form of testing, \u003Cstrong>evaluation\u003C\u002Fstrong>, and it has to be a delivery gate, not a one-off exercise before a demo.\u003C\u002Fp>\n\u003Ch2>Four layers of evaluation\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>1. Task quality.\u003C\u002Fstrong> Does the feature do its job? For extraction, field-level accuracy against labelled documents. For classification, precision and recall per class. For summarization, coverage of required facts. For retrieval, whether the right sources come back.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>2. Groundedness.\u003C\u002Fstrong> For anything generated from sources, is every claim supported by the retrieved material, and are citations correct? An ungrounded answer is a defect even when it happens to be true.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>3. Safety and policy.\u003C\u002Fstrong> Does the feature refuse what it should: out-of-scope questions, requests for data the user can&#39;t access, instructions hidden in documents (prompt injection)? Does it avoid prohibited content and claims?\u003C\u002Fp>\n\u003Cp>\u003Cstrong>4. Regression.\u003C\u002Fstrong> Every change to prompts, models, retrieval settings or content is re-evaluated against the same test sets, and the results are compared with the last accepted baseline.\u003C\u002Fp>\n\u003Ch2>Building the test sets\u003C\u002Fh2>\n\u003Cp>Good test sets come from the workflow, not from the vendor:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>real questions and documents from the pilot, anonymized where needed\u003C\u002Fli>\n\u003Cli>edge cases the business owner worries about\u003C\u002Fli>\n\u003Cli>known-hard cases collected from production feedback\u003C\u002Fli>\n\u003Cli>adversarial cases: injection attempts, ambiguous requests, missing data\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Every case has an expected outcome defined by a person who owns the domain.\u003C\u002Fp>\n\u003Ch2>Guardrails in the application\u003C\u002Fh2>\n\u003Cp>Evaluation tells you how the feature behaves. Guardrails constrain it in production:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Grounding rules:\u003C\u002Fstrong> answer only from retrieved, authorized sources, or say you don&#39;t know.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Output validation:\u003C\u002Fstrong> structured outputs checked against schemas and business rules before use.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Allow-lists:\u003C\u002Fstrong> an AI can only reference entities that exist. It can&#39;t invent a product, a customer or a case number.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Human checkpoints:\u003C\u002Fstrong> consequential outputs are drafts until a person accepts them.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Untrusted-input handling:\u003C\u002Fstrong> document and user content is treated as data, never as instructions.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Fallbacks:\u003C\u002Fstrong> if the model is unavailable or uncertain, the workflow continues deterministically.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Monitoring after launch\u003C\u002Fh2>\n\u003Cp>In production, keep measuring: acceptance and edit rates on AI drafts, user flags, drift in evaluation scores on a sampled stream, and latency and cost. Those signals feed the next round of test cases.\u003C\u002Fp>\n\u003Ch2>How the factory handles it\u003C\u002Fh2>\n\u003Cp>In our architecture, evaluation sits alongside the automated test suite. Every application foundation that includes AI features ships with an evaluation harness, and a Production Sprint doesn&#39;t close until the agreed evaluation thresholds are met. It&#39;s one of the \u003Ca href=\"\u002Fservices\u002Fai-production-sprint\">quality gates\u003C\u002Fa> we use to decide whether something is done.\u003C\u002Fp>\n\u003Cp>Related: \u003Ca href=\"\u002Fblog\u002Ffrom-ai-pilot-to-production-application\">from AI pilot to production application\u003C\u002Fa> and \u003Ca href=\"\u002Fblog\u002Fai-model-governance-as-an-application\">AI model governance as an application\u003C\u002Fa>.\u003C\u002Fp>\n\u003Cp>Have a pilot that&#39;s never been evaluated properly? \u003Ca href=\"\u002Fcontact\">Bring it to us\u003C\u002Fa>.\u003C\u002Fp>\n","Evaluation and guardrails: how to test AI features before production","AI features need evaluation as a delivery gate, just like tests: test sets, groundedness checks, safety checks and regression on every change.","ai-in-production",[33,46,15,47],"production","human-in-the-loop","2026-06-30T00:00:00.000Z",6,2,1791555300969]