Control Credibility
Are stated controls actually operating, or is the policy document carrying the load? Do observed system behaviors match what governance documentation asserts?
The Framework
Evidence over attestation. Structure over checklists.
Most AI compliance reviews lean on policy PDFs, vendor questionnaires, and management interviews. The result is gaps that surface later — in front of an auditor, a regulator, or an LP. Kaptrix replaces that with a structured, repeatable framework that separates what is implemented from what is documented, and turns the gap between the two into the central object of the review.
Six dimensions we evaluate
Are stated controls actually operating, or is the policy document carrying the load? Do observed system behaviors match what governance documentation asserts?
Which models, sub-processors, and third-party AI services sit in the data path? How much regulated data crosses those boundaries, and are the contracts, DPAs, and SCCs in place to support it?
Does data handling match the applicable regime — PII, PHI, financial, biometric, cross-border? Are training data, inference inputs, and outputs scoped, retained, and disposed of in line with the law that actually applies?
When the system behaves badly — bias, drift, hallucination, leakage — is there a named owner, an incident pathway, and a defensible record of action? Is oversight proportionate to the system's risk tier?
Are monitoring, logging, evaluation, and human-in-the-loop controls instrumented in production, or only described in policy? Will the evidence exist when an auditor opens the request list?
What remains unverified against the applicable frameworks? Where is compliance debt hiding, and which unknowns will a regulator press on first?
How we think about scoring
Evidence, not attestation
Every score is tied to an artifact — a control test, a log sample, a policy excerpt, a contract clause, a model card. If it cannot be cited, it cannot be scored.
Missing evidence is a finding
Unknowns are surfaced, not assumed compliant. A dense gap register is itself a signal — and the basis of the remediation plan.
Non-compensatory by design
Strength on policy cannot offset weakness in operating effectiveness. A pristine AI governance charter does not neutralize an unmonitored model in production. Guardrails catch what an average would hide.
Framework-aware, not framework-locked
The same control evidence is mapped against the regimes that actually apply to the system — NIST AI RMF, ISO/IEC 42001, EU AI Act, SOC 2, HIPAA, GDPR, sectoral rules. One assessment, multiple framework views, no double work.
Two reviewers, then calibration
Every evaluation is dual-scored. Where reviewers disagree, the reasoning is the deliverable — not a split-the-difference average.
How scoring works
The audit trail
Artifact
Policy, log, contract, model card
Proposal
Bounded ±0.5, named sub-criterion
Operator decision
Accept, reject, or revise
Score impact
Written with rationale
Base scoring is operator-owned
The AI does not assign base scores. This is an architectural constraint, not a configuration setting.
Each sub-criterion is scored 0.0–5.0 in 0.5 increments by a named reviewer, with rationale attached.
Rollup is failure-weighted
Tier-A signals carry the highest weight. Strong performance on lower-tier signals cannot offset weakness at Tier A.
Dimension scores roll up by mean. Composite scores apply failure weighting to claim integrity, regulated-data handling, and accountability.
Evidence generates proposals, not silent edits
Proposals are bounded — max ±0.5 per sub-criterion, no cross-dimension bleed — and require operator approval.
When an artifact contradicts, supports, augments, or exposes a gap, the engine produces a proposal naming the affected sub-criterion, the direction of pressure, the rationale, and the supporting artifacts.
Confidence is calculated independently
High score / low confidence and high score / high confidence are surfaced as distinct states.
Coverage of the model, source quality, recency, and signal consistency produce a confidence signal that qualifies the score. Confidence does not override the score — it tells the reader how much to trust it.
Every output is traceable
If a conclusion cannot be walked back to its source, it is not a conclusion the platform will present.
Artifact → proposal → operator decision → score impact. No hidden logic. No silent adjustments.
Surfaced as a caution — treat the score as provisional until coverage improves.
Audit-ready. Evidence coverage, recency, and consistency all support the conclusion.
What Kaptrix delivers
Dimensional scores mapped to the frameworks that apply, an evidence gap register, control-failure surfacing, and a clear, audit-ready recommendation.
Built for
Under NDA
Kaptrix is an AI compliance platform. The full scoring methodology — including sub-criteria, framework crosswalks, weighting tiers, and decision thresholds — is available under NDA.