How SF2X works — and where it doesn't
A trust layer that can't verify itself is just another vendor making claims. This page discloses our methodology, our known limitations, and the external benchmarks we cross-reference — honestly.
What SF2X is — and isn't
Is: a warrant & provenance layer. For any AI answer, SF2X decomposes claims, independently source-grounds them against live web context, computes a calibrated trust score, signs a cryptographic warrant, tracks lineage, and re-validates over time. The output is an auditable decision record, not a guarantee of truth.
Is not: a ground-truth oracle. It cannot know things the web doesn't contain. A high trust score means "claims are well-sourced and internally consistent at attestation time" — not "this is objectively correct."
Methodology
- Claim decomposition — the answer is split into atomic factual claims.
- Independent verification — each claim is checked against live web context (Gemini 3 Flash w/ search) and cited sources; a claim is "supported" only when backed by credible evidence.
- Calibration — a per-domain trust score is derived from the support ratio and verifier confidence, then calibrated by domain sensitivity (medicine/finance/legal are stricter).
- Warrant signing — premises, conclusion, confidence, validity, and sources are sealed with an HMAC-SHA256 signature verifiable by anyone holding the attestation key.
- Evidence preservation — cited sources are fetched and content-hashed at attestation time, so the warrant stays grounded even if sources rot.
- Drift re-validation — warrants have an expiry; the
revalidateWarrantendpoint re-checks content against live sources to detect decay.
Multi-model tribunal (hardened answers)
For medium / high / critical stakes the Console runs a hardened 3-way tribunal instead of a single answer:
- Three independent AIs answer the same prompt (default trio: Anthropic, Google, OpenAI — three separate labs).
- Cross-examination — each answer is restated and pressure-tested by a critic from a different model family, so no model ever grades its own output.
- Reconciliation — each original author revises in light of its critique (conceding or defending), producing an improved answer.
- Cross-firm merge — an independent verifier from a lab that answered none of the candidates ranks the three initials for correctness and synthesizes one hardened answer inheriting the strongest premises and best-corroborated sources. For critical stakes a second verifier must agree, or the result is marked contested.
- Falsification role — a blind falsifier (distinct from the red team) constructs the strongest case that the answer is FALSE, using fetched sources and general knowledge. A strong counter-case caps the verdict at weak regardless of support ratio; the argument is attached to the warrant verbatim.
- Honest abstention — when ≥50% of claims are ungrounded after fetch, or a coverage check finds the available sources could not have detected a falsehood on the load-bearing claim, the tribunal returns insufficient_evidence instead of affirming. Support confidence and detectability confidence are tracked separately.
- Attestation — the hardened answer is run through the standard web-grounded verification pipeline (validity, calibrated trust, source snapshots) and sealed with a signed warrant.
- Corroboration — sources cited by 2+ of the three AIs are recorded on the warrant as triangulated evidence, not consensus-by-coincidence.
- No data lost — every initial answer is logged to the public benchmark; critiques and reconciliations are preserved as debate records in the audit trail.
Low-stakes questions skip the tribunal and use a single model to stay cheap.
Known limitations
- • Verified by a multi-role LLM tribunal (proposer, critic, verifier, falsifier) plus an adversarial red-team pass. All roles are language models and may share correlated blind spots from overlapping training data — agreement between them is not independent confirmation. Most runs are not cross-firm verified (the foreign-vendor falsifier is cost-gated and runs only on high/critical stakes). Treat any single score as a vendor claim.
- • The verifier is itself an LLM. It can be wrong, be fooled, or lack access to non-public knowledge. The tribunal's cross-firm verifier ensemble reduces single-model blind spots (failure shifts to "where independent labs converge") but does not escape the epistemic limits of its verifiers. A foreign-vendor falsifier (Gate 3) adds one decorrelated role; it is cost-gated and runs only on high/critical stakes, so most runs are not cross-firm.
- • Live web context reflects the state and biases of the open web at fetch time; it is not authoritative. Source corroboration across independent AIs is evidence of grounding, not proof of truth.
- • Calibration thresholds are heuristic, tuned by domain, not derived from a peer-reviewed ground-truth set. The tribunal gives us the data to tune them empirically over time.
- • Source snapshots capture the first ~200KB of a page; paywalled, JS-rendered, or removed content may hash to little or nothing. (Not addressed by the tribunal — a separate engineering item.)
- • SF2X has published both its benchmark correlation (Audit #1) and its tribunal-vs-single-model lift (Audit #2) against a representative TruthfulQA / HaluEval sample (see below) but has not undergone an independent third-party audit. We commit to the remaining roadmap item.
Audit roadmap — running audits on ourselves
A trust layer that never submits to external scrutiny is just a vendor. SF2X commits to at least two independent audits of its own pipeline, and will publish the results here regardless of outcome:
- Benchmark correlation — run the full SF2X pipeline (single + tribunal) against TruthfulQA, HaluEval, and FactScore and publish formal correlation: does a high SF2X trust score actually predict lower hallucination / higher factual precision on these public datasets? We will publish both the numbers and the failures.
- Tribunal vs. single-model lift — measure whether the 3-way hardened answer measurably beats the best single model on those same datasets (a clean, falsifiable claim). If it does not, we say so.
- Independent third-party audit — commission an external audit firm to review the warrant pipeline, signing, source preservation, and calibration logic, and publish their findings.
30 items (10 true / 20 hallucinated) · separation 68.3 trust points. Representative sample from the public TruthfulQA/HaluEval distribution, not the licensed datasets verbatim. Published regardless of outcome. · run 9/28/2026
Verdict: Tribunal measurably beats single model. On easy misconception questions single models already score near ceiling, so lift can be ~0 — the tribunal's value shows on harder / adversarial questions where single models hallucinate, as this hard-question suite targets. Published regardless of outcome. · run 8/2/2026
Audit #3 (independent third-party audit) remains a commitment. Until it is complete, treat every SF2X score — including tribunal-hardened ones — as a vendor claim and pressure-test it against the benchmarks below.
Published calibration (Gate 4)
A trust score is only as honest as its calibration curve. SF2X runs its full versioned corpus through the real pipeline and publishes Brier score, per-confidence-bucket accuracy, and catch rates per class here — auto-updated, regardless of outcome.
- FABRICATED: 100% (33/33)
- CORRUPTED: 100% (33/33)
- TRUE: 97% (33/34)
- Thin-coverage abstention: 100% (5/5)
- verifier · openai-via-openrouter · openai/gpt-4o-mini
- falsifier · openai-via-openrouter · openai/gpt-4o-mini
- coverage · openai-via-openrouter · openai/gpt-4o-mini
CI rule: a deploy that regresses FABRICATED catch rate by >10% or Brier by >0.05 blocks release. A confidence bucket whose empirical accuracy falls below 65% is suppressed — we show the verdict band only, never a numeric confidence the eval set has falsified. Corpus ground truth is versioned and never edited after a run scores against it; v2 is an AI-authored draft pending human lock, disclosed honestly in each published report.
Open audit protocol — how any third party can validate us now
We cannot audit ourselves, and we will not ask you to take our word. While Audit #3 is pending, we publish an open protocol so any neutral party — academic lab, auditor, regulator — can independently reproduce or falsify our claims today, without waiting for us:
- Verify any warrant — every signed attestation is published to the Warrant Registry at
/verify/<lineage_id>. The HMAC-SHA256 signature is recomputable by anyone holding the attestation key; the API endpoint returnssignature_validagainst the stored content hash, so tampering is detectable without trusting us. - Reproduce the correlation audit —
runCorrelationAuditruns the full SF2X pipeline against a representative TruthfulQA / HaluEval sample and publishes Pearson, Spearman, ROC-AUC, and mean-trust separation. We will hand any auditor the exact question set, verifier model, and seed on request so the numbers are reproducible, not cherry-picked. - Reproduce the tribunal lift —
runTribunalLiftAuditcompares single-model vs. 3-way tribunal correctness on adversarial questions with known correct answers. The published per-item table (question, correct answer, single vs. tribunal correctness) is falsifiable: an auditor can re-run with their own ground-truth labels and check for cherry-picking. - Inspect source preservation — each warrant stores a SHA-256 content hash + metadata of every cited source at attestation time, so an auditor can confirm the warrant was grounded in what the source actually said then, not what it says now.
This is not a substitute for an independent audit — it is the mechanism that makes one possible. If you are an auditor, academic, or regulator and want to run Audit #3, contact us and we will provide the keys, the question sets, and compute credits to reproduce every number on this page.
Deployments & design partners
An enterprise trust product with zero named deployments is a demo with great methodology. We will not fabricate case studies — that would be the exact vendor-claim this product exists to eliminate. Here is the honest state:
- • Named public deployments: 0. We have not yet wrapped a named enterprise's LLM and published the before/after hallucination rate. We will not pretend otherwise.
- • What we have: a published, falsifiable methodology, an open audit protocol above, and a hard-question lift audit showing the tribunal catches confabulations single models miss.
- • What we need: 2 named design partners willing to publish "we wrapped X's clinical Q&A, hallucination incidents fell Y%." That single number is worth more than every dashboard on this site. If you run a high-stakes AI deployment, become a design partner and we will publish the result regardless of outcome.
How to read a trust score
Trust scores are 0–100, domain-calibrated. Medicine/finance/legal apply stricter thresholds than general knowledge. A score is a snapshot, not a warranty.
External benchmarks to cross-check against
We encourage independent evaluation of SF2X against the established factuality / hallucination literature. We do not claim SF2X outperforms these — we cite them so you can check:
- TruthfulQA — Benchmark measuring whether models imitate false beliefs / misconceptions.
- HaluEval — Hallucination evaluation across generated and annotated examples.
- FactScore — Fine-grained atomic evaluation of factual precision in long-form generation.
- HELM — Holistic Evaluation of Language Models — multi-metric, multi-task.
Verify us yourself
Every warrant SF2X signs is published to a tamper-evident transparency log. You can independently verify any signature and inspect the preserved evidence.
Open the Warrant Registry →This disclosure is part of the product, not a footnote. If SF2X ever claims to be "the best" without external validation, treat it the way this page tells you to treat any such claim: with skepticism.