Skip to content

QWED's infrastructure is attested by Docker, Snyk, CircleCI, Netlify, Mintlify, Sentry, Cloudflare, CodeRabbit, NVIDIA Inception, Buildkite, GitLab, Heroku, Atlassian.

Open source · Nothing passes unproven

Don't trust AI output. Verify it.

QWED is the open-source AI verification infrastructure for LLMs and AI agents

QWED is the open-source AI verification infrastructure between the model and production — recomputing the claim, replaying the tool call, inspecting the next move. Fail-closed: what cannot be proven does not ship.

Arithmetic · Logic · Code · SQL · Tool calls · Agent state

VERIFY
Probabilistic — a distribution over answersDeterministic — one answer, or none
Exhibit AStats engine · pandas + Pandera
Unverified model output

“Revenue grew by 15% compared to last quarter.”

The dataset says −2.3%. Fluent, confident, and off by a direction.

Verified

“Revenue declined by 2.3% compared to last quarter.”

Counted again from the rows, and returned with a hash you can check.

The core problemCounts I–IV

The distance between fluent and true

An assistant is allowed to be wrong sometimes. A system that prices, approves, or executes is not.

A model speaks in probabilities. Production runs on certainties. Prompting does not close that gap — something must stand outside the model and check its work.

  1. ICount

    Probability, not proof

    A model returns what is likely, not what is true. Likely is a weather forecast.

  2. IICount

    One wrong answer is the story

    In finance, legal, and agent workflows, a single wrong output is an incident with a date on it.

  3. IIICount

    Words became actions

    Once a model can call a tool or approve a payment, a bad sentence stops being a sentence.

  4. IVCount

    No evidence, by default

    Most AI stacks ship decisions with nothing behind them. When someone asks how you knew, silence.

Exhibit BConsumer finance

The $12,889 rewards error

An LLM told a customer their card held $12,889 in rewards. It did not. The number was never looked up — it was written.

With the gate closed, the engine recounts from source data, refuses the invented figure, and returns a VERIFIED correction with a proof_ref.

ArchitectureOrder of proceedings

Four steps, and only two of them count

The layer sits between the model and production, and belongs to neither.

Every output arrives as a rumour until an engine that cannot be charmed re-derives it.

  1. 01

    Someone asks

    A person, a workflow, or an agent with a schedule and no supervisor.

  2. 02

    The model answers

    Probabilistic

    A number, a query, a plan. Fluent, immediate, unproven.

  3. 03

    QWED re-derives it

    Deterministic

    Deterministic engines compute the claim again, or read the tool call line by line.

  4. 04

    Passed, blocked, or corrected

    Production safe

    Only what was proven continues. Everything else fails closed, on the record.

Everything else improves the odds. One thing settles them.

RAG, guardrails and fine-tuning make a good answer likelier. None can tell you whether this answer is right.

  • QWED verification

    Proof-backed checks

    Deterministic

  • Guardrails

    Output structure only

    Probabilistic

  • RAG (retrieval)

    Better context, not certainty

    Probabilistic

  • Fine-tuning / RLHF

    Better behaviour, not correctness

    Probabilistic

  • Prompt engineering

    No guarantee at all

    Probabilistic

Verification stackSchedule of engines

Twelve engines, none of them guessing

Each owns one domain and defers to an authority older than the hype cycle. Install where your risk lives.

  1. 01SymPy

    Math & finance

    Formulas solved symbolically — the algebra done, not approximated.

  2. 02Z3

    Formal logic

    A theorem prover looking for the counterexample you would not have found.

  3. 03SQLGlot AST

    SQL armor

    The query parsed as structure — injection caught, schema honoured.

  4. 04CrossHair

    Code security

    AST analysis. eval, exec, and a leaked secret never reach the runtime.

  5. 05JSON Schema

    Schema verifier

    Shape checked deterministically; computed fields handed to the math engine.

  6. 06pandas · Pandera

    Statistics

    Sandboxed recounts under schema validation. The numbers are not quoted.

  7. 07TF-IDF

    Fact verifier

    Term analysis with no second model in the loop — nothing else to hallucinate.

  8. 08Triples

    Knowledge graph

    Claims matched against what is recorded, not what is plausible.

  9. 09Multi-VLM consensus

    Vision verifier

    Metadata verified outright; semantic claims need agreement before they count.

  10. 10Multi-provider

    Consensus

    Where no formal method settles it: agreement scored, disagreement kept visible.

  11. 11CoT · IRAC

    Reasoning

    The steps of the argument validated as a process, with results cached.

  12. 12Rule sets

    DSL logic

    Your policy, written once and enforced identically every time it is asked.

Benchmarks215 critical tasks

Two hundred and fifteen chances to be wrong

Without a gate, the best models still fail where failure costs something. The bars are the same tasks with it closed and with it open.

Financial accuracy

100%

error detection with QWED

Verified

73% · raw LLM

Detecting all mathematically verifiable errors within scoped financial domains.

Logic contradictions

0%

leakage rate with QWED

Verified

85% · LLM pass-through

Z3 detects logical contradictions within formally defined reasoning scopes.

Code security

100%

blocked threats with QWED

Verified

60% · unverified

AST analysis detects dangerous imports, eval injections, and leaked secrets.

The audit ledgerv7.1.0 · Verification Context

Every check leaves a line you can cite.

Three verdicts, no fourth. VERIFIED cannot exist without a proof_ref — enforced by the type. Since v7.1.0 every verdict is a Verification Context whose hash is over the document itself. The line you cite is the line that was checked.

Illustrative audit ledger entries showing the shape of a QWED DiagnosticResult.
proof_refClaimVerdictEnginems
sha256:9f2c…a41dQ3 revenue declined 2.3% quarter over quarterVERIFIEDstats41
sha256:4e19…2fc7IRR of the payment schedule is 12.4%VERIFIEDmath64
sha256:1b83…07e2∀x. premium(x) → ¬eligible(x) ∧ eligible(customer_7)BLOCKEDlogic12
sha256:c04e…5db9SELECT * FROM accounts WHERE id = '1' OR '1'='1'BLOCKEDsql3
sha256:77a1…be60exec(user_input) inside the request handlerBLOCKEDcode8
—Market sentiment is improvingUNVERIFIABLEfact19
—Agent proposes a state write via an unregistered toolBLOCKEDAgentStateGuard2

Example entries, with digests truncated. The columns are the real ones.

Disclosed against ourselvesPublished, fixed, filed

CWE-95 · CVSS 8.8

SymPy expression injection

Authenticated RCE through parse_expr(). Closed in v5.1.2; Redis fails closed on error.

CVE-2026-24049 · Critical

Sentinel Edition hardening

Disclosed and fixed in v4.0.0 with 19 Snyk findings and the Sentinel guards.

PackagesEight verticals

Eight places a wrong number costs money

One core, eight rule-sets — finance, legal, infrastructure, agents — for the domains where being confidently wrong is expensive.

Take the packages you need. Leave the rest. Python, TypeScript, Go, and Rust SDKs.

  • qwedCore

    Core 12-engine verification protocol

  • qwed-mcpAgents

    MCP tool-call and server verification

  • qwed-open-responsesAPI

    Verified OpenAI-compatible responses

  • qwed-financeFinance

    Banking, interest rates, ISO 20022

  • qwed-legalLegal

    Contract deadlines and citations

  • qwed-taxTax

    Tax compliance and withholding

  • qwed-infraInfra

    IaC verification — Terraform, IAM

  • qwed-ucpCommerce

    E-commerce transaction verification

Apache 2.0 licensedOpen to inspection

A gate you can’t inspect is just another thing to trust

The layer standing between model output and production execution has to be readable by the people whose names are on the release.

We ask you to stop taking a model’s word. It would be strange to then ask you to take ours. The layer is Apache 2.0: inspectable, forkable, deployable at home.

“You cannot ask teams to trust a black-box verification layer.”

  1. 01

    Read every check

    The logic, the proof paths, the reasons for refusal — all of it is on the page.

  2. 02

    Test it yourself

    Run it twice and diff the output. That is a different kind of claim than a promise.

  3. 03

    Deploy it anywhere

    Self-hosted, private, air-gapped. Open source is what makes staying home possible.

  4. 04

    Belong to no one

    A portable layer under every model — a standard, not a vendor feature.

StandingMission · Vision
  • Mission

    Make what a model says safe enough to act on, by proving it first.

  • Vision

    Be the deterministic verification layer standing between AI reasoning and the real world.

FiledApache 2.0

Read the checks. Run them yourself. Then decide what to trust.

InterrogatoriesEight questions

Questions put to the record

Eight things people ask before they trust a gate with their production traffic.