Don't trust AI output. Verify it.
QWED is the open-source AI verification infrastructure for LLMs and AI agents
QWED is the open-source AI verification infrastructure between the model and production — recomputing the claim, replaying the tool call, inspecting the next move. Fail-closed: what cannot be proven does not ship.
Arithmetic · Logic · Code · SQL · Tool calls · Agent state
“Revenue grew by 15% compared to last quarter.”
The dataset says −2.3%. Fluent, confident, and off by a direction.
“Revenue declined by 2.3% compared to last quarter.”
Counted again from the rows, and returned with a hash you can check.
The distance between fluent and true
An assistant is allowed to be wrong sometimes. A system that prices, approves, or executes is not.
A model speaks in probabilities. Production runs on certainties. Prompting does not close that gap — something must stand outside the model and check its work.
- ICount
Probability, not proof
A model returns what is likely, not what is true. Likely is a weather forecast.
- IICount
One wrong answer is the story
In finance, legal, and agent workflows, a single wrong output is an incident with a date on it.
- IIICount
Words became actions
Once a model can call a tool or approve a payment, a bad sentence stops being a sentence.
- IVCount
No evidence, by default
Most AI stacks ship decisions with nothing behind them. When someone asks how you knew, silence.
The $12,889 rewards error
An LLM told a customer their card held $12,889 in rewards. It did not. The number was never looked up — it was written.
With the gate closed, the engine recounts from source data, refuses the invented figure, and returns a VERIFIED correction with a proof_ref.
Four steps, and only two of them count
The layer sits between the model and production, and belongs to neither.
Every output arrives as a rumour until an engine that cannot be charmed re-derives it.
- 01
Someone asks
A person, a workflow, or an agent with a schedule and no supervisor.
- 02
The model answers
ProbabilisticA number, a query, a plan. Fluent, immediate, unproven.
- 03
QWED re-derives it
DeterministicDeterministic engines compute the claim again, or read the tool call line by line.
- 04
Passed, blocked, or corrected
Production safeOnly what was proven continues. Everything else fails closed, on the record.
Everything else improves the odds. One thing settles them.
RAG, guardrails and fine-tuning make a good answer likelier. None can tell you whether this answer is right.
QWED verification
Proof-backed checks
Deterministic
Guardrails
Output structure only
Probabilistic
RAG (retrieval)
Better context, not certainty
Probabilistic
Fine-tuning / RLHF
Better behaviour, not correctness
Probabilistic
Prompt engineering
No guarantee at all
Probabilistic
Twelve engines, none of them guessing
Each owns one domain and defers to an authority older than the hype cycle. Install where your risk lives.
- 01SymPy
Math & finance
Formulas solved symbolically — the algebra done, not approximated.
- 02Z3
Formal logic
A theorem prover looking for the counterexample you would not have found.
- 03SQLGlot AST
SQL armor
The query parsed as structure — injection caught, schema honoured.
- 04CrossHair
Code security
AST analysis. eval, exec, and a leaked secret never reach the runtime.
- 05JSON Schema
Schema verifier
Shape checked deterministically; computed fields handed to the math engine.
- 06pandas · Pandera
Statistics
Sandboxed recounts under schema validation. The numbers are not quoted.
- 07TF-IDF
Fact verifier
Term analysis with no second model in the loop — nothing else to hallucinate.
- 08Triples
Knowledge graph
Claims matched against what is recorded, not what is plausible.
- 09Multi-VLM consensus
Vision verifier
Metadata verified outright; semantic claims need agreement before they count.
- 10Multi-provider
Consensus
Where no formal method settles it: agreement scored, disagreement kept visible.
- 11CoT · IRAC
Reasoning
The steps of the argument validated as a process, with results cached.
- 12Rule sets
DSL logic
Your policy, written once and enforced identically every time it is asked.
Two hundred and fifteen chances to be wrong
Without a gate, the best models still fail where failure costs something. The bars are the same tasks with it closed and with it open.
Financial accuracy
100%error detection with QWED
Verified
73% · raw LLM
Detecting all mathematically verifiable errors within scoped financial domains.
Logic contradictions
0%leakage rate with QWED
Verified
85% · LLM pass-through
Z3 detects logical contradictions within formally defined reasoning scopes.
Code security
100%blocked threats with QWED
Verified
60% · unverified
AST analysis detects dangerous imports, eval injections, and leaked secrets.
Every check leaves a line you can cite.
Three verdicts, no fourth. VERIFIED cannot exist without a proof_ref — enforced by the type. Since v7.1.0 every verdict is a Verification Context whose hash is over the document itself. The line you cite is the line that was checked.
| proof_ref | Claim | Verdict | Engine | ms |
|---|---|---|---|---|
| sha256:9f2c…a41d | Q3 revenue declined 2.3% quarter over quarter | VERIFIED | stats | 41 |
| sha256:4e19…2fc7 | IRR of the payment schedule is 12.4% | VERIFIED | math | 64 |
| sha256:1b83…07e2 | ∀x. premium(x) → ¬eligible(x) ∧ eligible(customer_7) | BLOCKED | logic | 12 |
| sha256:c04e…5db9 | SELECT * FROM accounts WHERE id = '1' OR '1'='1' | BLOCKED | sql | 3 |
| sha256:77a1…be60 | exec(user_input) inside the request handler | BLOCKED | code | 8 |
| — | Market sentiment is improving | UNVERIFIABLE | fact | 19 |
| — | Agent proposes a state write via an unregistered tool | BLOCKED | AgentStateGuard | 2 |
Example entries, with digests truncated. The columns are the real ones.
CWE-95 · CVSS 8.8
SymPy expression injection
Authenticated RCE through parse_expr(). Closed in v5.1.2; Redis fails closed on error.
CVE-2026-24049 · Critical
Sentinel Edition hardening
Disclosed and fixed in v4.0.0 with 19 Snyk findings and the Sentinel guards.
Eight places a wrong number costs money
One core, eight rule-sets — finance, legal, infrastructure, agents — for the domains where being confidently wrong is expensive.
Take the packages you need. Leave the rest. Python, TypeScript, Go, and Rust SDKs.
- qwedCore
Core 12-engine verification protocol
- qwed-mcpAgents
MCP tool-call and server verification
- qwed-open-responsesAPI
Verified OpenAI-compatible responses
- qwed-financeFinance
Banking, interest rates, ISO 20022
- qwed-legalLegal
Contract deadlines and citations
- qwed-taxTax
Tax compliance and withholding
- qwed-infraInfra
IaC verification — Terraform, IAM
- qwed-ucpCommerce
E-commerce transaction verification
A gate you can’t inspect is just another thing to trust
The layer standing between model output and production execution has to be readable by the people whose names are on the release.
We ask you to stop taking a model’s word. It would be strange to then ask you to take ours. The layer is Apache 2.0: inspectable, forkable, deployable at home.
“You cannot ask teams to trust a black-box verification layer.”
01
Read every check
The logic, the proof paths, the reasons for refusal — all of it is on the page.
02
Test it yourself
Run it twice and diff the output. That is a different kind of claim than a promise.
03
Deploy it anywhere
Self-hosted, private, air-gapped. Open source is what makes staying home possible.
04
Belong to no one
A portable layer under every model — a standard, not a vendor feature.
Watched by people who are not usParties of record
Mission
Make what a model says safe enough to act on, by proving it first.
Vision
Be the deterministic verification layer standing between AI reasoning and the real world.
Read the checks. Run them yourself. Then decide what to trust.
Questions put to the record
Eight things people ask before they trust a gate with their production traffic.
RAG feeds the model better documents. QWED reads the answer that comes back. One improves what goes in; the other decides what is allowed out. RAG adds knowledge, QWED adds certainty, and you probably want both.
Anything whose correctness can actually be established: math, logic, code, SQL, schemas, process flows, and agent tool calls. Every check ends in one of three verdicts, and the claim is approved, corrected, or blocked before it can execute.
Yes. QWED is model-agnostic and works with GPT-4, Claude, Gemini, Llama, Mistral, and whatever ships next month. It verifies outputs rather than trusting a provider, so changing models does not change what you can prove.
No. Fine-tuning shapes how a model behaves, its style, and its fit to a task. It still cannot tell you whether this particular answer is correct or safe to execute. QWED is the check after the behaviour, not a replacement for it.
Yes, Apache 2.0. A verification layer you cannot read is one more thing you are being asked to believe, so the core is inspectable, auditable, and self-hostable by design.
It depends on the engine. Symbolic checks over math, SQL, or logic usually finish in milliseconds; consensus and multi-model paths take longer. The more useful question is not how long a check takes, but what an unverified output costs on the day it is wrong.
It means the same input produces the same verdict, every time, with no sampling in between. The math, logic, SQL, code, and schema engines use symbolic solvers and program analysis rather than asking another model for its opinion.
Yes. QWED is built to run inside your own infrastructure, with your own model providers and your own verification policy. Nothing about checking an answer requires your data to leave the building.