In development · Pilot enquiries

The evidence gate for enterprise AI.

HalluciProof is designed to check each factual claim against its supporting evidence, apply a risk-based decision before delivery, and create a verifiable record of the result.

Explore the Platform
Individual claims, linked to evidence.
Decisions before delivery.
Portable assurance records.
Claim-level evidence review
Illustrative preview
Context: Draft generative answer for policy enquiry #POL-8492
Claim 1 · Coverage Status ALLOW
“Your comprehensive policy includes windshield repair and replacement subject to approved repairers.”
Evidence: Policy Section 4.2 matches wording and active customer schedule.
Claim 2 · Excess Amount BLOCK
“Your policy excess is £250 for accidental damage claims.”
Contradiction: Draft says “Your policy excess is £250”, but the customer schedule shows £350.
Claim 3 · Endorsement Freshness ESCALATE
“Courtesy car provision is guaranteed for up to 14 days during vehicle servicing.”
Freshness Alert: Claim relies on an endorsement that has passed its evidence freshness threshold.
Illustrative preview: This is sample content and does not perform verification or generate actual signed receipts.
Background & Governance

Addressing the Enterprise AI Verification Gap

Generative AI produces articulate answers, but lack of granular verification exposes organisations to compliance and remediation risk.

The Granularity Problem

Generative AI can produce fluent answers containing incorrect figures, unsupported conditions, unverifiable citations or statements based on outdated sources.

A single aggregate score for the whole answer may not show which individual statement is unsupported. When validation only happens post-delivery, organisations are left handling costly customer complaints, regulatory scrutiny and remediation.

The Pre-Delivery Control Layer

HalluciProof is developing a dedicated control layer that sits between generative AI systems and the end recipients of their answers.

Rather than inspecting output in retrospect, the platform isolates material factual claims, tests evidence sufficiency against verified records, and enforces organizational policy before answers are delivered.

HalluciProof complements existing AI evaluation and observability tools by introducing claim-level evidence lineage, independent adversarial challenge, and signed decision records.

Kainat Zanab

Founder and Chief Executive Officer · Lancaster, United Kingdom

Kainat Zanab holds an MSc in Digital Business Innovation and Management from Lancaster University, awarded with Distinction. The HalluciProof project draws directly on her postgraduate dissertation research into AI hallucinations and enterprise-safe AI deployment.

Relevant Background & Research:

  • An MSc consultancy engagement with SAP on AI governance, compliance automation and regulatory risk.
  • An IBM research collaboration delivered as part of her MSc.
  • An NHS Trust Lancashire AI in Healthcare university project.
  • A Cisco, Lancaster University Management School and NTNU university collaboration.
  • Enterprise Technology Consultant experience at BT Group, including participation in a pilot cohort trialling AI-assisted tools.
Note: The academic engagements above were conducted as university projects and research collaborations during MSc study. SAP, IBM, NHS, Cisco, and BT are not customers, investors, or partners of HalluciProof.

Kainat leads strategy, product definition, evidence methodology, customer discovery and commercial development, supported by a planned founding AI engineer.

System Architecture & Modules

Modular Architecture for Verifiable AI Delivery

A 10-module control system structured across four architectural layers, designed for granular evidence testing and auditability.

Layer 1

Input

Captures draft model responses, retrieved grounding context, metadata, and workflow identifiers from enterprise systems.

Layer 2

Core Verification

Decomposes text into discrete factual claims, builds evidence links, deterministically checks figures, and tests sufficiency.

Layer 3

Decision & Assurance

Applies customer risk thresholds, executes runtime gates, and generates tamper-evident cryptographic receipts.

Layer 4

Integration

Connects with vector stores, agent frameworks, SDKs, webhook sinks, and enterprise review consoles.

Module 01 Planned

Claim Decomposer

Separates material factual statements from tone, caveats and non-factual language. Categorises claims including eligibility, prices, dates, citations and general assertions.

Module 02 Planned

Evidence Graph Builder

Links claims to specific passages, records or tool outputs, retaining precise evidence references, source versions and cryptographic content hashes.

Module 03 Planned

Evidence Sufficiency Engine

Tests whether linked evidence materially supports the claim. Evaluates supported, partially supported, contradicted, missing or stale states with deterministic checks for numbers, dates and identifiers.

Module 04 Planned

Cross-Model Adversarial Verifier

Challenges high-risk claims that passed earlier checks using a model from an independent provider and architecture family, avoiding shared inductive bias with generating and sufficiency models.

Module 05 Planned

Risk Policy Engine

Applies customer-defined policies and thresholds by use case and claim type. Higher-risk statements (such as binding financial commitments or legal rights) face stricter evidence requirements.

Module 06 Planned

Runtime Decision Gate

Combines evidence results, challenge outcomes and customer policy to execute deterministic actions: allow, qualify, escalate or block the claim prior to end-user delivery.

Module 07 Planned

Evidence-Decay Monitor

Tracks source freshness over time and triggers re-verification of dependent claims when a grounding document, database record or guideline becomes stale or changes.

Module 08 Planned

Signed Evidence Receipt Service

Hashes, time-stamps and digitally signs claims, evidence references, challenge outcomes and gate decisions for portable and independent verification.

Module 09 Planned

Challenge and Test Generator

Builds targeted evidence-gap and adversarial test suites for each deployment and continuously feeds the customer workflow benchmark.

Module 10 Planned

API, SDK and Connectors

Provides planned interfaces for retrieval pipelines, agent frameworks and enterprise sources. Planned interfaces include REST and streaming APIs, Python and TypeScript SDKs, webhooks and connectors.

Supporting Capabilities

Enterprise controls and reporting utilities designed to accompany core gate verification:

Human review console
Customer-specific benchmark reports
Assurance reports
Policy versioning
Evidence source registry
Stale-source & decision-rate alerts
Receipt export & verification utility
Product Status: Concept & Architecture (TRL 2)

Planned Development Roadmap

Capabilities are currently planned or in development. Month 1 is modelled as January 2027 following endorsement and formal incorporation.

Month 3
Core Batch Pipeline

Delivery of core batch processing pipeline and execution of the first evidence-assurance assessment pilot.

Month 5
Runtime MVP

Deployment of runtime MVP and basic human review console for pilot partner testing.

Month 12
Adversarial Verifier GA

General availability of Cross-Model Adversarial Verifier add-on for high-risk claims.

Month 15
Signed Receipts GA

General availability of Signed Evidence Receipt Service with cryptographic hashing.

Month 21
Decay Monitor GA

General availability of continuous Evidence-Decay Monitor and stale source alerting.

Years 2–3
Sector Packs & Scale

Agentic trace coverage, regulated sector packs, and enterprise private deployment options.

Later Roadmap
Advanced Ecosystem

Opt-in federated benchmark exchange, rotating challenger model pool, and multimodal claim verification.

Roadmap Note: All milestone dates and timelines are planned projections based on business plan modelling. Product capabilities are described as planned and in development. Certifications (e.g. ISO) and patent filings are planned future objectives and are not claimed as completed achievements.
Operational Workflow

Eight Steps From Draft to Verifiable Decision

How an answer moves from raw generative output to isolated claims, sufficiency checks, and pre-delivery gate decisions.

1

Capture the draft

Receive the AI answer, retrieved context, and workflow identifier directly from the generation pipeline.

2

Separate the claims

Identify and classify the individual material factual statements, separating facts from tone and filler.

3

Connect the evidence

Link each claim to supporting passages, enterprise records or tool outputs with content hashes.

4

Test evidence sufficiency

Check support, partial support, contradiction, missing evidence, and freshness with deterministic validation.

5

Independently challenge high-risk claims

Use a separately sourced model family to challenge claims that passed initial checks and identify ungrounded assumptions.

6

Apply the customer’s risk policy

Evaluate against configured thresholds to return allow, qualify, escalate, or block.

7

Record and review

Create decision records and, once released, signed receipts. Route escalations to reviewers and log outcomes.

8

Monitor evidence freshness

Once the decay monitor is released, recheck dependent claims when their supporting evidence becomes outdated.

Selective Challenge

The independent challenge is selective and targeted at high-risk claims, not mandatory for every low-risk statement.

Direct Block on Missing Evidence

Missing or contradicted evidence can trigger block or escalation immediately without requiring an independent challenge.

Customer Policy & Oversight

The customer defines the risk tolerances, sets policy thresholds, and retains full human oversight at all times.

Pre-Delivery Gate Decisions

Every evaluated claim resolves to one of four deterministic gate outcomes:

ALLOW
Release Claim

Release the supported claim under the configured policy when evidence fully corroborates the statement.

QUALIFY
Conditional Release

Release the claim accompanied by necessary conditional phrasing, context, or required qualifying disclaimers.

ESCALATE
Human Review

Send the claim, linked evidence, and reason to a human reviewer prior to delivery for manual verification.

BLOCK
Halt Delivery

Stop the claim from being delivered to prevent unverified, contradictory, or high-risk claims from reaching users.

Customer Pilot & Adoption Journey

Adoption is designed to follow a phased, low-friction validation path:

1. Scope one workflow → 2. Agree success criteria → 3. Build a benchmark → 4. Assess claims → 5. Review results → 6. Integrate & monitor

Early pilots assess logged responses in batch to measure baseline verification rates. The initial pilot focuses on assessment and does not promise the immediate deployment of the complete future runtime platform.

Commercial Structure

Target Markets & Launch Pricing

Proposed Year 1 launch pricing from the business plan. Subscriptions are invoiced annually in advance. All amounts in GBP, excluding VAT.

Initial Focus

Financial Services

Customer assistants and adviser tools producing factual statements about products, pricing, cover, and eligibility terms.

Initial Focus

Legal & Professional Services

Legal research, document drafting, regulatory citations, and structured contract or clause summarisation.

Follow-on

Healthcare

Bounded knowledge assistants using provider-supplied guidance and clinical protocols. HalluciProof supports provider safety processes; it does not provide clinical advice.

Follow-on

AI Vendors

Embedding assurance gates directly into third-party AI software products via API/SDK and white-label licensing.

Market Rollout Strategy: Financial services and legal services represent the initial focus for Year 1. Healthcare workflows and vendor licensing follow as the platform matures. Insurance intermediaries and public-sector suppliers represent later expansion markets.
Fixed-Scope Engagement

Evidence-Assurance Pilot

£15,000 fixed fee, excluding VAT

A focused engagement to measure factual support and baseline gate performance on real logged responses.

  • Scope: One bounded workflow over 2 to 4 weeks.
  • Dataset: Normally 300–500 human-reviewed benchmark claims.
  • Criteria: Success criteria agreed formally before the pilot commences.
  • Deliverable: Comprehensive benchmark report covering agreed performance measures.
  • Methodology: Early pilots assess logged responses in batch.
Commercial Terms

Payment Schedule: 50% payable on signature and 50% payable on completion.

Initial Year 1 Design-Partner Offer: The full £15,000 pilot fee can be credited against the first year’s subscription if the customer signs within 60 days of pilot completion.
Starter
£1,750 / month equivalent
£21,000 per year, invoiced annually in advance
  • One gated workflow
  • Core verification engine
  • Review console
  • Standard receipts when released
  • Email support
Enterprise
£9,500 / month equivalent
£114,000 per year, invoiced annually in advance
  • Unlimited workflows
  • Private deployment options according to the roadmap
  • Custom policies & thresholds
  • Audit exports
  • Priority support

Optional Add-ons & Vendor Licensing

Specialised capabilities and partner programme pricing (all prices exclude VAT):

Cross-Model Adversarial Add-on

£950 per customer / month

Billed monthly in arrears. Premium verification for high-risk claim categories using independent model architectures. Planned general availability: Month 12.

Regulated-Sector Assurance Pack

£500 per customer / month

Billed monthly in arrears. Pre-configured policy packs for financial services, legal, or healthcare governance. Planned rollout in Years 2–3.

API/SDK & White-Label Licence

£3,500 / month equivalent

£42,000 per year, invoiced annually in advance. For AI vendors embedding verification into their products. Planned vendor programme from Year 2.

Pricing Assumptions Note: These figures represent proposed Year 1 launch assumptions from the business plan, subject to empirical validation through pilots. Prices remain unchanged in Year 2. Planned price adjustments start in Year 3 and apply to existing customers at renewal. Subscriptions are invoiced annually in advance; monthly figures are provided as equivalents for comparison only.
Frequently Asked Questions

Frequently Asked Questions

Key details on platform capabilities, architecture design, data governance, and pilot engagements.