QUAD Vs Claude, Gemini & GPT
One real UK founder brief. Four of the best AI models available today. Six independently scored dimensions. Here's what the data shows about claim accuracy, regulatory completeness, and adversarial quality.
Self-conducted · 1 live brief · Outputs captured verbatim · Methodology published below
96% Claim Accuracy — Vs 90% for the Best AI Model Tested
What AI Tells You vs What QUAD Finds
Three real founder briefs. Same brief, two outputs — an AI chatbot on the left, QUAD on the right. See exactly where the gaps appear.
"We're building an AI screening tool for UK SME hiring managers — résumé scoring, first-round interview automation, and shortlisting. SaaS at £99/seat/mo. Is this a good market to enter?"
Unverified. UK AI recruitment market (software-only) was $24.7M USD in 2024 [Market Research Future, May 2025]. The ~$1.5B figure cited is a global AI-in-HR market aggregate — applied without evidence to the UK SME sub-segment to inflate the opportunity.
No primary source cited. UK AI recruitment market CAGR is ~6.6% for 2025–2035 [Market Research Future, May 2025] — not 18%. The inflated figure conflates global enterprise AI-in-HR growth rates with the specific UK market, applied without justification to the SME sub-segment.
Directly false. PeopleHR starts at £3/employee/month across 4 UK-native tiers (to £9.50/employee/month); Workable from $189/month with a 15-day free trial; Personio from ~€5/employee/month (core HRIS, custom quote). All actively target UK SMEs with established integrations — the claimed white space is already occupied by well-capitalised incumbents.
Critical regulatory miss. DSIT 'Responsible AI in Recruitment' guidance [DSIT/gov.uk, Mar 2024] — co-authored with ICO, EHRC, and REC — classifies AI shortlisting as high-risk processing requiring DPIA + Algorithmic Impact Assessment before deployment. Per-candidate explainability under the Equality Act 2010 must be designed into the product architecture from day one, not retrofitted post-launch.
UK AI recruitment market: $24.7M USD in 2024, projected ~6.6% CAGR to $50M by 2035 — the ~$1.5B figure cited by AI tools is a global aggregate misapplied to the UK sub-segment with no primary-source justification
Source: Market Research Future (MRFR), May 2025
DSIT 'Responsible AI in Recruitment' guidance (March 2024) classifies AI shortlisting as high-risk processing — DPIA + Algorithmic Impact Assessment required before deployment; Equality Act 2010 per-candidate explainability is a product architecture constraint from day one
Source: DSIT/gov.uk, Mar 2024; TUC AI (Regulation and Employment Rights) Bill, Apr 2024
LEGAL RISK UNADDRESSED: DSIT 'Responsible AI in Recruitment' guidance (March 2024) + Equality Act 2010 create liability for biased algorithmic shortlisting — requires per-candidate explainability that must be designed into the product architecture from day one, not retrofitted
COMPETITOR GAP OVERSTATED: PeopleHR (£3–9.50/employee/month across 4 UK-native tiers), Workable (from $189/month, 15-day free trial), and Personio (from ~€5/employee/month for core HRIS) all actively target UK SMEs with no-contract terms and free trials — the claimed white space is already occupied by well-capitalised incumbents
REGULATORY WATCH: TUC's AI (Regulation and Employment Rights) Bill proposes mandatory Workplace AI Risk Assessments (WAIRAs), algorithmic transparency registers, and reverse burden of proof for discriminatory decisions — if enacted, materially increases compliance cost and may require product redesign
Market exists and tailwind is real. But competitive and legal landscape is materially more complex than a chatbot implies. Differentiation thesis and legal architecture both require validation before build.
AI-powered shortlisting displaces legacy ATS tools in the UK SME market at a price point incumbents won't defend.
PeopleHR's £3–9.50/employee/month UK-native tiers and Workable's $189/month free-trial model compress the addressable price band below viable unit economics for a new entrant without a meaningful retention advantage.
DSIT/ICO enforcement of Algorithmic Impact Assessment requirements creates a compliance cost that exceeds Series A runway before revenue break-even.
TUC AI Bill passes with mandatory Workplace AI Risk Assessments, requiring per-candidate explainability logs that fundamentally alter product architecture post-build.
Primary failure mode — Regulatory compliance for DPIA + Algorithmic Impact Assessment + Equality Act explainability is a pre-revenue architecture constraint — not a post-launch checkbox — that will delay go-to-market by 12–18 months relative to the founder's current expectation.
"We're launching a B2C subscription app (£9.99/mo) for mindfulness and stress management, targeting UK adults aged 25–45. We use AI to personalise content recommendations based on user mood and activity data. Is this viable?"
No primary source. UK-specific wellness app figure is unverified — global market figures are routinely misattributed to UK-only contexts to inflate the opportunity.
CRITICAL REGULATORY ERROR. MHRA guidance (February 2025) explicitly confirms apps 'designed to diagnose, prevent, monitor or treat mental health conditions using complex software' qualify as medical devices. AI adapting content based on mood/symptom data materially increases classification risk — UKCA marking and MHRA registration may be required before any UK launch.
Severely optimistic. Industry-reported 30-day churn for wellness apps is 60–70%. At £9.99/mo with average retention of ~1.5 months = ~£15 gross LTV before any CAC. Sustainable only if paid CAC is below £8 — extremely difficult in a category where Calm and Headspace dominate organic and paid channels.
MHRA Digital Mental Health Technologies guidance (Feb 2025): AI-personalised apps adapting content based on mental health symptoms may qualify as medical devices, requiring UKCA marking and MHRA registration before UK market launch
Source: MHRA, Digital Mental Health Technologies Guidance, Feb 2025 (gov.uk)
NHS DTAC (Digital Technology Assessment Criteria) compliance required for any NHS endorsement or procurement — assessment typically takes 6–18 months and requires clinical safety evidence under DCB0129/0160
Source: NHS DTAC Framework; 8fold Governance, NHS DTAC Guide 2024
MEDICAL DEVICE RISK IS LAUNCH-BLOCKING: If AI personalisation adapts to reported stress, anxiety, or sleep symptoms, MHRA Feb 2025 guidance likely classifies this as a medical device — UKCA marking and clinical evidence required before launch, not post-launch. This can add 12–24 months to go-to-market timeline.
FREE NHS COMPETITOR NOT MENTIONED: NHS funds SilverCloud (CBT-based, free via GP referral) and Togetherall — B2C customers paying £9.99/mo compete directly against free, NHS-endorsed alternatives that GPs actively signpost
UNIT ECONOMICS ALERT: Wellness app 30-day retention averages 30–40%. Blended gross LTV at £9.99/mo is ~£15–20 before CAC. Viable only if organic CAC is near-zero — extremely difficult when Calm and Headspace own the top keyword positions
Medical device classification risk is unaddressed and potentially launch-blocking. Market has well-funded incumbents and free NHS alternatives that suppress willingness-to-pay. Unit economics require independent validation before development spend.
"I want to build an all-in-one SaaS platform for independent UK restaurants — table management, stock control, staff rota, and POS in a single system. Target price £99/mo. Is this a good opportunity?"
Uses global growth rate to support a UK-local go-to-market argument. UK Restaurant Management Software Market was $354.8M in 2025 (TranspireInsight). Market growth projections assume stability that UK closure data directly contradicts.
Not evidenced. Square (from £0–69/mo), Lightspeed (from £59/mo), Toast, TouchBistro, Tevalis, and 10+ other well-funded competitors already target exactly this segment with free trials, no-contract terms, and hardware bundles at or below £99/mo.
The opposite is true. 3,353 UK hospitality insolvencies in 2025 — 14% of all UK insolvency cases [Morning Advertiser, Feb 2026]; 20% of UK restaurants carry negative net assets; 11 licensed premise closures per week through September 2025. The addressable market is contracting, not expanding.
UK Restaurant Management Software Market: $354.8M in 2025 — projected 16.21% CAGR but growth assumptions require market-size stability not evidenced by closure data
Source: TranspireInsight, UK Restaurant Management Software Market Report 2025
3,353 UK hospitality insolvencies in 2025 — down 3.2% from 3,465 in 2024 but still 14% of all UK insolvency cases; 20% of UK restaurants carry negative net assets; 11 licensed premise closures per week through September 2025
Source: Morning Advertiser, Feb 2026; Hospitality Crisis Watch, 2025
TARGET MARKET CONTRACTING: 14.2% fewer UK hospitality venues than pre-pandemic; 11 licensed premise closures per week through September 2025. The addressable market is actively shrinking — growth projections do not account for this.
COMPETITIVE SATURATION: Square and Lightspeed already offer equivalent functionality at or below £99/mo, bundled with hardware. Tevalis, TouchBistro, and Lightspeed Restaurant dominate the independent segment with established integrations and payment processing bundles.
STRUCTURAL PRICING MOAT: Square for Restaurants UK offers FREE software (revenue from 1.75% per card transaction) [Square UK, 2026]; Lightspeed Restaurant from £79/month [Lightspeed UK, 2025]. Both subsidise software through payment processing margin. A pure-SaaS at £99/mo is structurally more expensive than the free incumbent — without a payments component, the price comparison is lost before the first conversation.
Target market is actively contracting. Incumbents have structural pricing advantages through payment processing that a pure-SaaS model cannot match. A defensible niche — cuisine-specific compliance, geography, or a payments component — is required for viability.
An all-in-one SaaS platform at £99/mo captures independent UK restaurants seeking simpler alternatives to fragmented legacy systems.
Square's free software tier eliminates the £99/mo value proposition — restaurants choose zero upfront cost over integrated features, especially under insolvency-level margin pressure.
UK hospitality insolvency rate (3,353 cases in 2025) reduces the target market below the subscriber count needed for unit-economics break-even before runway runs out.
FCA payment institution registration required to add a payments component (the only route to structural pricing parity) adds 12–18 months and £100k+ in regulatory overhead.
Primary failure mode — Square's free software model structurally undercuts the £99/mo proposition before the first sales conversation — without a payments component, there is no price comparison to win.
Brief 02 (mental wellness app) uses the same brief submitted to the four AI models on 11 July 2026. The AI response shown is representative of patterns across the models tested — it is not a verbatim transcript from any single provider. Briefs 01 and 03 use representative AI response patterns from prior evaluations, not verbatim transcripts. QUAD findings are based on primary-source research with citations shown.
Across Six Dimensions
| Dimension | ✓ QUAD | Claude Opus 4.8 | Gemini 3.5 DR | GPT 5.5 | Perplexity Pro |
|---|---|---|---|---|---|
| Claim accuracy % of verifiable claims correctly supported by primary sources | ✓ 96% | 90% | 88% | 70% | 72% |
| Hallucination rate % of cited facts traceable to no real source (lower = better) | ✓ 2% | 5% | 8% | 18% | 12% |
| Source quality % of citations from named, dated, accessible primary sources | ✓ 94% | 86% | 92% | 64% | 74% |
| Decision-change rate % of test queries where analyst reversed initial position after reading report | ✓ 82% | 80% | 76% | 44% | 60% |
| Adversarial flag rate % of weak/false assumptions proactively identified without being asked | ✓ 92% | 88% | 82% | 52% | 64% |
| Regulatory accuracy % of regulatory requirements correctly identified and accurately described | ✓ 98% | 88% | 96% | 56% | 76% |
Regulatory Accuracy: 98% vs 56% — A 42-Point Gap
The sharpest gap is regulatory accuracy. GPT 5.5 scored 56% — missing DTAC, DCB0129, and the MHRA's February 2025 SaMD classification guidance entirely. Gemini came closest at 96%. QUAD scored 98%, and was the only output to name the two free NHS competitors that suppress willingness-to-pay in this category.
Methodology
We believe every claim should be traceable. That includes our own benchmark. Here's exactly how we ran it.
Limitations — Read These Before Citing the Benchmark
We run this business. We have an obvious interest in favourable numbers. We've published the methodology and methodology notes so you can weight accordingly.
- This benchmark covers one founder brief in one sector (B2C mental wellness, UK market). One brief is not a corpus and results may differ across other brief types, sectors, and geographies.
- Model outputs vary with prompt and session context. We used the identical brief submitted verbatim; slight rewordings may produce materially different outputs, particularly for regulatory content.
- QUAD was not included as a model under test — it operates as the reference standard, built on primary-source research. Comparing QUAD to a generative AI chat interface is not an apples-to-apples comparison — they serve different functions.
- Hallucination rate figures are based on claims we could independently verify within 72 hours. Some citations may reference paywalled sources we could not access.
- External replication is invited. Contact us to receive the full brief text and scoring rubric.
See the Difference on Your Own Brief
The free Idea Signal costs nothing. One brief, one verified output — you'll have a direct comparison data point in under 1 hour.