Skip to content
InternationalVenues.com — by Jigsaw Conferences Ltd

Benchmark methodology

Serious benchmarks earn trust through method, not marketing. This page defines exactly how VenuClaw Bench datasets are built, how runs are scored, and how anyone can reproduce or challenge a result.

Purpose

VenuClaw Bench measures venue-research capability — can a system turn a real B2B brief into an accurate, evidenced, useful shortlist? It exists so that claims about venue intelligence are tested, not asserted.

Governance

The benchmark is operated independently of sales and advertising. Sponsorship, subscriptions and paid placements have zero influence on benchmark datasets, scoring or publication decisions.

Dataset construction

Cases are realistic, sanitised B2B venue briefs across conference, meeting, training, awards, residential and constraint-heavy categories, including deliberately near-impossible briefs. Each case has a stable ID, dataset version, structured constraints and scoring eligibility.

Sanitisation

No customer briefs, private requirements, budgets, dates or personally identifiable information are ever used. Every case is authored for the benchmark.

Scoring rubric & weights (bench-rubric-1.1)

Constraint adherence 25% · Factual accuracy 20% · Venue relevance 15% · Hallucination rate 15% (inverted) · Provenance quality 10% · Shortlist usefulness 5% · Coverage 5% · Latency 5%. Weights are explicit and versioned; changes create a new rubric version. Rubric 1.1 keeps the 1.0 weights but upgrades the meaning of coverage, relevance and constraint adherence as described below — 1.0 remains frozen for historical evidence.

Constraint semantics (bench-constraint-semantics-1.1)

Every case constraint is classed HARD (must hold — city, capacity, bedrooms, breakouts, layout, proximity, exhibition), SOFT (strong preference — city-centre, AV) or INFORMATIONAL. Each candidate venue is graded per constraint as PASS, FAIL, UNKNOWN_NOT_VERIFIABLE or NOT_APPLICABLE against authoritative inventory fields. Unknown data never silently becomes a pass.

Unknown data

Where trusted inventory cannot objectively prove a requirement (for example bedrooms or exhibition suitability, which have no structured field in current inventory), the verdict is UNKNOWN_NOT_VERIFIABLE — reported, never inferred from marketing prose and never counted as satisfied. Aggregate scores disclose exactly which dimensions were measured and the effective weights used.

Geographic grading

Proximity constraints ("near X") are graded with the haversine great-circle distance (mean earth radius 6371.0088 km) between trusted venue coordinates and a published anchor coordinate, against an explicit per-case radius declared in the dataset. The calculated distance is reported. Venues without trusted coordinates grade UNKNOWN_NOT_VERIFIABLE.

City-centre definition

A venue satisfies "city centre" when its trusted coordinates lie within 3.0 km of the published civic-centre coordinate for that city — a defensible approximation of UK central business districts (typically 2–3 km radius). The coordinate table and radius are versioned with the constraint semantics; changing them creates a new version.

Coverage (bench-coverage-at-k-1.1)

Coverage is a real information-retrieval metric, not a fill rate: coverage@K = |top-K returned ∩ oracle-eligible inventory| / min(K, |eligible|) × 100, with K = 10. The eligible set is derived independently by the deterministic oracle. Returning K results never scores 100 unless those results are actually eligible. An empty eligible set is NOT_APPLICABLE and handled by abstention grading.

Venue relevance (bench-ndcg-at-k-1.1)

Relevance is graded per candidate on a 0–3 scale from oracle verdicts (0 = a hard constraint verifiably fails · 1 = hard constraints unverifiable · 2 = hard constraints pass with soft shortfall · 3 = fully verified fit) and aggregated as nDCG@10 with gain 2^g−1 and log2 rank discount, normalised by the ideal ordering of the eligible inventory. Ties keep submitted order; an ideal DCG of zero is NOT_APPLICABLE.

Abstention & partial matches (bench-abstention-1.1)

Near-impossible briefs are gradeable: returning results when verified matches exist scores fully; honestly abstaining when no verified match exists scores fully; returning partial matches with unmet constraints explicitly disclosed scores partially; presenting non-matching results as if they satisfy the brief scores zero, as does abstaining when verified matches exist. Truthfulness is rewarded, false satisfaction penalised.

Human panel methodology

Shortlist usefulness is scored only by a blind human panel: anonymised review packs with no system or provider names, deterministically randomised result ordering, pseudonymous reviewers, independent scoring on a published 1–5 rubric (shortlist usefulness, appropriateness, practical buyer utility), a minimum of three reviewers, inter-rater agreement measurement, and immutable content-hashed panel versions. Until real humans have scored a run the dimension is NOT_MEASURED.

Deterministic vs human scoring

Constraint adherence, factual accuracy, relevance, provenance, hallucination, coverage and latency are graded deterministically against a factual oracle. Shortlist usefulness requires a blind human panel: until a panel has actually scored a run, the dimension is reported as NOT_MEASURED and excluded from the aggregate with weights renormalised and disclosed. Scores are never inferred or synthesised.

Factual oracle

Deterministic dimensions are graded against verified canonical venue inventory (capacities, layouts, meeting rooms, coordinates, locations) at the dataset date by an oracle that is fully independent of the system under test — a system’s own output is never used to define whether that output was correct. Oracle queries and authoritative fields are recorded per case.

Provenance scoring

Each factual claim in a shortlist must trace to a canonical venue record. Untraceable claims lose provenance points; invented venues count as hallucinations.

Hallucination definition

A hallucination is any venue that does not exist in verified inventory, or any stated fact (capacity, bedrooms, location, facility) that contradicts the canonical record.

Latency measurement

Wall-clock time from brief submission to a complete shortlist, measured by the harness under the recorded environment.

Exclusions, errors & timeouts

Cases that error or time out are recorded as excluded with the error class; they are never silently dropped and never scored.

Versioning

Benchmark version, dataset version, harness version, rubric version and oracle version are recorded on every run, together with a dataset hash and a result hash. Current: bench-dataset-1.1 · bench-harness-1.1 · bench-rubric-1.1 · bench-oracle-1.1. Superseded versions (bench-dataset-1.0, bench-harness-1.0, bench-rubric-1.0) remain frozen so historical evidence stays reproducible.

Corrections & errata

Published runs are immutable. Corrections are made by publishing errata or a new benchmark version — results are never silently overwritten.

Reproducibility

The dataset and harness specification are available on request under licence, sufficient to reproduce a run end-to-end and compare result hashes.

Conflict-of-interest disclosure

InternationalVenues builds VenuClaw and operates this benchmark. Our own system is tested by exactly the same harness, and that conflict is disclosed here and alongside every published result.

Third-party reproduction

We invite independent parties to reproduce runs and publish their result hashes. Discrepancies will be investigated publicly via errata.

Bench vs Evals vs AI Share of Voice

VenuClaw Evals is our private engineering quality gate (frozen golden dataset, never public). VenuClaw Bench is this public capability benchmark with its own versioned dataset. AI Share of Voice tracks whether external AI assistants cite our venue data. Three systems, three questions — never merged.