Skip to content
InternationalVenues.com — by Jigsaw Conferences Ltd

International Venue Research Institute™

VenuClaw Bench

An open benchmark for testing how accurately AI venue-search systems understand real event briefs.

Measure constraint compliance, evidence quality, geographic accuracy, honest uncertainty and buyer usefulness against reproducible venue-search tasks.

Public results pending validated benchmark run

Why this benchmark exists

Typical AI venue recommendations can appear plausible while failing on capacity, room layout, breakout rooms, location radius, bedrooms, evidence, current inventory, uncertainty and citation integrity. VenuClaw Bench measures whether a system understands the brief, finds valid options, shows why they match, and admits when it cannot verify something.

What the benchmark measures

Twelve dimensions. Final public weightings are set at the validated run — no weighting is claimed before it exists.

01

Brief Understanding

Does the system parse the full brief — headcount, layout, dates, location, budget signals — before searching?

Missed constraints produce confident but useless shortlists.

Deterministic + judgement

02

Hard-Constraint Compliance

Does every returned venue satisfy the stated non-negotiables?

One non-compliant venue in a shortlist wastes a buyer’s time and trust.

Deterministic

03

Capacity Accuracy

Do stated capacities match structured, layout-specific venue records?

Capacity is the most common silent failure in AI venue answers.

Deterministic

04

Layout Accuracy

Is theatre / cabaret / boardroom capacity evidence layout-correct, not a single vague number?

A room that seats 300 theatre-style may seat 120 cabaret.

Deterministic

05

Breakout-Room Evidence

Are breakout requirements matched with actual room-level evidence?

Plausible-sounding “has breakout space” claims often have no data behind them.

Deterministic

06

Geographic / Distance Accuracy

Are radius and proximity claims measured from verified coordinates?

Distance guesses fail airport and multi-site briefs.

Deterministic

07

Accommodation / Bedroom Honesty

Does the system admit when bedroom data cannot be verified rather than inventing it?

Unverifiable bedroom claims are a leading hallucination class.

Deterministic

08

Venue Validity / Hallucination Control

Is every venue a real, currently listed entity — no invented or defunct venues?

Hallucinated venues are the hardest failure for buyers to detect.

Deterministic

09

Evidence & Citation Integrity

Is every material claim traceable to a canonical venue record?

Uncited claims cannot be audited or trusted in procurement.

Deterministic

10

Honest Uncertainty / Abstention

Does the system say “cannot verify” or ask for clarification when the data does not support an answer?

Refusing to guess is a capability, not a weakness.

Deterministic + judgement

11

Tool / Search Selection

Does the system choose the right retrieval strategy for the brief instead of a generic lookup?

Wrong tooling produces shallow, non-compliant candidate pools.

Deterministic + judgement

12

Buyer Usefulness

Would a real organiser act on this shortlist?

The final test: decision support, clarity and practical value.

Blind human panel (pending)

Public benchmark cases

Representative public-safe cases. The full versioned task set, sanitisation and anti-overfitting policy are described in the methodology. Private regression fixtures are never published.

Brief

A one-day conference in London for 200 delegates, theatre-style, with organiser-verified capacity evidence.

Hard constraints

  • City: London
  • Capacity: ≥200 theatre-style
  • Layout-specific evidence required

Expected evidence: Structured theatre capacity ≥ 200 from the venue’s canonical record for every returned venue.

Fail conditions: Any venue below 200 theatre; capacity quoted without layout; unverifiable capacity claims.

Brief

A training programme needing a main room for 120 plus at least 4 separate breakout rooms.

Hard constraints

  • Main room ≥120
  • ≥4 distinct breakout rooms
  • Room-level evidence

Expected evidence: Named or counted breakout rooms in structured records — not a generic “flexible space” claim.

Fail conditions: Breakout claims with no room-level data; double-counting the main room.

Brief

A meeting venue within 10 km of Heathrow Airport for 60 delegates.

Hard constraints

  • ≤10 km measured from trusted airport coordinates
  • Capacity ≥60

Expected evidence: Haversine distance from verified venue coordinates to the trusted anchor; distance shown per venue.

Fail conditions: Any venue outside the measured radius; city-name matching used as a proxy for distance.

Brief

A residential conference for 500 with 300 bedrooms on site — where bedroom data is not verifiable.

Hard constraints

  • Capacity ≥500
  • 300 on-site bedrooms
  • Bedroom evidence required

Expected evidence: An explicit statement that bedroom counts cannot currently be verified, with the compliant non-bedroom facts still delivered.

Fail conditions: Invented bedroom counts; silently dropping the bedroom constraint.

Brief

An intentionally impossible brief: 5,000 theatre-style in a small market town within 5 km of its rail station.

Hard constraints

  • Capacity ≥5,000
  • Specific small-town radius

Expected evidence: A clear, early statement that no compliant venue exists, with the nearest verifiable alternatives labelled as non-compliant.

Fail conditions: Returning plausible-looking but non-compliant venues as if they matched.

Brief

A combined brief: Manchester, 200 cabaret-style, ground-floor access, within a defined budget band.

Hard constraints

  • City + layout capacity + accessibility + budget

Expected evidence: Every dimension answered with record-level evidence, or explicitly flagged as unverifiable.

Fail conditions: Answering the easy constraints and silently ignoring the hard ones.

Brief

An ambiguous brief (“somewhere impressive for an important client dinner”) requiring clarification.

Hard constraints

  • Under-specified: no city, headcount or date

Expected evidence: Targeted clarifying questions or an explicit human-handoff — not a confident generic list.

Fail conditions: A confident shortlist produced from guessed constraints.

Deterministic checks

Graded automatically against structured venue records: capacity compliance, radius compliance, valid venue identity, citation presence, hard-constraint failure, hallucination and structured evidence integrity. Same input, same grade — every run.

Human-judgement checks

Usefulness, clarity, decision support, appropriateness and practical buyer value are scored only by a blind human panel. Human panel validation is pending — until real humans have scored a run, these dimensions are reported as NOT_MEASURED and no comparative results are published. AI-assisted shadow scoring is never presented as human scoring.

Public results

Pending validated benchmark run

The methodology and benchmark structure are public. Comparative scores will be published only after the validation protocol and human-review stage are complete. No leaderboard, ranking or score appears here until then.

Machine-readable status: publication_status: results_not_published at /api/bench/spec

Reproducibility & versioning

Versioned tasks

Every case set is content-hashed and versioned; results always cite the exact benchmark version.

Deterministic graders

Objective dimensions grade against a factual oracle built from canonical venue records.

Publication policy

Results appear only after a complete reproducible run — never estimated, never placeholder.

Change policy

Task or rubric changes create a new version; published runs are immutable with errata noted.