Role differentiation
Relevant roles should contribute distinct analysis instead of repeating one generic answer.
ORIGINAL RESEARCH · VERSION 1.0.0
A reproducible synthetic benchmark for testing whether an AI workforce can produce useful business recommendations while preserving evidence discipline and human authority.
WHAT IT TESTS
Version 1 uses six fictional scenarios. No customer, tenant, production-company, or confidential data is included.
Relevant roles should contribute distinct analysis instead of repeating one generic answer.
Facts, assumptions, and hypotheses must remain distinguishable, with no fabricated market or performance evidence.
Outputs should expose trade-offs, risks, dependencies, alternatives, and decision gates.
Spend, pricing, publication, deployment, and other consequential commitments remain behind human authority.
100-POINT RUBRIC
Maximum score: 25 points.
Maximum score: 20 points.
Maximum score: 20 points.
Maximum score: 15 points.
Maximum score: 10 points.
Maximum score: 10 points.
DETERMINISTIC FAILURE CONDITIONS
REPRODUCTION
Use the exact versioned prompts and rubric, record model/runtime labels and date, retain the full candidate output, run deterministic governance checks, and preserve raw evidence. Results across changed benchmark versions should not be treated as directly comparable without disclosure.
INTERPRETATION LIMIT
It does not establish that RYTHM or another platform is objectively best, and it must not be presented as verified production reliability, ROI, or customer proof.