A studio's hit rate is mostly determined by two things: how many ideas it can test, and how well it decides which ones to continue. The first is a production problem — we have written about our prototype pipeline. The second is a decision problem, and it deserves as much structure.
This is the framework we use to decide the fate of a prototype. It is deliberately simple. Complex scoring models feel rigorous but tend to hide the fact that someone already decided.
Principle 1: Thresholds before data
Every prototype brief states, before a line of code is written, the metrics that will be measured and the numbers below which the project stops. This is not because the numbers are perfect. It is because a threshold set after seeing the data is not a threshold; it is a rationalization.
Thresholds come from three places, in order of preference:
- The publisher's or studio's own portfolio benchmarks for the genre and market.
- Public genre benchmarks, adjusted for the test market.
- A conservative guess, flagged as such, to be revised after the first two prototypes.
Principle 2: Few metrics, clearly defined
We gate on a small set:
- D1 and D3 retention — the first-session and first-return signals available within a short window.
- D7 retention — when the test window allows; the strongest early predictor of a game.
- Average session length and sessions per day — whether the game fits into a routine.
- Early monetization signal — rewarded video engagement and, if a shop is live, first-purchase conversion. Not revenue; too noisy at prototype scale.
- CPI from creative tests — because a great retaining game nobody can acquire is not a business.
Each metric has a written definition ("D1 = share of installs with a session starting 24–48h after install, test-market cohort, organic excluded") that is used in every readout. Ambiguity in definitions is where decisions get quietly bent.
Principle 3: Enough data, not perfect data
Prototype tests run at small scale, so confidence intervals are wide. We handle this in two ways. First, we set a minimum cohort size for a readout to count at all, and extend the test rather than decide on a thin cohort. Second, we look at the direction and consistency across daily cohorts rather than a single aggregate — a metric that improves each day as levels are fixed is a different story from one that is flat.
We do not pretend to statistical rigor we do not have. We do insist on not deciding from three days of data because a deadline arrived.
Principle 4: Three outcomes, not two
Binary go / kill decisions push teams to argue for "go" because "kill" feels final. We use three outcomes:
- Go. All gate metrics at or above threshold. The title moves to a defined next phase — usually soft launch scope — with new, higher thresholds.
- Iterate. One or two metrics below threshold and the data points to a specific, fixable cause: a level cluster with abnormal drop-off, an onboarding step, a creative that misrepresents the game. One more sprint with a specific hypothesis, then a final go / kill. Iterate cannot be chosen twice in a row.
- Kill. Below threshold without a clear cause, or after a failed iterate. The mechanic, levels and learnings are archived in the framework's library.
The "iterate cannot repeat" rule is the most important sentence in this framework. It is what keeps iterate from becoming a slow-motion go.
Principle 5: The readout is a document, not a meeting
Before the decision meeting, the producer writes a one-page readout:
- The brief's hypothesis and thresholds, verbatim.
- Each gate metric with its value, cohort size and trend.
- Per-level funnel highlights: the three worst levels and why.
- Creative results: best CPI, retention by creative.
- A recommendation with the specific reason.
The meeting then discusses the recommendation rather than rediscovering the data. Meetings that start with a dashboard end with the loudest opinion.
What this looks like with a publisher
When we prototype with a publisher, the brief and thresholds are written together, the readout is a shared document and the decision meeting has both teams in it. That removes the most common friction in prototype deals: each side measuring different things and being surprised by the other's conclusion.
The honest part
This framework does not make the decision easy. It makes it legible. A killed prototype still hurts. But when the team can point to a number they agreed on months earlier, the hurt is about the game, not about the process — and the next brief gets written the following week.
Our analytics stack is built to produce exactly these readouts for every title on the framework. Read the case study.



