The rubric decides more than the judges do.
This is the complete pitch competition judging rubric that ships as the default in EvalLens: six weighted dimensions, an anchor description for every band, and a disagreement threshold with a defined action. Take it and run your event with it, no signup and no tool required. The research below explains why the rubric, not the composition of your jury, is the variable that moves score agreement.
Breaking and gymnastics have the same caliber of judge. Only one has reliable scores.
A study of Olympic breaking at the Paris 2024 games found judge agreement between 0.21 and 0.45 on loosely defined criteria, while artistic gymnastics, which enumerates every observable element, reaches 0.94 to 0.98. The judges are world class in both sports. The rubric is the difference. A pitch jury scoring “Team” and “Market” on a bare 1 to 10 column sits on the breaking side of that gap.
Six dimensions, weighted, with anchors for 3 versus 7.
These are the default dimensions and weights of the EvalLens pitch panel. The anchor columns describe the two bands juries argue about most: what a middling 3 looks like against a strong 7. Weights are yours to edit before scoring starts. Once scoring begins they lock, so every submission is ranked on the same standard.
| Dimension | Weight | A 3 looks like | A 7 looks like |
|---|---|---|---|
| P1Problem significanceRed flag: “Huge market therefore big problem” hand-waving; a solution in search of a problem. | 0.15 | No real problem articulated, or the pain is vague and unsubstantiated. | A specific, frequent, costly pain with a clearly identified user and a credible reason it matters now. |
| P2Solution differentiationRed flag: Feature soup with no clear core; differentiation that is purely branding. | 0.15 | Solution is unclear, disconnected from the problem, or a thin wrapper over an existing tool with no real advantage. | A coherent solution with a clear mechanism and a genuine, defensible difference from alternatives. |
| P3Market attractivenessRed flag: Top-down “huge TAM therefore success” with no entry segment. | 0.20 | No real market reasoning, no segment defined, or an obviously implausible market claim. | A well sized, reachable market with a clear segment, a credible entry motion, and a believable path to first customers. |
| P4Business model / GTMRed flag: “We will figure out pricing later” as the entire monetization story. | 0.15 | No monetization logic, or pricing that ignores the buyer. | Clear monetization, sensible pricing for the buyer, and a credible go-to-market motion with a beachhead. |
| P5Team / founder fitRed flag: Credentials substituting for evidence of execution; critical roles missing with no plan. | 0.20 | No meaningful information about the team, or an obvious mismatch between the team and what the venture requires. | A capable, reasonably complete team with clear relevant experience and good fit to the problem. |
| P6Feasibility / readinessRed flag: A roadmap that assumes no setbacks; ignored dependencies that gate the plan. | 0.15 | The plan is implausible, internally inconsistent, or absent. | A credible, well sequenced plan with resources that broadly match the ambition and risks acknowledged. |
The full version has four bands per dimension. In the product, each dimension carries anchors for 0 to 3, 4 to 6, 7 to 8 and 9 to 10, plus the red flags above, and the top band is reserved for what is demonstrated, not merely asserted. The two columns here are the working core: if your judges can tell a 3 from a 7 the same way, most of the disagreement problem is already gone.
Five rules that turn a scale into a
The rubric table is half the tool. The other half is the procedure judges follow to land on a number. These five rules are the scoring procedure EvalLens hard-codes for its own AI panel, and they work just as well for a human one.
- 01
Evidence before numbers
Every claim a judge relies on gets a slide reference. No slide, no claim. This one habit converts “I liked it” into a record another judge can check, and it is the difference between a score and an opinion.
- 02
Collect both directions, then the gaps
Before scoring, write down what supports a higher band, what pulls toward a lower band, and what the deck simply does not establish. Missing evidence is recorded as missing, never invented and never silently forgiven.
- 03
Name the band before the number
The judge states which band the evidence lands in, in the form “this falls in the 7 to 8 band because…”, and only then picks a score inside that band. The decision is the band. The number just lives in it.
- 04
On the boundary, take the lower band
When the score sits between two bands and material evidence is missing, the rule is the lower band. Conservative by design: incomplete decks stay comparable instead of drifting up on benefit of the doubt.
- 05
Check the spread, do not average it away
Per dimension, Spread is the highest score minus the lowest across the reads that cover it. Below 1.5 is consensus, 1.5 to 2.99 is a split, and 3.0 or more is a conflict that goes to manual review. Averaging a conflict hides exactly the case a jury exists to discuss. The full aggregation math is published on our methodology page.
One core, three formats.
Printable scorecard
One page per submission: six rows for P1 to P6, the weight next to each, one score field and one evidence line per row, and a spread check at the bottom of the stack. If a judge cannot cite where a score came from, the blank line makes that visible in the room, not after the results email.
Demo day
Keep the same six dimensions. A cohort that just finished a program has had equal coaching on story, so many organizers shift weight toward Team / founder fit (P5) and Feasibility (P6) and away from pitch polish. Decide the weights before scoring starts and lock them, so the field is ranked on one standard end to end.
Hackathon (H1 to H6)
Hackathons are execution-first, so the rubric changes shape: Execution and Demo carries 0.30 and Technical Depth 0.20, both protected, with Problem Impact and Innovation Divergence at 0.15 and UX Clarity and Delivery Readiness at 0.10. The full version lives on our hackathons page.
Running a pitch competition end to end? The pitch competitions page shows how this rubric drives a full first read of the field, and the methodology page publishes the aggregation and Spread math behind it.
Before you copy it into your
- Can I use this rubric without EvalLens?
- Yes, and that is the point of publishing it. The dimensions, weights, anchors and the Spread rule work on paper, in a spreadsheet, or in any scoring tool. EvalLens becomes useful when the field is bigger than your judges' hours: an AI panel runs the first read on this same rubric and your judges decide with the evidence in front of them.
- Why these weights, and can I change them?
- The weights are the defaults of the EvalLens pitch panel: Market attractiveness and Team carry 0.20 each, the other four dimensions 0.15 each, reflecting what early stage juries most often argue about. They are meant to be edited before scoring starts. The one rule worth keeping is the lock: once scoring begins, weights freeze, so every submission is ranked on the same standard.
- What does the Spread threshold actually do?
- Spread per dimension is the highest score minus the lowest across the judges who cover it. Below 1.5 reads as consensus, 1.5 to 2.99 as a split, and 3.0 or more as a conflict. A conflict does not change anyone's score. It routes that dimension to a human conversation instead of letting an average bury the disagreement.
- How is this different from the university scorecard PDFs?
- Most published scorecards list criteria names and a 1 to 10 column, which is exactly the setup that produced 0.21 to 0.45 agreement among Olympic breaking judges. This rubric adds the three things that move that number: an anchor description for each band, an evidence line per score, and a disagreement threshold with a defined action.
- Does an AI panel score with this rubric too?
- In EvalLens, yes. The same anchors and red flags drive a panel of AI reviewers that reads every submission before your judges do, and the same Spread rule flags where those reads conflict. It has carried 1,000+ evaluation runs. The output is advisory: AI prepares the analysis, the judges decide.
Want the rubric to run itself?
Use the rubric freely. When the field outgrows your judges' hours, EvalLens runs the first read on it: an AI panel scores every submission against these anchors, flags every Spread conflict, and hands your jury the evidence. AI prepares the analysis, the judges decide.