From the outside, judging a stack of pitch decks looks rigorous. A panel reads each deck, assigns scores, publishes a leaderboard. Clean.

The weak part is the part you can't see. Why did this deck score a 7 and that one a 6? What in the deck actually backed it? Where did the judges disagree? And the ranking that came out — did an AI decide it, or did you?

Most tools answer none of those. They hand you a number and hope you don't ask. EvalLens is built to answer all four, on every deck in the batch. That is what lets you screen a hundred decks in an afternoon and still defend every call.

The short version

Deck
→ Decoder
→ 6 Pitch AI judges (independent)
→ P1–P6 dimension scoring
→ Deterministic math (AI Total Score, advisory)
→ Human Review
→ Jury Score
→ Leaderboard

AI prepares the evaluation. You make the final scoring decision. Everything below is how that line is enforced, step by step — not as a slogan.

Step 1 — Decode the deck

Before any scoring starts, the system reads the deck as it is. The goal is to capture what the deck actually says — not to guess what the startup probably meant.

The Decoder turns any format into structured material for the judges:

  • slide content and visible text;
  • visual meaning and likely slide purpose;
  • missing or unclear information;
  • source references back to specific slides.

So a thin deck reads as a thin deck. A gap is flagged as a gap, not quietly filled in with a generous assumption. This matters more than it sounds: most "AI feedback" tools paper over missing information because a confident paragraph reads better than an honest blank.

Step 2 — Two things people confuse: dimensions and judges

This is where most explanations get fuzzy, so here is the distinction plainly.

A Pitch Competition deck is scored on six dimensions — the things being measured:

CodeDimensionDefault weight
P1Problem significance0.15
P2Solution differentiation0.15
P3Market attractiveness0.20
P4Business model / GTM0.15
P5Team / founder fit0.20
P6Feasibility / readiness0.15

Market and team carry the most weight by default — and you can adjust the weights before judging starts, then they lock.

Then there are six judges — the perspectives doing the measuring: Problem, Solution Logic, Business Value / Market, Pitch Quality, Team Readiness, and Feasibility. A judge isn't tied to a single dimension. Each one contributes to the dimensions it's responsible for through a routing matrix — full weight where it's the primary reviewer, less where it's a secondary one.

Dimensions keep the review from collapsing into one vague impression. A startup can be strong on team and thin on market evidence, and the report keeps those signals apart. A good story can't paper over a weak number — and you see exactly where each deck stands.

Step 3 — Six judges that don't talk to each other

The six judges run in parallel, and none of them sees the others' scores while it works. That independence is the point. The moment judges can see each other, they converge — the loud one anchors the room, and you get consensus that looks like agreement but is really just deference.

For each dimension it covers, a judge returns a score from 0 to 10, a confidence level, the evidence it leaned on, strengths, weaknesses, red flags, and the questions it would ask the team live. Six perspectives stay genuinely separate — the way a real panel is supposed to work, and rarely does.

Step 4 — The numbers come from math, not mood

Once the judges are done, a deterministic function does the arithmetic. It applies the routing matrix and the dimension weights and computes the per-dimension scores, the advisory AI Total Score, and the disagreement signal. No model is asked how it feels about the result.

Two consequences fall out of that:

  • It's reproducible. Run the same deck twice and you get the same numbers, because the math is a fixed function, not a fresh opinion. A ranking holds up when someone asks how it was reached.
  • The narrative can't invent numbers. The written report is generated after the math and has to build around it. If a dimension scored 7.2, the text cannot describe it as weak. Words explain the score; they never set it.

That order — math first, prose second — is the difference between an evidence-based report and a confident guess.

Step 5 — Disagreement is a feature, not noise

Averages hide conflict. A dimension where every judge gives 7.5 is not the same as one where one gives 5 and another gives 10 — but the average is identical.

So EvalLens measures the spread between judges on each dimension and labels it:

SpreadStatusWhat it means
below 1.5Consensusjudges agree
1.5 to 2.99Splitmoderate disagreement
3.0 and upConflictserious disagreement

A conflict isn't a problem to be smoothed over. It's a map. It points your attention to the exact decks and dimensions where the judges don't agree — which is precisely where your human judgment is worth the most. You spend your review time where it actually changes the ranking.

Step 6 — You make the final call

The governing rule is short:

AI Total Score is advisory.
The leaderboard ranks on your Jury Score.

You read the AI report, add live Q&A and the context only you have, then submit the final scores. The AI baseline stays on screen as explanation — it never becomes the decision. The heavy lifting is automated; the call, and the ranking, are yours.

Why this matters

EvalLens isn't trying to make pitch judging sound more scientific with heavier language. It solves a plain problem: evaluating startups at volume needs structure, traceability, and a result you can stand behind.

A score becomes credible when it's expensive to fake — when you can see what was measured, what evidence backed it, where the judges split, and who signed off. Across 1,000+ evaluation runs, that's the line that separates a ranking you can defend from a number you have to trust on faith.

A good evaluation system should show:

  • what was evaluated;
  • what evidence was used;
  • which dimensions were strong or weak;
  • where the judges disagreed;
  • who made the final decision.

That is the core of how EvalLens works — and why the lens stays in your hands.

The full rubric, judge routing, and scoring weights are documented on Methodology.

Want to see it run on your own batch? Book a demo.

Common questions

How does EvalLens score a pitch deck? A deck goes through a fixed pipeline. The Decoder reads the deck as it is, six independent AI judges score it across six dimensions, and deterministic math turns those judgments into an advisory AI Total Score. Then a human reviews the report and sets the Jury Score, which is what the leaderboard actually ranks on.

Does the AI decide the final ranking? No. The AI Total Score is advisory, and the leaderboard ranks on the Jury Score a human submits after reading the AI report and adding live Q&A and context. The AI baseline stays on screen as explanation, never as the decision.

What dimensions does EvalLens use to evaluate a pitch deck? Six: problem significance, solution differentiation, market attractiveness, business model and GTM, team and founder fit, and feasibility. Market and team carry the most weight by default, 0.20 each against 0.15 for the rest. Organizers can adjust the weights before judging starts, and then they lock.

Are AI pitch deck scores reproducible? In EvalLens, yes. The aggregation is a fixed mathematical function, not a fresh model opinion, so running the same deck twice returns the same numbers. The written report is generated after the math and has to build around it, which means the narrative can never invent a score.

What happens when the AI judges disagree? EvalLens measures the spread between judges on each dimension and labels it: below 1.5 is consensus, 1.5 to 2.99 is a split, and 3.0 or more is a conflict. A conflict is not smoothed over. It points you to the exact decks and dimensions where your human judgment changes the ranking most.