The scores are decided before the first pitch starts.
What a judge receives before the event moves score quality more than anything you can do after it. This page is the full briefing pack template: the six blocks to include, the 11 minute calibration with the strongest published result behind it, and the arithmetic for how much judging your panel can actually absorb. Take it and use it, no signup required.
Feedback after the event is the one fix that measurably fails.
Trained NIH reviewers scoring the same proposals agreed at essentially zero. Telling reviewers about their disagreement afterwards changed nothing. An 11 minute briefing before scoring nearly closed the gap. Three published results, one conclusion: the briefing pack is not paperwork, it is the intervention.
Six blocks, in the order a judge reads them.
Send it as one document, a week before the event, with the video linked at the top. Everything else about your event belongs in a separate logistics email, so the pack stays about one thing: how to score.
The rubric, with anchors
Not a list of criteria names. Each dimension needs an anchor description per band, so a judge can tell what a 3 looks like against a 7 before the first pitch. Our full pitch rubric with weights and anchors is published free on the rubric page, ready to drop into the pack.
The scoring procedure
Three sentences that do most of the work: cite the evidence behind every score, name the band before the number, and when the score sits on a boundary with evidence missing, take the lower band. Judges who follow this produce records, not impressions.
The disagreement rule
Define, in the pack, what happens when scores diverge. In EvalLens the rule is Spread: highest score minus lowest per dimension, where under 1.5 is consensus, 1.5 to 2.99 is a split, and 3.0 or more is a conflict that goes to discussion instead of an average. Judges who know disagreement is expected stop softening their scores toward the middle.
Conflicts of interest, upfront
A line asking each judge to declare investments, employment, mentorship or personal ties to any team, and the recusal mechanics when one exists. Collect it before scores exist, because a recusal after the leaderboard looks like damage control. A log of who scored what makes any recusal verifiable later.
The load and the schedule
Tell each judge exactly how many submissions they will read, how long one takes, and when breaks land. The arithmetic below shows why this block is not a courtesy: unplanned load is where scoring quality quietly dies.
The calibration exercise
An 11 minute video and one practice submission, done before event day. This is the highest ROI block in the pack, and the next section shows exactly how to build it.
Eleven minutes that moved agreement from 0.61 to
A randomized trial with 75 professors tested exactly one intervention: an 11 minute video explaining what each scale value means and what an inaccurate score costs (Sattler et al., PLOS ONE, 2015). Here is how to build yours.
- 01
Record one 11 minute video
Walk through your scale, one band at a time, in your own words with your own anchors. In the Sattler trial this format alone moved reviewer agreement from ICC 0.61 to 0.89. It costs a volunteer judge less time than watching one extra pitch.
- 02
Score one sample submission on camera
Take a real deck from a past event, walk the rubric, cite the evidence, name the band, pick the number. Judges copy what they see far better than what they read. One worked example beats three pages of instructions.
- 03
Show the cost of a wrong band
Trained reviewers in the same trial picked the correct rating band 74% of the time against 35% untrained, and spent more time on the criteria. Show one concrete example of a misplaced score changing a ranking, and the scale stops being decorative.
- 04
Lock independent scores before any discussion
Research on live review panels found scores shifting with the room, with the chair's interventions and even laughter measurably moving numbers. The fix is procedural: every judge submits scores before deliberation, the chair speaks last, and discussion resolves flagged conflicts instead of manufacturing agreement.
Promise your judges a number, not an evening
- 4 minwhat the MLH organizer guide budgets per project per judge: 2 minutes of demo, 1 for questions and scoring, 1 to walk to the next tableMLH organizer guide
- 13judges the MLH formula J = ceil(P x n x t / T) demands for its own worked example of a 500 attendee event with a 120 minute window, at three rounds per project. Plus 2 to 3 extra judges as a no-show bufferMLH organizer guide
- 30 minTechnovation's estimate for reviewing one submission properly, with judges committing to at least 5 submissions and around 3 hours including trainingtechnovation.org
- 1.5 hper business plan when written feedback is part of the ask, at Project ECHO's competition. Judges take 3 plans for roughly 6 hours total. This is the price of feedback, and the reason most events stop giving itprojectecho.org
When the formula asks for judges you do not have, the options are fewer rounds, a longer window, or a first read that lands before your judges start. EvalLens runs that first read: an AI panel scores every submission against your rubric and briefs each judge on what to verify, so panel hours go to decisions instead of triage. How the scoring holds up is documented on the methodology page, and the full pitch competition setup on the use case page.
What organizers ask before sending the
- How long should the briefing pack be?
- Short enough to be read: two pages plus the rubric, and an 11 minute video. The evidence is specific on this. An 11 minute training video raised reviewer agreement from ICC 0.61 to 0.89 in a randomized trial, while longer, unstructured briefings have no comparable result attached. Every block that does not change how a judge scores belongs in a separate logistics email.
- Do we still need a live calibration meeting?
- Usually not. The measured gain came from a recorded video, which every judge can watch on their own time, and a live room adds its own risks: panel studies found scores shifting with the chair's interventions and the room's reactions. If you do meet live, keep the rule that independent scores are locked before any discussion starts.
- How many judges do we actually need?
- Run the MLH formula against your own numbers: projects, times rounds per project, times minutes per project, divided by the judging window. For its worked example of a 500 attendee event and a 120 minute window it lands on 13 judges, plus a 2 to 3 judge buffer. If the result is more judges than you can recruit, the honest options are fewer rounds, a longer window, or a first read that arrives before judges do. That last one is the job EvalLens does.
- What if judges skim the pack and score how they always have?
- Make the pack load-bearing instead of attached. Put the anchors on the scorecard itself rather than in an appendix, require an evidence line next to every score, and open the event by scoring one practice submission together. A judge can skip a PDF, but not a form field that asks which slide the score came from.
- Can we fix scoring after the event with feedback instead?
- The evidence says no. In a randomized trial across two review years, reviewers who were shown how their scores diverged from others did not agree more the following year, with agreement stuck around ICC 0.30 to 0.40. What did improve was agreement on well defined, checkable questions. Both findings point upstream: calibrate before scoring, and write criteria a judge can verify rather than feel.
Send your judges a briefing, not a pile.
Build the pack from this page and it works on its own. Add EvalLens and every judge opens the event with the field already read: scores with evidence behind them, conflicts flagged by the Spread rule, and questions worth their minutes. AI prepares the analysis, the judges decide.