We sat down to write seven landing pages, one for each kind of program we work with. Accelerators. Grants and prizes. Corporate innovation. VC open calls. Pitch competitions. Angel networks. Crowdfunding platforms under the EU regime.
By the fourth page, something was off. The vocabulary kept moving. Cohort. Award cycle. Deal flow. Screening night. Underneath it, one sentence refused to change: more submissions arrive than the team can read properly, and somebody has to explain the outcome.
That was irritating for about a day, and then it was useful. When seven buyers who never meet each other describe the same week, they are not shopping for seven products. They are hiring one job, in seven vocabularies.
The short version

many submissions × limited review capacity × a decision you have to defend
All three, or it is not this job. Miss one and you usually need something else: another pair of hands, a different meeting, a shorter form. Everything below is how to run that test on your own program.
If you are arriving cold: EvalLens reads a batch of pitch decks and prepares an evidence-linked first read of each one, with reports, gaps, scores and questions. Your jury owns the final ranking. The test below is the same one we use to tell whether that will help your program at all.
Your segment is not your problem
Buyers shop by category. A grant officer searches for grant review software. An accelerator ops lead searches for application screening for accelerators. It is a reasonable instinct, and it is the wrong axis.
Category predicts your vocabulary. It does not predict your pain.
Two accelerators, one with 120 applications and one with 1,400, are running different problems with the same job title. An accelerator with 1,400 applications and a foundation with 1,400 applications are running the same problem in different clothes. What sorts programs into groups is not what they are called. It is the shape of the week before the decision.
That reframe cost us four landing pages to notice, so we may as well pass it on.
Seven rooms, one job
Same job, different moment. The moment is what makes it feel unique from the inside.
| The room | The moment | The job |
|---|---|---|
| Pitch competitions | Before finals day | Turn open submissions into a ranked finalist board |
| Accelerators | Before cohort selection | Compare applicants on one standard and defend the cohort |
| VC open calls | Before the pipeline meeting | Turn inbound decks into a partner-ready first read |
| Angel networks | Before diligence night | Know which decks deserve the room's time |
| Grants and prizes | Before funding decisions | Score against fixed criteria and keep a record that survives an appeal |
| Corporate innovation | Before the stakeholder review | Separate real partnership potential from innovation theatre |
| Crowdfunding platforms | Before a project is onboarded | Apply one documented standard to every applicant |
Our product docs carry two more rooms that behave the same way: hackathons before live judging, and universities before demo day. Same test, same answer.
Notice what the third column never says. It never says pick the winner. Every one of those jobs ends with a person in a room, holding a decision they have to say out loud.
Condition one: volume the calendar cannot absorb
This is the condition everybody leads with, and the only one people measure.
Size of the pile is the wrong tell. Watch what your reviewers do when it grows. Under pressure, careful readers do not read faster. They read less. DocSend's 2024 data clocks the average investor's first read at 2 minutes and 30 seconds per deck (DocSend), and their earlier research found that on a deck heading for a no, investors give up at 2
(TechCrunch).Two and a half minutes is what the calendar allows. Nobody chooses to skim slide nine.
Condition two: capacity that is fixed and expensive
Review capacity does not stretch. It is senior partners, volunteer judges, faculty, or a screening committee that meets on a weeknight because its members have day jobs.
You cannot buy more of that in the four weeks before a deadline. You can only spend it differently. The standard coping move is to split the pile: partners take thirty decks each, or an intern runs the first cut. That gets you through the batch by guaranteeing that no two submissions get the same reading, which is a debt you repay in condition three.
Meanwhile, of the hours your best reviewers give you, how many go to reading the eleven submissions that were never going to make it?
Condition three: exposure
This is the condition that decides fit, and the one buyers leave out of the brief.
Exposure means somebody can ask you why. An applicant, a board, a trustee, a regulator, an LP. And we liked them is not an answer you can give that person.
The reason this matters is that human review is less consistent than most programs assume. Across 23,414 ratings at a national science fund, agreement between independent reviewers came out at an intraclass correlation of about 0.26 (PLOS One). Two qualified experts, same proposal, different numbers. Unstructured expert judgment does that.
Training moves the number. In a controlled study, reviewers who went through a training program reached inter-rater reliability of 0.89, against 0.61 without it (PLOS One).
So consistency lives in the process, not in the people. Hiring better judges will not close a gap that opens every time the same proposal meets two readings. Giving every submission the same reading will.

When the answer is no
A test that cannot fail is marketing. Four honest cases where we are the wrong call.
You get five decks a week. Then you do not have a screening problem, you have a calendar problem. Read them.
Nobody asks you why. If your decision surprises no one and answers to no one, the record we produce is overhead you will not use.
You need the last mile, not the first. We prepare a screening read: the gate that decides whether something is worth real time. Deal terms, financial diligence, reference calls and the investment committee memo are a different document, further down, one deal at a time. If that is your bottleneck, this is not your tool.
You want an arbiter. Some buyers want the AI to decide so nobody has to own the call. We built it the other way around on purpose. The AI never picks the winner. If that is a dealbreaker, better to find out now than in month three.
What changes between rooms, and what does not
Once a program passes the test, most of the setup is the same. The differences are real but narrow.
What changes. The mode and its panel. A pitch competition is thesis-first: six judges across the six dimensions, demo optional. A hackathon is execution-first, with its own panel, and no demo means the entry is incomplete rather than badly scored. The criteria weights are yours, and they apply at the leaderboard, not inside the judges' reading. Intake changes too: your team can add entrants by hand, or you can open a public page and let teams upload their own materials.
What does not. Judges read independently and never see each other's scores. Every score stays connected to the evidence in the deck it came from. The AI Total Score is advisory, always. The leaderboard ranks on the Jury Score, which a human sets after reading the report. And the whole path stays reconstructible afterwards, which is the part that matters when someone asks about entry number eleven.
That reconstructible path is what programs are buying from us. The rest is configuration.

Run the test on your own program
Five things you can do this week, none of which involve talking to us.
- Count the batch and the weeks. Submissions in the cycle, divided by working days before the decision. That is your real reading time per submission, and most teams find it lower than they guessed.
- Count the readers, in hours. Not people. Hours of senior attention that actually exist, from people who will not quietly hand it to an intern.
- Name who can ask why. Write the actual names or roles. If the list is empty, stop here.
- Write the sentence you would have to say. "We ranked this one above that one because ___." If you cannot finish it with something a stranger could check, you are not shopping for speed. You are shopping for defensibility, and those are different purchases.
- Pick the batch you already argued about. Every program has one cycle where the shortlist was contested. That is the batch worth re-running, because you already know the answer and can see whether a structured pass would have surfaced it.
What we are not claiming
We have run more than 1,000 evaluations across programs, and we can show you the engine, the report and the audit trail on your own decks. What we cannot show you yet is a customer logo or a named case study, because we do not have one to show. When we do, it will be on this blog with numbers in it.
Until then, the honest offer is a retro-test on a batch you already decided, where you get to check our work against your own conclusions. That is a better proof than a testimonial anyway.
Through August 31, that retro-test is free, for batches of up to 10 decks. Not free-trial free, with a card on file and a timer. Free because right now your verdict on our work is worth more to us than the invoice.
Common questions
Does EvalLens work for my segment? Ask the three conditions rather than the segment: many submissions, limited review capacity, and a decision you have to defend. All three means the fit is real, whatever your program is called. Two out of three usually means something else would help more.
What does the retro-test cost? Through August 31, nothing, for batches of up to 10 decks. You bring a batch you already decided, we run it, and you compare our reports against the conclusions you already paid for the hard way. After August, it moves to a stated fixed fee.
Do you have a different product for each program type? No. One engine, configured per mode. What changes is the judge panel, the criteria weights and the intake path. What stays is independent judges, evidence-linked scores and a human-owned final ranking.
Does the AI decide who wins? No. The AI Total Score is advisory and the leaderboard ranks on the Jury Score that you set. AI prepares the analysis. You make the final call.
What stops our reviewers from rubber-stamping the AI score? The leaderboard is built from the Jury Score your reviewers enter, so there is no ranking until a person has made one. And the report hands them evidence to argue with, not just a number to accept: strengths, weaknesses, gaps, and the questions worth asking live. Disagreeing with the AI read is part of the design, not a malfunction.
We only run one cycle a year. Is that worth it? Often yes, because annual programs carry the highest exposure per decision and the least practice at defending it. Run the retro-test on last year's batch and decide from evidence.
If your program passes the three-condition test, the useful next step is not a demo of our features. It is a batch of yours. Book a demo, bring the cycle you already argued about, and run it as a free retro-test before August 31, up to 10 decks. See what a second reading finds.



