The spreadsheet was right, until batch 200.
A form, a spreadsheet and a ChatGPT tab is the most common application review stack in the world, and for good reasons. EvalLens exists for the season it stops working: the batch outgrows the sheet, the prompt returns a different score every run, and someone finally asks why a deck got a 4.2. AI prepares the analysis, your team decides.
The DIY stack is not a mistake. It is stage-appropriate.
Most programs we talk to run exactly this: a form for intake, a sheet for scores, a chat tab for a second opinion. Before anyone tells you to replace it, here is what it gets genuinely right.
It costs nothing
A form, a sheet and a chat tab are free or already paid for. For a first season with forty applications, that is not a compromise. It is the correct engineering decision, and we would make the same one.
It bends to any process
No vendor rubric, no onboarding call, no waiting for a feature. Add a column, change a formula, rewrite the prompt over lunch. While your process is still being invented, that flexibility beats any product.
Everyone already knows it
Zero training. Every reviewer on earth can open a spreadsheet, and by now most have a ChatGPT tab open next to it. A 2026 Affinity survey of roughly 300 dealmakers found 85% already use AI daily.
Every number here comes from someone who hit the wall
- 200applications at which Kyle Taylor of Launchpad LA looked at his spreadsheet and, as he told Zapier's blog, "knew that was not a good system"zapier.com/blog
- 60%score variation between runs when grading vendor MarkInMinutes sent the exact same essay and rubric to ChatGPT three timesmarkinminutes.com
- 85%of dealmakers already use AI daily, per the 2026 Affinity survey of roughly 300 respondents. The question is no longer whether AI reads applicationsAffinity 2026, via developmentcorporate.com
- 12%of institutional funds have AI screening actually running in production. That gap between daily use and production process is this whole pagedevelopmentcorporate.com
Three seams, and they all tear the same season.
The stack does not fail loudly. It fails as a copy-paste error here, a re-sorted sheet there, and a score nobody can reconstruct at the exact moment somebody asks for it.
The seam between the tools is a person
The form, the sheet and the chat window do not talk to each other. Someone exports the CSV, fixes the columns, pastes each application into the prompt and copies the answer back. An operator of a nonprofit grant program wrote in the Airtable community in June 2026 that her Zapier to Airtable connection "does break fairly regularly" and she has to go reconnect it. Telling detail from the same thread: nobody in it had a scoring workflow at all, only deadline tracking. The integration layer of the DIY stack is a human, and humans take weekends.
The sheet has a ceiling, and it is around 200 rows
At Launchpad LA, applications were readable one by one, then suddenly they were not. "People were putting unstructured notes next to the application. As you try to scale that, it becomes impossible," Kyle Taylor told Zapier's blog. The failure modes are boringly specific. ScoreJudge, a judging software vendor, keeps a whole catalog of them: two judges typing in the same cell where whoever saved last wins, formulas that silently break when a column shifts, a judge accidentally sorting the sheet, no per-judge breakdown in the export. None of these are exotic. All of them happen the week the batch doubles.
ChatGPT scores with confidence and leaves no trace
MarkInMinutes, a grading tool vendor, ran a documented test: the same essay and rubric, submitted three times, came back with scores varying by up to 60%. Justifications were of the "demonstrates good understanding" kind, citing nothing, and some feedback referenced content that was not in the work at all. Practitioners on Hacker News made the same point in July 2026: an LLM judge hallucinates facts in both directions, inventing flaws and missing real ones. And as Development Corporate put it the same month, a ChatGPT prompt in a private Slack channel has no versioning, no calibration data and no audit trail. The output looks like a decision. Nothing behind it can be reproduced, checked or defended.
Same operations, different machinery.
Not a feature checklist. These are the six operations every screening season actually consists of, and what carries each one in either setup.
| Operation | Forms + Sheets + ChatGPT | EvalLens |
|---|---|---|
| Intake | The form exports a CSV. Someone pastes it into the sheet, fixes the columns by hand and chases missing attachments over email. | Applications land as one batch with completeness checked automatically. Your existing form stays. EvalLens is a layer on top of intake, not a replacement for it. |
| One rubric for every application | The rubric lives inside a prompt that gets edited, shortened and re-pasted. Application 8 and application 180 are quietly graded against different questions. | Criteria and weights are set before the run and lock when it starts. Every application in the batch is scored against the same frozen rubric, and the weights are visible on the leaderboard. |
| Reading every application in full | Reviewers are sharp for row 3 and skimming by row 150. The tail of the sheet gets a different quality of attention than the top. | A panel of independent AI reviewers gives submission 1 and submission 200 the same full read under identical rules. Order and fatigue stop being scoring factors. |
| Evidence behind the score | A 7 with "seems strong" in the next cell, or a ChatGPT justification that cites nothing from the actual document. | Every score is tied to quotes pulled from the submission itself. When a claim has no evidence in the materials, the report says so instead of inventing some. |
| Re-scoring after a rubric change | Re-paste everything and get new numbers with no way to tell whether the rubric changed or the model's mood did. | Re-run the batch under the new weights. The aggregation math is deterministic, so what changed in the scores traces back to what changed in the rubric. |
| The trail behind "why did this score 4.2" | Scroll a chat history and hope the thread was not deleted. There is no per-application record to show a founder, a board or a sponsor. | A per-criterion record: scores, the evidence behind them, where reviewers disagreed, and the human decision logged on top. The answer takes five minutes, not an excavation. |
What does not change:people make every decision. The panel’s output is a screening memo per application, a first gate that gets each one a full and consistent read. The shortlist is built from human calls, with the advisory AI score logged next to them.
When the spreadsheet is still the right answer.
Keep the sheet if all of this is true.
If that is your program, close this tab with our blessing. Bookmark it for the season the batch doubles, because per the operators quoted above, that is roughly when the sheet stops being a system.
What teams on the DIY stack ask us
- Is EvalLens just ChatGPT with a nicer interface?
- No, and the difference is exactly the failure modes on this page. A single chat prompt gives one pass with, per MarkInMinutes' test, up to 60% score variation between runs on the same input. EvalLens runs a panel of independent AI reviewers against a rubric that locks before the run, ties every score to quotes from the submission, records where reviewers disagreed instead of averaging it away, and aggregates with deterministic math. Same batch, same rubric, same math, and a record you can open afterwards.
- Our prompt works fine. Why change anything?
- It probably does produce plausible output, which is exactly the trap. As Development Corporate wrote in July 2026, a standing prompt in a private Slack channel has no versioning, no calibration data and no audit trail, but it produces the same-looking output as a governed process. The day a founder, a board member or a sponsor asks why an application scored what it scored, plausible is not the standard. Reproducible is. You are not replacing your prompt because it is bad. You are replacing it because it cannot testify.
- Do we have to abandon our form and our spreadsheet?
- No. EvalLens sits on top of the intake you already run. Keep your form, keep the sheet as your working view if you like. Applications come in as a batch, the panel does the first read, and results are yours to export. What goes away is the copy-paste seam and the untracked scoring, not your tools.
- Does the AI decide who advances?
- No, and there is no mode where it does. The panel produces a screening memo per application: scores, evidence, and flagged disagreements. It is a first gate that gets every application a full, consistent read. The advisory AI score sits next to the human score, and the shortlist is built from human decisions. AI prepares the analysis, your team decides.
- How do we test this without committing to anything?
- Send a batch you already reviewed on your spreadsheet and compare the panel's read against the decisions you know. The first run is free through August 31, for batches up to 10 decks. If the sheet was giving you the right answers, you will see that too, and this page told you to keep it.
Test it against the sheet you already trust.
Send a batch you already reviewed and compare the panel's read against the decisions you know. The first run is free through August 31, for batches up to 10 decks. AI prepares the analysis, your team makes the call.