AI prepares the analysis · People decide

Every project gets a full read before the expo floor opens.

EvalLens is hackathon judging software that runs the first pass. An AI panel reads every submission against your rubric, scores execution and technical depth before anyone walks the tables, and hands each judge a briefing instead of a blank scorecard. Your judges still pick the winner.

The judging table, honestly

Your judges cannot read what your judges never see

  • 4 minall the time the MLH organizer guide budgets per project per judge: 2 minutes of demo, 1 for questions and scoring, 1 to walk to the next tableguide.mlh.com
  • 18judges the MLH formula demands for 175 projects in a two hour expo, at three rounds each: J = ceil(P x n x t / T)MLH organizer guide
  • 5%share of the projects the average judge actually saw at HackMIT, where 100 judges covered more than 200 projectsanishathalye.com
  • 70%share of a standard Devpost rubric riding on technical execution and innovation, assessed from a 3 minute demo video nobody is required to watch to the endinfo.devpost.com
Monday, in the Discord

“How did that project win?” cc: the sponsor.

A team that shipped a working build lost to a team that demoed well. Today you answer with a shrug, because four minutes at a table is genuinely not a review. With a record, the thread ends in one reply.

Score
8.1
on Execution and Demo. H1 carries 0.30 of the default weight
Finding
Two of three feature claims are demonstrated. The third is described, not shown.
Quote
“…deployed on a public endpoint, 40 test users…” · slide 6
Disagreement
Judges split on Technical Depth. The report flags the spread instead of averaging it away.
Human score
Your organizer's Jury Score, logged next to the AI read. The leaderboard is built from the human number.

The first submission and the last submission are read under exactly the same rules, which is more than any four minute table visit can claim. The AI score is advisory. The ranking comes from your judges.

How it works

Submissions close, judging starts already read

The AI panel does the first pass. Your judges keep the floor, the questions and the final call.

  1. 01

    Your rubric and tracks, locked

    Criteria, weights and tracks configured per event, plus a methodology line you can publish in the rules where fairness claims belong. You get: a rulebook your judges and sponsors can read before the doors open.

  2. 02

    Submissions land on your event page

    A public link or QR with a deadline and live statuses, or a manual batch you upload yourself. Completeness is checked automatically, so staff chases exceptions instead of the pile. You get: a clean field the minute the deadline hits.

  3. 03

    The hackathon panel does the first read

    Five reviewer roles, independent AI reads rather than people, score every submission on execution, technical depth, problem impact, innovation, UX clarity and delivery readiness. Execution and technical depth are weight protected, so a polished story cannot outrank a working build. You get: the whole field pre-read in hours.

  4. 04

    Every judge walks in with a briefing

    Per team: scores with the evidence behind them, quotes tagged to the slide they came from, what to verify at the table, and three questions worth the four minutes. You get: table visits that test the build instead of the pitch.

  5. 05

    The expo runs exactly as designed

    Same tables, same judges, same closing ceremony. Judges score as usual, and where AI reviewers disagreed the report says so, so deliberation starts at the real argument. You get: your judges' leaderboard, better informed.

  6. 06

    Leaderboard, then feedback for every team

    The ranking is built from human Jury Scores and your criteria weights. Structured feedback is drafted from the evidence and approved by your staff before it goes out. You get: teams that come back next year and tell people why.

Under the hood

A judge built for what teams actually shipped.

Five reviewers, execution weighted

Innovation, Technical Execution, Business Value, Pitch Quality and Feasibility read every submission independently, across six dimensions. Execution and Demo carries 0.30 of the default weight and Technical Depth 0.20, and both are protected. Weights are yours to set before the run and lock when it starts.

Today it reads the submission you already collect

The deck, the project description and the team's own notes, in the same intake you already run. Nothing changes for participants, and no judge loses a role.

Repository and live URL are next, and we say so

Reading a repo and a running demo end to end is the next build on the roadmap, not a claim we make today. When it ships you will hear it from us before you read it on a slide.

The gap the rubric never admits

A standard hackathon rubric puts most of the weight on technical execution and innovation, then asks a judge to grade both from a 3 minute video and a table visit. The first pass closes that gap before your judges ever have to.

When the Discord asks

You get the script, not just the software.

The “what we tell everyone” kit, included.

Hackers notice everything and post about all of it. The risk is never the tool. The risk is defending the tool with no script. So the setup ships with ready language for every audience.

The opening ceremony sentence“Every submission gets a full read under identical rules, and humans decide every placement.”
The rules page paragraphA methodology statement for your event rules: what the AI panel assists with, what the judges decide, and how a team can ask about its own record.
The submission form linePlain language on the form itself, so nobody discovers AI involvement after the results are announced.
The conflict of interest noteJudge conflicts and recusals stay your policy. The record simply logs who scored what, which is what makes a recusal verifiable later.
Data & team IP

The block your legal team reads first.

Never trained on

Team submissions are processed only for your event's evaluation and never used to train models. Contractual.

The event owns the record

Reports, scores and the decision log belong to your program. Retention and deletion follow your policy, and a DPA is available. Student data handling is structured to support your institution's obligations.

Procurement-ready

PO and invoice accepted, vendor registration forms and security questionnaires supported, public sub-processor list at /subprocessors, education discount for university programs.

Try it on last year’s field first. Send us a batch you already judged and compare the AI read against the placements you know. The first retro-test run is free through August 31, for batches up to 10 decks. Send us your batch or see pricing.

FAQ

What your judges and your Discord will ask

Does the AI pick the winners?
No, and there is no mode where it does. The AI panel produces an advisory read per submission. Your organizer sets the Jury Score per dimension, and the leaderboard is built from those human scores and your criteria weights. The AI number sits next to the human one for reference, and it never ranks anyone.
Does it review our teams' code and live demos?
Not yet, and we will not pretend otherwise. Today the panel reads the submission materials you already collect: the deck, the project description and the team's notes. Repository and live URL review is the next build on the roadmap. What you get today is a complete, consistent first pass across the whole field and a briefing that tells each judge exactly what to verify at the table.
Our judges are sponsors and alumni. We are not cutting them.
Nothing here cuts a judge. Judge count is a sponsor perk and a program KPI, so what changes is the ask, not the headcount. Sponsors came to meet builders and be seen, not to speed read 175 submissions on a Sunday. They keep the floor and arrive briefed.
We run five tracks with different criteria.
Tracks, criteria and weights are configured per event, and each track scores against its own rubric. Weights can be edited right up until the run starts, then they lock so the field is scored on one standard end to end.
Will our teams' projects train your model, and who sees them?
Never trained on, contractually. Submissions are processed only for your event, inside a closed perimeter with no public links. The sub-processor and model provider list is available for your IT review, and every report and score belongs to the event.
Can a team ask why it placed where it placed?
Yes. Every submission has a record: scores per dimension, the evidence and quotes behind them, where the AI reviewers disagreed, and the human score that decided the placement. Whether you run open appeals stays your policy. The record turns that conversation into five minutes instead of an archaeology dig.
Next step

Run the first pass on a field you already judged.

Send last year's submissions, or this year's batch before the expo. The panel reads every one on your rubric while your judges do what they always do. The first run is free through August 31, for batches up to 10 decks. AI prepares the analysis, your judges make the call.