Every conversation about EvalLens starts with pitch decks. This one starts with squats.
We signed a paid pilot with an operating fitness platform: roughly 1,000 active users across two coaching programmes, built and run by its founder, a coach with a methodology her users actually follow. Production integration is underway, and the pilot itself runs for four months starting September 30. We are not naming the platform yet. That announcement is hers to make, on her timeline. What we can talk about is what we are building, because it says a lot about what EvalLens actually is.
Coach's methodology → Rubric → Evidence-linked progress reviews → AI coach (her digital twin) → Human sign-off
AI prepares the analysis. The coach makes the call. Same sentence we always say, new room.
Why a pitch evaluation company is scoring workouts
EvalLens is evaluation infrastructure: a rubric that encodes how an expert judges, evidence attached to every conclusion, and a human who owns the final decision. The easy read on us is "a tool that scores pitch decks." We understand the read. It is also wrong in a useful way. A pitch deck happens to be the first thing we pointed the infrastructure at. But the shape of the problem is everywhere: work gets submitted, a standard exists, and someone has to stand behind the judgment.
A fitness programme has exactly that shape. A user logs workouts and check-ins. The coach has a methodology that defines what good progress looks like: which sessions count, how consistency is judged, when to progress and when to ease off. And somebody has to look at the data and say, honestly, how it is going. The catch is scale. Industry benchmarks put day-30 retention for a typical fitness app near 10 percent, and coached programmes beat that number for one reason: accountability from a person who actually reads your work. Which is exactly the part that does not scale. One coach's attention does not divide by 1,000.
That is not a fitness problem. That is an evaluation problem.

What we are building
Six layers, the same architecture we run on pitch decks, tuned to a new subject:

The one people ask about first is the twin.
A twin, not a replacement
The AI coach does not have opinions of its own. It has hers. Its guidance comes from her methodology and from the evidence-linked review of each user's actual progress, and it works inside hard guardrails: training and habit guidance only, no medical or injury advice, and a defined escalation path to the human coach for anything outside that line.
Here is what that looks like for one member. She completes her workouts, logs how the body feels, and misses two check-ins. The system does not invent a verdict. It reads that record through the coach's rubric, prepares a review where every conclusion points at something in the log, and then does one of two things: offers guidance it can back with that record, or routes the case to the coach because it crossed the agreed line. The member gets attention every cycle. The coach gets her hours back for the cases that actually need her, and the final word on anything that matters stays hers, on the record.
The point is not to automate coaching. It is to make her coaching standard available to every user on the platform without diluting it. Until now, the ceiling of her attention was the ceiling of her business. The pilot is designed to raise that ceiling without touching the signature.
The pilot starts with a representative cohort of 100 users drawn from both programmes, with room to scale to the full base. And rather than describe the bar in adjectives, here is the actual bar. These are the benchmarks the pilot has to clear:
| What we measure | The bar |
|---|---|
| Users receiving a structured review in each cycle | 90% or more |
| Programme adherence vs. the pre-pilot baseline | +15% |
| Coach time per user review | 5 minutes or less |
| Assessments reconstructible from the record | 100% |
| Outputs rejected by the human reviewer | 5% or less |
| End-user rating of the AI coach | 4.0 out of 5 or higher |
Targets, not results. We publish them now so that when the report card comes, there is something to check it against.

What this means for EvalLens
This is our first commercial pilot outside startup evaluation, and it changes the way to read the company.
Pitch decks are still the wedge. The 1,000+ evaluation runs behind us, the live events, the fund work: none of that goes anywhere. But the machinery underneath was never deck-specific. Encode a methodology, attach evidence, keep a human on the final call. That works wherever expert judgment needs to scale without becoming a black box, and a fitness platform with a real base and a real methodology is a very honest place to prove it.
There is a simple test for whether any of this applies to you. Three conditions: you have a methodology that is genuinely yours, not a content library; your users submit real work, whether that is logs, check-ins, drafts or decks; and somewhere in your product a judgment gets made that a person has to stand behind. If all three hold, you have the same problem shape this pilot is built on.
Commercial terms after the pilot are usage-based, so the economics scale with actual use rather than promises. That part we designed as carefully as the pipeline.
Common questions
Can EvalLens evaluate things other than pitch decks? Yes. EvalLens is evaluation infrastructure: an expert methodology encoded into a rubric, evidence-linked AI analysis, and a human who makes the final decision. The first commercial pilot outside startup evaluation applies this to user progress on an operating fitness platform with about 1,000 users.
Who is the pilot with? An operating fitness platform with roughly 1,000 active users across two coaching programmes. We are not naming it yet; the announcement belongs to its founder.
Does the AI coach replace the human coach? No. The AI coach is grounded in the human coach's own methodology, works inside strict guardrails, and escalates anything sensitive to her. No assessment with a material consequence for a user is issued without human review, and that sign-off is recorded.
When will EvalLens share results? The pilot runs four months from September 30. After it closes, we will publish what the targets were and how the numbers came out, in this Newsroom.
The quiet part
Somewhere in the last year, "we put AI in the product" stopped being news. The question that matters now is who answers for what the AI says. In this pilot the answer is written into the architecture: a named coach, her methodology, her sign-off, on the record.
If your platform has a methodology and an audience, and you want the analysis without giving up the final call, book a demo. Bring your rubric. We will bring the receipts.
And if you are just here for the scoreboard, stay close to this Newsroom. The pilot closes in late January, and the numbers land here first, against the exact table above.



