Product research · Startup · 2025 · Early access

Growth Genie

Product teams test an idea on AI digital twins of their users and get a scored verdict in minutes.

0:35
correlation with live-market results
>95%
faster product decisions
10×
cost saving over live tests
90%
per experiment, versus $10k+ studies
$10–50

The problem

Product teams tested an idea by recruiting panels, running surveys or launching a live A/B test. Each study cost $10,000 or more and took weeks, and a bad idea could surface only after launch.

Who it’s for
Founders and product, growth and marketing teams who want to test an idea on AI twins of their users before they build it.
My role
Sole backend and AI engineer
Team
2 engineers
Timeline
9 months, Nov 2025 – Sep 2026

Stack

Agents

  • LangGraph
  • GPT-4o mini
  • LangChain

Services

  • FastAPI
  • NestJS
  • Socket.IO

Research

  • GPT-4o vision
  • Playwright
  • Serper
  • Tavily

Data

  • PostgreSQL
  • Redis
  • SQLite

Platform

  • Docker
  • Railway
  • Next.js
  • Cloudinary

How it works

App

  • Web app
  • NestJS API
  • Live updates

Understand

  • Intent agent
  • Fetcher agent
  • Clarifying question

Panel

  • Panel generator
  • User approval
  • Market research

Review

  • Persona reviewers
  • Referee
  • Scoring engine

Results

  • Report
  • PostgreSQL
  • Redis pub/sub
  1. 01 A team submits an idea

    Text, files and links go through the NestJS API, which replies at once and streams each reviewer’s status back over a WebSocket.

  2. 02 Agents read what was sent

    An intent agent works out what to test and asks one question if it’s ambiguous; a fetcher agent crawls pages and reads images.

  3. 03 A panel is drawn and approved

    A generator builds 3 to 12 reviewers who each hold a seat in that business; the user approves them, and research checks the claims.

  4. 04 Every reviewer judges in parallel

    Each persona writes a review with a verdict per criterion, and a referee sends weak reviews back for revision, up to three times.

  5. 05 Python does the maths

    The engine checks each verdict’s quote against the submission and research, scores each reviewer 0–100, and the API averages the panel.

  6. 06 The report lands live

    The report is saved to PostgreSQL, and each reviewer’s status goes out over Redis to the user’s screen as it finishes.

Decisions

01

A score anyone can re-derive

Chose A Python score computed from per-criterion verdicts over letting the model write its own score out of 100.

A score scraped from the model’s prose couldn’t be audited or explained. Now it’s arithmetic on verdicts and evidence that anyone can re-check.

Trade-off: Rubrics, verdict bands and weights became code to maintain, covered by 113 tests.

02

Evidence, or it counts for less

Chose Fuzzy matching of every quote a verdict cites over trusting the evidence the model says it found.

Strong verdicts often arrived with no evidence at all. Now a verdict whose quote isn’t in the submission or research is discounted.

Trade-off: An honest paraphrase fails the match too, so it’s discounted along with invented quotes.

03

Same judges for every revision

Chose The original reviewer panel for every revision over drawing a fresh panel on each run.

Each reviewer re-weights the rubric, so a new panel changes the maths and the score change stops meaning better or worse.

Trade-off: A pivoted idea keeps reviewers that no longer fit, so users got an opt-out to draw a new panel.

What I’d do next

  • I’d check scores against a fixed set of ideas rated by human experts on every change, not only against the scoring arithmetic.
  • I’d use database migrations from day one: TypeORM’s schema sync once reset every execution ID to null on a restart.
  • I’d run those 113 tests in CI on every push; today they run only by hand.

Results

  • Cost per experiment

    Before: $10,000+After: $10–50

  • Time to a validated idea

    Before: weeksAfter: minutes

  • Ideas tested before build, a month

    Before: 2After: 30

More work