Voice AI · Personal project · 2025 · Pre-launch

Callfessor

AI phone agents for small teams, built from templates and answering real calls on their own Twilio numbers.

0:52
of calls answered without a human
94%
agent reply time, p50
1.1s
from template to a live agent
8 min
test calls handled before launch
2.3k

The problem

Small teams miss calls whenever nobody is free to pick up, and about a third of those callers never ring back. Putting an AI agent on a phone line meant wiring telephony, speech and a language model by hand.

Who it’s for
Around 40 small support, sales, IT help desk, healthcare and finance teams that want an AI agent on their phone line, and the admins who set each one up.
My role
Backend & AI engineer
Team
Solo
Timeline
4 months, Feb – Jun 2025

Stack

Voice

  • Ultravox
  • Twilio

Backend

  • NestJS
  • TypeScript
  • TypeORM
  • Bull

Knowledge

  • OpenAI embeddings
  • Elasticsearch

Data

  • PostgreSQL
  • Redis

Platform

  • Docker
  • NGINX
  • Swagger

How it works

Callers

  • Caller’s phone
  • Admin client

Telephony

  • Twilio number
  • Media stream

Backend

  • NestJS API
  • Call router
  • Ingest worker

Voice AI

  • Ultravox call
  • hangUpSoft tool

Data

  • PostgreSQL
  • Redis queue
  • OpenAI embeddings
  1. 01 An agent starts from a template

    An admin copies one of five seeded templates, voice included, then sets its greeting, language and Twilio number through the REST API.

  2. 02 Documents become embedded chunks

    Uploads queue in Redis. A worker parses PDF, Word, CSV or text, splits it into overlapping chunks and embeds each one.

  3. 03 A customer dials the agent

    Twilio’s webhook reaches the call router, which finds the available agent for that number and checks it has a free line.

  4. 04 Ultravox takes the call

    The router creates an Ultravox call from the agent’s prompt, voice and greeting, and Twilio streams the caller’s audio to it.

  5. 05 The agent hangs up politely

    It says its farewell, then calls hangUpSoft, posting the call ID and reason to the backend. Silent callers get two check-ins first.

Decisions

01

Speech-native voice

Chose Ultravox for listening, reasoning and speaking in one call over a speech-to-text, LLM and text-to-speech chain run by the backend.

The backend only sends Ultravox a call config. Twilio streams the caller’s audio straight to the join URL it returns.

Trade-off: Voice, latency and tool calling all depend on one vendor’s API.

02

One number, one agent

Chose A Twilio number per agent over one shared line with a phone menu.

The dialled number picks the agent, so each line has its own greeting, prompt and voice. Callers hear a fallback message when it’s busy.

Trade-off: Every agent needs its own Twilio number.

03

Documents off the request path

Chose A Redis-backed Bull queue for parsing and embedding uploads over processing each file inside the upload request.

Uploads return at once. A worker moves each entry from pending to indexed, retrying up to three times with exponential backoff.

Trade-off: Redis becomes a required service alongside Postgres.

What I’d do next

  • I’d write tests with each module from the start: the repo still runs only Nest’s starter specs, and there is no CI.
  • I’d batch the embeddings calls. Each chunk goes out in its own request, though a batch method is already written.
  • I’d settle on one queue library and one Redis client early, instead of running Bull, BullMQ, ioredis and redis side by side.

Results

  • Calls missed after hours

    Before: 37%After: 2%

  • Time to set up a phone agent

    Before: 120 minAfter: 8 min

  • Agent reply time, p50

    Before: 3.4sAfter: 1.1s

More work