Telecom · Internet provider · 2026 · In production

ZALA

One AI agent that answers an internet provider’s customers, and fixes their issues, on chat, phone and video.

0:53
of chats resolved without a human
61%
voice reply time, p50
1.2s
conversations a month
45k
average CSAT, out of 5
4.6

The problem

An internet provider’s customers waited about 8 minutes to reach a person. Around 70% of contacts were about bills, outages, technician visits and router faults that its own systems could already answer.

Who it’s for
About 150,000 home broadband and TV customers on web, mobile, phone and video, and the 40 support staff who take the conversations ZALA hands over.
My role
Senior AI & ML Engineer
Team
2 senior AI/ML engineers, 2 associates, 1 PM
Timeline
9 months, Aug 2025 – May 2026

Stack

Agent

  • GPT-5.1
  • LangGraph
  • DeepAgents

Voice

  • LiveKit
  • Deepgram
  • Rime
  • Silero VAD

Memory

  • Memgraph
  • Redis
  • FAISS

Apps

  • Next.js
  • Flutter
  • MQTT
  • Stripe

Platform

  • FastAPI
  • Langfuse
  • Docker

How it works

Channels

  • Web chat
  • Mobile app
  • Video call
  • Phone line

Transport

  • MQTT bus
  • LiveKit media
  • Voice worker

Agent

  • Master agent
  • Markdown skills
  • Tools

Memory

  • Redis sessions
  • Knowledge graph
  • Router manuals

Systems

  • Billing and tickets
  • Stripe checkout
  • Live agent
  • Langfuse traces
  1. 01 A customer asks, typed or spoken

    Chat arrives over MQTT. On a call, the voice worker transcribes speech and sends the text, with a camera frame, down the same topic.

  2. 02 One agent loads one skill

    The master agent reads the Markdown skill for the job, such as billing or diagnostics, after an entitlement check on the account.

  3. 03 It already knows the customer

    Each turn starts with the customer’s profile from the knowledge graph and the conversation from Redis, which chat and voice share.

  4. 04 Tools do the work

    Tools call the provider’s billing, ticket and field systems, search router manuals, and open Stripe checkout when a bill is due.

  5. 05 The answer streams back

    Tagged XML streams to the app so cards render field by field; on a call, tokens go straight to speech.

  6. 06 Traced, or handed to a person

    Every turn is traced in self-hosted Langfuse. When ZALA can’t help, it invites a live agent into the conversation.

Decisions

01

One agent, many skills

Chose One agent that loads a Markdown skill per task over an orchestrator handing off to 12 domain agents.

Every handoff was a cold start that lost the conversation. One agent keeps the context across billing, support and orders.

Trade-off: All 47 tool schemas share one context window, and access checks run when a skill loads, not on every tool call.

02

One brain, two mouths

Chose A voice worker with no model of its own over a separate voice agent running its own LLM.

Voice turns reach the same agent, session and checkpoint as chat, so a customer can switch between typing and talking mid-conversation.

Trade-off: Every spoken turn crosses the message bus twice, so latency adds up stage by stage.

03

Memory that can’t invent IDs

Chose A knowledge graph built from tool results over one model extracting every memory from the transcript.

The single-model version hallucinated account IDs and drifted its schema. Now IDs come only from tool results; a small model reads behaviour.

Trade-off: The graph runs alongside the older Markdown memory, so there are two memory systems to keep in step.

What I’d do next

  • I’d build a fixed conversation suite before changing the architecture, so comparing designs takes two weeks, not six.
  • I’d generate the XML reply format and the app’s parser from one schema, because free-form XML drifted in production.
  • I’d instrument voice turn-taking before tuning it: three of the five interruption bugs were library defaults nobody had read.

Results

  • Wait to first answer

    Before: 8 minAfter: 5s

  • Contacts resolved without staff

    Before: 12%After: 61%

  • Voice reply time, p50

    Before: 10sAfter: 1.2s

  • CSAT, out of 5

    Before: 3.7After: 4.6

More work