Telecom · Enterprise platform · 2026 · Under development

Zala Studio

A no-code studio where enterprises build AI agents and deploy each one into its own isolated runtime in seconds.

0:54
from Deploy click to a live agent
7–9s
engineering releases per agent change
0
from Publish to serving traffic
5 min
to bring a new tenant live
1 day

The problem

One hard-coded master agent served every tenant: a single prompt, twelve built-in skills and about 45 Python tools. Every new tenant was a fork in code and every new skill a release, so no agent changed without an engineering deploy.

Who it’s for
Tenant admins and operators at telecom and utility companies who build agents without code, and the customers those agents serve on chat and voice.
My role
Senior AI & ML Engineer
Team
3 leads, 5 devs
Timeline
May 2026

Stack

Runtime

  • DeepAgents
  • Firecracker
  • LangGraph

Studio

  • Flutter
  • Riverpod
  • Material 3

Backend

  • FastAPI
  • OPA
  • Keycloak

Data

  • PostgreSQL + pgvector
  • Redis
  • Memgraph
  • SQLite
  • MinIO

Channels & ops

  • LiveKit
  • Langfuse
  • Kubernetes

How it works

Studio

  • Create with AI
  • Canvas builder
  • Playground

Control plane

  • Admin API
  • Agent registry
  • Deploy factory

Channels

  • Web chat
  • Voice (LiveKit)

Runtime

  • Agent microVM
  • Deep Agents
  • OPA policy
  • State + memory

Ops

  • Approval inbox
  • Langfuse traces
  1. 01 Describe the agent in chat

    A tenant admin describes the job. Create with AI asks clarifying questions, then drafts the agent as an eight-node graph on the canvas.

  2. 02 Try it in the playground

    The playground boots the same bundle in its own pre-warmed microVM, so a test run behaves exactly like production.

  3. 03 Publish a versioned bundle

    Publishing saves AGENTS.md, config.json, tools.json and each SKILL.md as a new version: a Postgres row plus files in S3.

  4. 04 Deploy into its own microVM

    The factory injects the definition into a golden image and boots one Firecracker microVM per agent, ready in about 7–9 seconds.

  5. 05 A customer starts talking

    Chat or a LiveKit voice call reaches the agent’s endpoint. Deep Agents plans the turn and keeps each user’s thread and memory separate.

  6. 06 Every tool call is checked

    OPA allows, blocks or pauses every tool call. Paused calls wait in the inbox for a person, and Langfuse traces each turn.

Decisions

01

Agents as data

Chose Versioned agent definitions as Postgres rows and S3 bundles over a Git repository and CI pipeline per tenant.

Publishing becomes a row write instead of a release. Git per tenant meant per-tenant CI fan-out that the platform’s scale didn’t justify.

Trade-off: Git’s history and diffs don’t come free, so the studio needs its own version diff and rollback.

02

Golden image, config at boot

Chose One golden runtime image that loads each agent’s definition at boot over building a container image for every agent.

A deploy becomes a microVM boot, ready in about 7–9 seconds, and there is one hardened image to patch and scan, not one per agent.

Trade-off: The runtime has to interpret every spec, and tenants can’t ship their own dependencies.

03

Eight nodes, one status bar

Chose Eight canvas nodes and an always-visible status bar over a ninth canvas node for cost and evaluation.

Cost, P95 latency and eval pass rate stay in view on every screen instead of behind a node the builder must open.

Trade-off: It departs from the PRD’s nine nodes, so the spec stands in as an ADR until the PRD is amended.

What I’d do next

  • I’d connect the studio to the real runtime sooner. Create with AI still runs on a scripted stream and the playground on a mock runtime, so integration risk is still ahead.
  • I’d keep one architecture source of truth from day one: the README still describes a Django backend while the v2 design is FastAPI, and that drift misleads new engineers.
  • Next I’d build the form editor and ZIP export and import, so one bundle round-trips between the canvas, the form and a developer’s own editor.

Results

  • Agent change to live traffic

    Before: 2 weeksAfter: 5 min

  • New tenant to first conversation

    Before: 6 weeksAfter: 1 day

  • Deploy, definition to running agent

    Before: 15 minAfter: 8s

More work