Case study · 2026 to now · Founding engineer
Samora
A voice agent has to know when to hand the phone to a person, and survive the hand-off.
Samora (YC W26) runs AI voice agents for businesses across voice, SMS, WhatsApp, email and web chat, and keeps a human operator one click from any conversation. I’m on the early engineering team, working on the layer where software hands a customer to a person (takeover, handback, escalation) and on the reliability and observability that make live calls dependable.
- Channels
- 5: voice, SMS, WhatsApp, email, web chat
- Voice transports
- SIP · WebRTC · LiveKit
- Long-running work
- Temporal, idempotent transitions
- Platform scale
- 1M+ calls, per samora.ai
The problem
An AI agent can handle the easy bulk of calls. The remainder is where a business needs a person, and where naive systems fall apart: the caller is on hold, the operator is mid-call elsewhere, the offer rings out, the transfer leg drops, and nobody knows.
The job is to make “a person steps in” a first-class state of the system, with the same guarantees as everything else: it can be retried, recovered, observed and audited.
How it’s built
A simplified view of the parts I work in. Dashed borders are the recovery paths. They matter more than the happy path.
Scroll the diagram sideways →
- Hand-off is a state machine
- offer → ring → accept or time out → takeover → handback. Every transition is idempotent, so a retry or a duplicate webhook can’t assign one call twice.
- Escalations ride a Postgres-backed queue
- A ring-timeout reaper recovers offers nobody answered. Lost isn’t a possible outcome, only late.
- Long-lived work lives in Temporal
- Contact sequences run as durable workflows, with reconciliation logic that revives stale jobs, abandoned sessions and half-finished runs.
- One identifier joins everything
- A single call ID threads through Go services, the Python voice worker and the serverless fleet, so a bad call can be followed from first ring to last log line.
Decisions that cost something
01
Recover, don’t prevent.
A hand-off crosses a phone network, a browser tab and a human. Something will be lost. Rather than chase a protocol that can’t lose a message, I built a reaper that makes losing one cheap: find the stale offer, re-offer or release it, idempotently.
The costA stuck call can take one reaper interval to heal. In exchange, nothing stays stuck.
02
Make the teardown order explicit.
On hang-up, recordings and transcripts raced against media teardown. The fix wasn’t a retry; it was ordering: let egress settle, drain the transcript, then finalise artifacts durably.
The costA slightly longer tail on every call, in exchange for never losing the last sentence.
03
Evidence, or it didn’t happen.
The post-call evaluator is an LLM, so it’s held to a rule: every answer must cite the exact transcript segment behind it, and results are published as idempotent, typed records.
The costFewer clever answers, much easier audits.
What went wrong, and what it taught me
One day, a run of reliability fixes.
Calls were occasionally finishing without their final transcript, or losing a transfer leg mid-hand-off. In a single day I shipped the teardown-path fixes together: durable call-artifact finalisation, egress settling before media teardown, transcript draining, and recovery of lost transfer legs. The common thread was ordering, not retries. The lasting habit is to instrument the path end to end before changing it.