דלג לתוכן הראשי

Automation & AI

From pilot to production: building a voice AI lead pipeline you can trust

Daniel Eliyahu Bellelli··11 min read

A practical architecture for voice-agent leads: webhook verification, structured extraction, ICP scoring, retries, privacy and failure alerts.

A voice agent that sounds impressive in a demo is not yet a system. The system is what happens after the call: whether the lead is saved, whether its details are extracted correctly, what happens when a provider returns 429, and whether the team learns about a broken flow before a prospect notices that nobody called back.

A production pipeline needs explicit boundaries between collection, verification, extraction, storage and notification. The most important decisions often become visible only when the flow encounters real traffic and partial failures.

The pipeline stages

  1. Conversation: a Hebrew voice agent collects intent, business context and contact details.
  2. Webhook: the completed conversation is sent to a dedicated public endpoint.
  3. Verification: validate the HMAC signature before processing; reject invalid signatures with 401.
  4. Extraction: convert the transcript into structured fields using a strict schema.
  5. Scoring: assign a 0–10 fit score and an operational category.
  6. Storage: write the record under explicit row-level access policies.
  7. Notification: send the team a push or email notification with a direct link to the record.

Webhook verification is not optional

A public endpoint that accepts JSON and creates records is a data-contamination risk. Logging a warning when a signature is missing does not provide verification. Missing or invalid signatures must stop processing before any record is created.

Calculate the signature over the raw request body before parsing JSON: reformatting changes the bytes and invalidates the comparison. Compare signatures in constant time where supported, and apply the provider's timestamp rules to prevent replay of old requests.

Structured extraction: the schema comes first

Spoken Hebrew includes self-corrections, numbers expressed as words and phonetically transcribed names. A language model can handle this, but it needs a defined output schema rather than a request for free-form prose. Validate the result on the server and reject invalid output instead of silently guessing how to repair it.

Separate name, phone, email, organization, sector, main pain point, urgency and any budget actually mentioned. Missing information should be explicit null, not an empty string or an invented value. A transcript is evidence, not permission to fill in plausible details.

ICP scoring turns a list into a work queue

A team receiving dozens of leads needs a treatment order, not merely a chronological list. Define a 0–10 fit score using clear criteria: sector fit, organization size, a concrete pain point, urgency and decision-making ability. Add a short category such as ready for a demo, needs clarification or not relevant.

A useful score is explainable. If the team cannot see why a lead received 8 rather than 5, it will stop trusting the ranking.

Reliability: every external call eventually fails

Transcription, model and email providers will eventually return 429 or 503. Retry transient failures with exponential backoff and jitter. Honor Retry-After when present, and do not retry ordinary invalid-input or permission errors: an incorrect request does not improve on its next attempt.

A final failure must trigger an alert in a channel someone actually reads, including a safe correlation identifier and the failed stage. Keep lead storage independent of notification delivery so an email outage does not lose a valid lead.

Privacy: what must stay out of logs

Lead-processing logs can expose phone numbers, email addresses and full conversations. Mask personal information before logging: retain only a limited phone suffix or an abbreviated email local part when needed for diagnosis. Keep the complete transcript in a controlled location with explicit access and retention rules.

What to measure after launch

  • The share of completed conversations that result in a saved record.
  • The share of extracted results rejected by schema validation.
  • Median time from conversation completion to team notification.
  • Score distribution, checked periodically against human judgment.
  • Final failures after retries, grouped by provider and pipeline stage.

Conclusion

The difference between a demo and a production system lies at the boundaries: verification, validation, retries, privacy and alerts. Those shared safeguards make the pipeline dependable without someone inspecting it manually every morning.

More articles