AI Therapy Microservice
An AI service for an Australian government healthcare client that turns therapy session audio into transcripts and summaries, 300+ sessions a day, on CPU only.
The idea
The problem
Therapy sessions are recorded as audio. Turning an hour of speech into a transcript and a summary is slow, and a request that waits for the answer would time out. The data is also private, so it has to be handled with care at every step.
The answer
An event-driven service. The API accepts the audio and returns at once. A queue carries the work to background workers that transcribe and then summarise, so the slow part never blocks anyone.
How a session travels
Audio goes in one end and a transcript plus a summary comes out the other. Every step can fail and be retried on its own.
The API verifies the caller's token against the identity provider's JWKS, stores the audio encrypted at rest, and answers straight away.
A message on Azure Service Bus describes the job. This decouples receiving audio from processing it.
A worker runs Faster-Whisper on CPU to turn speech into text.
Phi-3.5-Mini, quantised to INT8, writes the summary from the transcript.
The result is stored and the caller is told. MLflow records how each run behaved.
The hard parts
Running language and speech models on CPU, at volume, with private data.
Faster-Whisper and Phi-3.5-Mini (INT8) run CPU-only on 15 GB pods in Azure Kubernetes Service. Quantising to INT8 is what makes that fast and small enough.
Threads fought over locks and ONNX ran out of memory under load. A multi-processing worker pool gave each worker its own memory and removed the contention.
Jobs that fail are retried, and the ones that keep failing land in a dead-letter queue, where they wait to be looked at instead of being lost.
JWKS token verification on every call, audio encrypted at rest, and MLflow observability so that problems show up as numbers before they show up as complaints.
Built with
| Compute | Azure Kubernetes Service, CPU-only pods with 15 GB of memory |
|---|---|
| Models | Faster-Whisper for speech, Phi-3.5-Mini (INT8) for summaries |
| Messaging | Azure Service Bus, with dead-letter queues and retry logic |
| API | FastAPI, JWKS token verification |
| Observability | MLflow |