Abdul HaqueHire me
Back to my work
WRK-02 · Dec 2025 – Present · client work

AI Therapy Microservice

An AI service for an Australian government healthcare client that turns therapy session audio into transcripts and summaries, 300+ sessions a day, on CPU only.

Azure AKSFaster-WhisperPhi-3.5-MiniService Bus
Daily sessions300+through the pipeline
Memory per pod15 GBCPU-only, no GPU
Model precisionINT8quantised for CPU speed

The idea

The problem

Therapy sessions are recorded as audio. Turning an hour of speech into a transcript and a summary is slow, and a request that waits for the answer would time out. The data is also private, so it has to be handled with care at every step.

The answer

An event-driven service. The API accepts the audio and returns at once. A queue carries the work to background workers that transcribe and then summarise, so the slow part never blocks anyone.

How a session travels

Audio goes in one end and a transcript plus a summary comes out the other. Every step can fail and be retried on its own.

UPLOADQUEUETRANSCRIBESUMMARISEDELIVER
1. UPLOAD

The API verifies the caller's token against the identity provider's JWKS, stores the audio encrypted at rest, and answers straight away.

2. QUEUE

A message on Azure Service Bus describes the job. This decouples receiving audio from processing it.

3. TRANSCRIBE

A worker runs Faster-Whisper on CPU to turn speech into text.

4. SUMMARISE

Phi-3.5-Mini, quantised to INT8, writes the summary from the transcript.

5. DELIVER

The result is stored and the caller is told. MLflow records how each run behaved.

The hard parts

Running language and speech models on CPU, at volume, with private data.

No GPU

Faster-Whisper and Phi-3.5-Mini (INT8) run CPU-only on 15 GB pods in Azure Kubernetes Service. Quantising to INT8 is what makes that fast and small enough.

Lock contention and out-of-memory

Threads fought over locks and ONNX ran out of memory under load. A multi-processing worker pool gave each worker its own memory and removed the contention.

Failures are normal

Jobs that fail are retried, and the ones that keep failing land in a dead-letter queue, where they wait to be looked at instead of being lost.

Private by default

JWKS token verification on every call, audio encrypted at rest, and MLflow observability so that problems show up as numbers before they show up as complaints.

Built with

ComputeAzure Kubernetes Service, CPU-only pods with 15 GB of memory
ModelsFaster-Whisper for speech, Phi-3.5-Mini (INT8) for summaries
MessagingAzure Service Bus, with dead-letter queues and retry logic
APIFastAPI, JWKS token verification
ObservabilityMLflow