Build real-time voice AI agents on Azure OpenAI β ideal for call centers, contact centers, IVR replacements, and conversational AI. Two production-ready implementations with zero API keys, managed identity auth, and full deployment guides.
Created by Vinayjain@microsoft.com / Vinex22@gmail.com
A complete sample for building real-time voice AI agents using the Azure OpenAI GPT Realtime API. Talk to GPT models in real-time β like a phone call with AI. Perfect for:
- π Call center / contact center AI agents β replace or augment IVR with GPT
- π€ Conversational AI voice assistants β custom personality, real-time responses
- π’ Enterprise voice bots β VNet-compliant, managed identity, no API keys
- π§ Customer support automation β voice-in, voice-out, with tool/function calling
Two architectures included:
| π WebRTC β Direct | π WebSocket Proxy β VNet-Safe | |
|---|---|---|
| Audio path | Browser β Azure directly | Browser β your server β Azure |
| Server role | Mint token, serve page, go idle | Active relay for every audio frame |
| VNet / Private Endpoint | β Audio bypasses it | β Fully respected |
| Latency | ~200-400ms | ~300-600ms |
| Server framework | Flask | FastAPI + uvicorn |
| Function calling / Tools | β No server in the loop | β Server can intercept and execute |
| Best for | Low latency, demos, internal tools | Production, compliance, regulated industries |
The server mints a disposable token, then the browser talks directly to Azure OpenAI over WebRTC. Server goes idle.
All audio flows through your server. Browser never contacts Azure. Compatible with Private Endpoints and VNet injection.
- Prerequisites
- Project Structure
- Configuration
- Version 1: WebRTC (Direct)
- Version 2: WebSocket Proxy (VNet-Safe)
- Proving No Direct Communication
- Debug Tool (inspect_session.py)
- Troubleshooting
- Azure OpenAI / Foundry resource in East US 2 or Sweden Central
- A realtime model deployment (e.g.
gpt-realtime-1.5,gpt-4o-realtime-preview,gpt-realtime-mini) - Cognitive Services User role assigned on the resource for your identity (or managed identity)
- Python 3.12+
- az login completed (for local dev)
realtime-agent/
βββ .env # shared config (both versions read this)
βββ .env.example # template
βββ .gitignore
βββ README.md
βββ inspect_session.py # debug tool β fetches token from webapp, opens WS, prints events
βββ webrtc/ # Version 1: WebRTC (direct browser-to-Azure)
β βββ token_service.py # Flask server
β βββ requirements.txt
β βββ architecture.drawio # draw.io architecture diagram
β βββ arch-webrtc.png # PNG export of architecture
β βββ static/
β βββ index.html # browser client
βββ websocket/ # Version 2: WebSocket Proxy (all traffic through server)
βββ ws_proxy.py # FastAPI server
βββ requirements.txt
βββ architecture.drawio # draw.io architecture diagram
βββ arch-websocket.png # PNG export of architecture
βββ static/
βββ ws.html # browser client
Copy .env.example to .env and fill in:
AZURE_RESOURCE=your-resource-name # the <resource> in https://<resource>.openai.azure.com
REALTIME_DEPLOYMENT=gpt-realtime-1.5 # your realtime model deployment name
REALTIME_VOICE=marin # alloy|ash|ballad|coral|echo|sage|shimmer|verse|marin
REALTIME_INSTRUCTIONS=You are a helpful assistant. # system promptBrowser ββGET /tokenββ> Flask Server ββPOST /client_secretsββ> Azure OpenAI
Browser <ββek_tokenβββ< Flask Server <ββek_tokenβββββββββββββ< Azure OpenAI
Browser βββββββββββββββ WebRTC audio (direct) ββββββββββββββββ> Azure OpenAI
The server mints an ephemeral token (~60s lifetime), hands it to the browser, then the browser connects directly to Azure over WebRTC. Server is not involved after that.
cd webrtc
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txt
az login
python token_service.py| Route | Method | Purpose |
|---|---|---|
/ |
GET | Serves static/index.html |
/config |
GET | Returns { "azureResource": "..." } |
/token |
GET | Mints ephemeral token via /openai/v1/realtime/client_secrets |
/connect |
POST | Optional SDP proxy + background WebSocket observer |
# 1. Create resource group (skip if exists)
az group create -n realtime-agent -l centralindia
# 2. Create App Service Plan (Linux)
az appservice plan create -g realtime-agent -n rta-plan --sku B2 --is-linux
# 3. Create web app with user-assigned managed identity
az webapp create -g realtime-agent -p rta-plan -n <app-name> --runtime "PYTHON:3.12"
# 4. Create and assign user-assigned managed identity
az identity create -g realtime-agent -n rta-identity
MI_ID=$(az identity show -g realtime-agent -n rta-identity --query id -o tsv)
MI_CLIENT=$(az identity show -g realtime-agent -n rta-identity --query clientId -o tsv)
MI_PRINCIPAL=$(az identity show -g realtime-agent -n rta-identity --query principalId -o tsv)
az webapp identity assign -g realtime-agent -n <app-name> --identities $MI_ID
# 5. Grant Cognitive Services User on your OpenAI resource
SCOPE=$(az cognitiveservices account show -g <openai-rg> -n <openai-resource> --query id -o tsv)
az role assignment create \
--assignee-object-id $MI_PRINCIPAL \
--assignee-principal-type ServicePrincipal \
--role "Cognitive Services User" \
--scope $SCOPE
# 6. Configure app settings
az webapp config appsettings set -g realtime-agent -n <app-name> --settings \
AZURE_CLIENT_ID=$MI_CLIENT \
AZURE_RESOURCE=<openai-resource> \
REALTIME_DEPLOYMENT=gpt-realtime-1.5 \
REALTIME_VOICE=marin \
"REALTIME_INSTRUCTIONS=You are a helpful assistant." \
SCM_DO_BUILD_DURING_DEPLOYMENT=true \
ENABLE_ORYX_BUILD=true
# 7. Configure startup command, WebSockets, HTTPS
az webapp config set -g realtime-agent -n <app-name> \
--startup-file "gunicorn --bind=0.0.0.0:8000 --workers=2 --timeout=120 token_service:app" \
--web-sockets-enabled true \
--always-on true \
--min-tls-version 1.2
az webapp update -g realtime-agent -n <app-name> --https-only true
# 8. Deploy
cd webrtc
Compress-Archive -Path token_service.py, requirements.txt, static -DestinationPath deploy.zip -Force
az webapp deploy -g realtime-agent -n <app-name> --src-path deploy.zip --type zip
# 9. Verify
curl https://<app-name>.azurewebsites.net/tokenBrowser ββwss /wsββ> FastAPI Server ββwssββ> Azure OpenAI
Browser <ββaudioβββ< FastAPI Server <ββaudioββ< Azure OpenAI
Every audio frame passes through your server. The browser never talks to Azure directly. No Azure credentials or endpoints are exposed to the browser.
cd websocket
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txt
az login
uvicorn ws_proxy:app --port 5001| Route | Method | Purpose |
|---|---|---|
/ |
GET | Serves static/ws.html |
/ws |
WebSocket | Bidirectional relay: browser audio <-> Azure OpenAI |
# 1. Create resource group (skip if exists)
az group create -n realtime-agent -l centralindia
# 2. Create App Service Plan (Linux, B2 or higher for WebSocket perf)
az appservice plan create -g realtime-agent -n rta-plan --sku B2 --is-linux
# 3. Create web app
az webapp create -g realtime-agent -p rta-plan -n <app-name> --runtime "PYTHON:3.12"
# 4. Create and assign user-assigned managed identity
az identity create -g realtime-agent -n rta-identity
MI_ID=$(az identity show -g realtime-agent -n rta-identity --query id -o tsv)
MI_CLIENT=$(az identity show -g realtime-agent -n rta-identity --query clientId -o tsv)
MI_PRINCIPAL=$(az identity show -g realtime-agent -n rta-identity --query principalId -o tsv)
az webapp identity assign -g realtime-agent -n <app-name> --identities $MI_ID
# 5. Grant Cognitive Services User on your OpenAI resource
SCOPE=$(az cognitiveservices account show -g <openai-rg> -n <openai-resource> --query id -o tsv)
az role assignment create \
--assignee-object-id $MI_PRINCIPAL \
--assignee-principal-type ServicePrincipal \
--role "Cognitive Services User" \
--scope $SCOPE
# 6. Configure app settings
az webapp config appsettings set -g realtime-agent -n <app-name> --settings \
AZURE_CLIENT_ID=$MI_CLIENT \
AZURE_RESOURCE=<openai-resource> \
REALTIME_DEPLOYMENT=gpt-realtime-1.5 \
REALTIME_VOICE=marin \
"REALTIME_INSTRUCTIONS=You are a helpful assistant." \
SCM_DO_BUILD_DURING_DEPLOYMENT=true \
ENABLE_ORYX_BUILD=true
# 7. Configure startup command, WebSockets, HTTPS
az webapp config set -g realtime-agent -n <app-name> \
--startup-file "gunicorn -k uvicorn.workers.UvicornWorker --bind=0.0.0.0:8000 --workers=2 --timeout=120 ws_proxy:app" \
--web-sockets-enabled true \
--always-on true \
--min-tls-version 1.2
az webapp update -g realtime-agent -n <app-name> --https-only true
# 8. Deploy
cd websocket
Compress-Archive -Path ws_proxy.py, requirements.txt, static -DestinationPath deploy.zip -Force
az webapp deploy -g realtime-agent -n <app-name> --src-path deploy.zip --type zip
# 9. Verify
curl https://<app-name>.azurewebsites.net/For full VNet compliance, add these after the basic deployment:
# 1. Create VNet and subnets
az network vnet create -g realtime-agent -n rta-vnet --address-prefix 10.0.0.0/16
az network vnet subnet create -g realtime-agent --vnet-name rta-vnet \
-n app-subnet --address-prefixes 10.0.1.0/24 \
--delegations Microsoft.Web/serverFarms
az network vnet subnet create -g realtime-agent --vnet-name rta-vnet \
-n pe-subnet --address-prefixes 10.0.2.0/24
# 2. Integrate App Service with VNet
az webapp vnet-integration add -g realtime-agent -n <app-name> \
--vnet rta-vnet --subnet app-subnet
# 3. Create Private Endpoint for Azure OpenAI
az network private-endpoint create -g realtime-agent -n openai-pe \
--vnet-name rta-vnet --subnet pe-subnet \
--private-connection-resource-id $SCOPE \
--group-id account \
--connection-name openai-pe-conn
# 4. Create Private DNS Zone
az network private-dns zone create -g realtime-agent -n privatelink.openai.azure.com
az network private-dns link vnet create -g realtime-agent \
-n openai-dns-link --zone-name privatelink.openai.azure.com \
--virtual-network rta-vnet --registration-enabled false
az network private-endpoint dns-zone-group create -g realtime-agent \
--endpoint-name openai-pe -n openai-dns-group \
--private-dns-zone privatelink.openai.azure.com \
--zone-name openai
# 5. Route all outbound through VNet
az webapp config appsettings set -g realtime-agent -n <app-name> \
--settings WEBSITE_VNET_ROUTE_ALL=1Now the WebSocket proxy resolves <resource>.openai.azure.com to the private IP
and all audio traffic stays inside your VNet.
Fetches a token from the deployed WebRTC webapp, opens a WebSocket to Azure, sends a text message, and prints every event (including raw audio deltas).
# Set WEBAPP_URL to point at your deployed WebRTC app
$env:WEBAPP_URL="https://rta-25966.azurewebsites.net"
python inspect_session.pyFive ways to verify the WebSocket proxy version has zero direct browser-to-Azure traffic:
Open F12 β Network tab β start a session β filter by WS:
| WebRTC version | WebSocket Proxy version | |
|---|---|---|
Requests to *.openai.azure.com |
β
POST /realtime/calls |
β None |
| WebSocket connections | None | wss://<your-server>/ws only |
chrome://webrtc-internals |
Active RTCPeerConnection | Completely empty |
View source on ws.html β there is no Azure URL anywhere:
Select-String -Path websocket/static/ws.html -Pattern "openai|azure|foundry"
# Returns nothing. It only connects to ${location.host}/ws β your own server.Add to your hosts file temporarily:
127.0.0.1 foundry-multimodel.openai.azure.com
- WebSocket proxy: still works β (browser never calls Azure)
- WebRTC version: breaks immediately β (browser can't reach Azure for SDP)
Every audio frame is logged server-side:
[relay] browser -> Azure: input_audio_buffer.append (your mic audio)
[relay] Azure -> browser: response.output_audio.delta (model's voice)
If it were direct, the server would see nothing after initial connect.
| Evidence | WebRTC | WebSocket Proxy |
|---|---|---|
Network tab shows *.openai.azure.com |
Yes | No |
chrome://webrtc-internals has connections |
Yes | Empty |
| Blocking Azure DNS breaks it | Yes | No |
| Server logs show every audio frame | No | Yes |
| Browser JS contains Azure URLs | Yes | No |
| Problem | Fix |
|---|---|
401 on /token or WebSocket |
Ensure Cognitive Services User role is assigned on the OpenAI resource |
| 403 Forbidden | Resource must be in East US 2 or Sweden Central |
ModuleNotFoundError on App Service |
Set ENABLE_ORYX_BUILD=true and SCM_DO_BUILD_DURING_DEPLOYMENT=true |
| No audio output in browser | Check browser isn't muting the tab; click page first for autoplay policy |
| Audio overlapping | Update to latest ws.html with stopPlayback() barge-in fix |
| WebSocket closes immediately | Ensure --web-sockets-enabled true on App Service |
| High latency on WebSocket proxy | Use B2+ plan; consider region closer to Azure OpenAI resource |

