Building a real-time chat widget seems straightforward until you introduce multiple browser tabs, spotty mobile networks, server restarts, and human agent takeover.
When designing the live handover architecture for mchatly, the primary requirement was zero dropped messages during handoff from an automated AI agent to a live human operator. Here are the core patterns that made it resilient.
The Dual-Channel State Problem
In a standard AI chat session, HTTP streaming (Server-Sent Events) is often enough. But the moment a customer requests a human agent, the paradigm shifts from request-response to bi-directional event orchestration.
To keep things decoupled:
- AI Chat Mode: Standard SSE stream with client-side optimistic rendering.
- Live Handover Mode: Upgrades connection to stateful WebSockets via Ably / Socket.io with Redis Pub/Sub backplane.
1. Decoupling WebSockets with Redis Pub/Sub
Never store active WebSocket connection handles inside application memory if you plan to run multiple server instances. If Customer A is connected to Server 1, and Agent B is connected to Server 2, direct memory messaging fails.
By introducing a Redis Pub/Sub backbone:
- Every room / conversation ID is a Redis channel.
- Any server instance can broadcast an event to the room without knowing which specific server holds the customer’s active socket.
- If a server crashes, the client reconnects to another instance behind the load balancer, resubscribes to the channel, and resumes immediately.
2. Idempotent Message Sequences
Mobile connections drop constantly (switching between Wi-Fi and LTE, entering tunnels, closing background tabs).
To guarantee delivery without duplicates:
- Every message payload contains a client-generated UUID (
idempotency_key) and an incremental integer sequence number. - When the socket reconnects, the client emits a
SYNCevent withlast_received_sequence_id. - The server responds with all messages newer than that sequence number from Redis cache, preventing gaps in the conversation log.
3. Graceful Human Handover Protocol
When transitioning a session from AI to human:
- The AI engine detects intent or triggers an escalation rule.
- An atomic Redis lock is acquired on the conversation state.
- Incoming customer messages are temporarily queued in Redis Streams.
- The dashboard alerts available operators; the first to accept claims the session lock.
- The queue drains to the operator console in exact order, and the customer receives an alert that a representative has joined.
Takeaway
Real-time infrastructure succeeds not when connections stay alive, but when your protocol gracefully handles connections dying and recovering without losing state.