Case study 10 · Engineering Lifecycle Services

Live run monitoring and cancellation across server instances

Step-by-step run logs streamed to the browser as they are written, and cancellation that reaches whichever server is running the pipeline.

Partner
A cross-border consulting firm
Period
October 2024 to January 2026
Photo: Steven Wei on Unsplash
Live streams
2
Whole run, and each step's input or output
Instance
Any
Cancellation reaches the server running the work
On connect
Replay
Earlier logs sent before live ones

The challenge

Pipelines can run for minutes. Users needed to watch progress step by step and stop a run that was going wrong, even though the platform runs on several server instances.

What we built

A database trigger that announces each new log entry, a listener that forwards entries to browsers over WebSockets, and a Redis channel that carries cancellation requests to the instance that owns the run and reports success back.

What was delivered

  • Live run and per-step views over WebSockets
  • Earlier log entries replayed when a viewer connects
  • Cancellation across instances, recorded as a Cancelled status

Partner background

Our partner is a cross-border consulting firm whose teams work with large volumes of documents, spreadsheets, databases and email. It wanted one internal platform where those teams could build data pipelines and AI agents themselves, instead of commissioning a new application for each need. The platform had to run inside the partner's Microsoft 365 and Azure environment, keep each team's work separate, and move work from experiment to production through controlled environments.

The challenge

Waiting without feedback

Without live progress, users refreshed pages or started duplicate runs to check whether a pipeline was still working.

More than one server

The request to view or cancel a run can arrive at a different instance from the one executing it.

Large outputs

A step's output can be too large to send as a single message.

Objectives

  • Show each step's progress as it happens
  • Let viewers who join late see what already happened
  • Cancel a run from any session, whichever instance runs it
  • Deliver large step outputs safely to the browser

Our role

CharCentric provided technical leadership and architecture within a multidisciplinary engineering team, and contributed directly to implementation. The platform was built over 16 months, from October 2024 to January 2026, as a Python and FastAPI backend on Azure.

Scope and timeline

The log trigger, streaming endpoints and cancellation were built between December 2024 and August 2025.

Delivery timeline from the project history, against the overall platform build
Delivery timeline from the project history, against the overall platform build

Approach

Let the database announce changes

A PostgreSQL trigger on the run log table sends a NOTIFY message for every new entry. Any instance can listen, so logs do not depend on where the run executes.

Listen only when needed

The listener starts when the first viewer subscribes and stops when the last one leaves, so idle instances hold no open listeners.

A shared channel for control

Cancellation requests are published on a Redis channel. The instance holding the run cancels its task, marks the run Cancelled and publishes a confirmation.

Implementation

Architecture overview
Architecture overview

Run stream

A WebSocket endpoint per run first sends stored entries, then live entries, and closes when the run finishes.

Step stream

A second endpoint sends a chosen step's input or output, split into size-limited batches.

Cancellation

Each instance keeps a record of the runs it is executing and listens for cancellation only while it has active runs.

How it works, step by step
How it works, step by step

Tools and technologies

ToolPurpose
PostgreSQL LISTEN and NOTIFYLog change notifications
asyncpgNative listener connection
Redis publish and subscribeCancellation requests and confirmations
FastAPI WebSocketsBrowser streaming
AlembicTrigger migration

What was delivered

WebSocket streams
2
Shared cancel channel
1
Listeners
On demand
Large outputs
Batched
  • Live, replayable run and step logs in the browser
  • Cancellation that works across server instances
  • Listeners that run only while someone is watching
  • Large step outputs delivered in batches

Why it matters

Operating a system includes seeing what it is doing and stopping it when needed. Designing for several instances from the start avoids behaviour that only works on one server.

If your organization is planning a platform of this kind, or needs a specific part of one designed and delivered, we would be glad to discuss it.