
Running user code and parsing documents in isolated services
Two separate services that let pipelines run user-written Python and turn documents into structured text, without putting the main platform at risk.
- Partner
- An international consultancy
- Period
- October 2024 to January 2026
- Memory per run
- 1 GB
- Runs above the limit are stopped
- Isolated services
- 2
- Code execution and document parsing
- Results
- Streamed
- Parsed as they are produced, returned in batches
The challenge
Users needed to write custom Python transformations and bring PDFs and office documents into pipelines and knowledge bases. Running either inside the main backend would let one heavy job slow or stop the platform for everyone.
What we built
A code execution service that runs each job in its own process under a memory guard, with concurrency set by available memory, and a Docling based parsing service that accepts files or URLs and returns structured text. Both communicate with the platform over WebSockets.
What was delivered
- User code isolated from the platform in separate processes
- Concurrency calculated from available memory at 1 GB per run
- Document parsing for pipelines and knowledge sources
Partner background
Our partner is an international consultancy whose teams work with large volumes of documents, spreadsheets, databases and email. It wanted one internal platform where those teams could build data pipelines and AI agents themselves, instead of commissioning a new application for each need. The platform had to run inside the partner's Microsoft 365 and Azure environment, keep each team's work separate, and move work from experiment to production through controlled environments.
The challenge
Custom logic is unavoidable
No library of predefined steps covers every transformation. Users need to write code, and that code can be slow or memory hungry.
Shared resources
A runaway job inside the main backend would affect every user of the platform.
Documents are not data
PDFs and office files must be converted to structured text before a pipeline or knowledge base can use them.
Objectives
- Run user code outside the main platform process
- Limit the memory any single run can use
- Admit only as many runs as the service can hold
- Convert documents to structured text for pipelines and agents
Our role
CharCentric provided technical leadership and architecture within a multidisciplinary engineering team, and contributed directly to implementation. The platform was built over 16 months, from October 2024 to January 2026, as a Python and FastAPI backend on Azure.
Scope and timeline
The code execution service was built between March and July 2025, and the document parsing service between April and October 2025.

Approach
Separate services, separate failures
Both capabilities run as their own containers. A failure or overload in one does not reach the platform API.
Admission by memory
The number of concurrent runs is calculated from available memory at 1 GB per run. Further requests wait for a free slot instead of overloading the host.
Stream instead of buffer
Inputs and outputs move in batches over WebSockets, and results are parsed with a streaming JSON parser, so large datasets do not have to fit in memory at once.
Implementation

Code execution
Code and input arrive in batches and are written to a working file. The run starts in its own process while a monitor checks memory use and stops it above 1 GB. Results stream back to the pipeline in batches.
Document parsing
The service accepts file content or a list of URLs, validates URL formats, keeps a cache of parsed documents and uses Docling to produce structured text.
Platform use
The code step uses the execution service. The document reader step and knowledge loaders use the parsing service.

Tools and technologies
| Tool | Purpose |
|---|---|
| FastAPI WebSockets | Service interfaces |
| Python subprocesses | Run isolation |
| psutil | Memory monitoring |
| ijson | Streaming JSON parsing |
| Docling | Document parsing |
| Docker | Service containers |
What was delivered
- Services
- 2
- Memory limit per run
- 1 GB
- Input and output
- Batched
- Parsed documents
- Cached
- An isolated execution service for user-written Python
- Memory limits and memory-based admission control
- A document parsing service for pipelines and knowledge bases
- Streaming, batched data transfer between services
Why it matters
Flexible platforms let users bring their own logic and files. Isolating that work, and limiting what it can consume, keeps flexibility from becoming a reliability risk.
If your organization is planning a platform of this kind, or needs a specific part of one designed and delivered, we would be glad to discuss it.