
Building a test suite that makes releases predictable for a growing AI and data platform
Unit, integration and end-to-end tests for a Python backend with no existing test suite, covering its databases, orchestration, storage and language model integrations.
- Partner
- A growing AI and data platform provider
- Period
- 3 months
- Test layer
- Unit
- API endpoints, database operations and core utilities
- Test layer
- Integration
- Authorization, Prefect, language models and Azure Blob Storage
- Test layer
- End to end
- Complete workflows from input to stored and displayed result
The challenge
The partner's Python backend connected to PostgreSQL, ClickHouse, Prefect, Azure Blob Storage, internal authorization services and language models. No test suite existed, so every release relied on manual checks and confidence fell as the platform grew.
What we built
A Pytest framework with three layers of tests: unit tests for the backend's own logic, integration tests for each external dependency, and end-to-end tests that run real workflows, including the platform's AI content generation and meeting summarization features.
What was delivered
- Automated tests across three layers
- Every external dependency covered by integration tests
- Key AI features verified as complete workflows
Partner background
Our partner is a growing AI and data platform provider that delivers backend systems to enterprises. As the platform gained features and integrations, its backend came to depend on several databases, a workflow orchestrator, cloud storage, internal security services and language models. The partner wanted to release frequently without losing confidence in what it was shipping.
The challenge
No safety net
The backend handled API requests, database work and external integrations, but no automated tests existed. Every change carried unknown risk.
Many external dependencies
Authorization services, Prefect, PostgreSQL, ClickHouse, Azure Blob Storage and language models each had to be exercised, because most failures appear at those boundaries.
Workflows, not just functions
The platform's value lies in multi-step workflows. Testing individual functions would not show whether a full workflow produced the right result.
A growing surface
New features and integrations were arriving regularly, so the test approach had to absorb them without being redesigned.
Objectives
- Automated tests at every layer of the backend and its dependencies
- Stable service behaviour under the conditions the platform meets in use
- Release confidence backed by automated verification
- A structure that new features and integrations can extend
Our role
CharCentric designed the test framework and wrote the unit, integration and end-to-end test suites over three months.
Scope and timeline
The work covered three layers: unit tests for individual services and modules, integration tests for interactions with external services, and end-to-end tests that simulate user journeys through the whole system. It was delivered in three months.

Approach
One framework, three layers
Pytest was chosen because it is widely used in the Python ecosystem and supports fixtures, parametrization and plugins for all three layers. The same conventions apply throughout, so contributors move between layers easily.
Test the boundaries
Integration tests run against each external dependency, because that is where data formats, permissions and timing most often break.
Follow the user
End-to-end tests follow complete journeys, from orchestration through storage to the result a user sees, including features that rely on language models.
Implementation

Unit tests
Unit tests cover API endpoints, database operations against PostgreSQL and ClickHouse, and the utilities that process, transform and validate data.
Integration tests
Integration tests cover communication with internal authorization and authentication services, Prefect task orchestration, language model calls and Azure Blob Storage uploads, retrieval and processing.
End-to-end tests
End-to-end tests run complete data flows from Prefect orchestration to storage in PostgreSQL and ClickHouse, the AI content generation feature, and the meeting summarization feature from recorded input to delivered summary.
Tools and technologies
| Tool | Purpose |
|---|---|
| Pytest | Test framework for all three layers |
| PostgreSQL and ClickHouse | Databases under test |
| Prefect | Workflow orchestration under test |
| Azure Blob Storage | File storage under test |
| Language model APIs | AI features under test |
| Internal auth services | Security integrations under test |
What was delivered
- Framework
- Pytest
- Unit, integration, end to end
- 3 layers
- Dependency tested
- Every
- Delivery
- 3 months
- A Pytest framework organized in three layers
- Integration coverage for every external dependency named in scope
- End-to-end coverage of the main data flows and two AI features
- A structure that new features can extend with the same conventions
Why it matters
Test counts describe effort. What matters is that each release can be checked automatically against the behaviour that partners depend on, which is what lets a growing platform keep releasing often.
If your organization is facing a similar challenge, we would be glad to discuss it.