Case study 06 · AI Transformation

Applying LLMs row by row to business data, with structured and predictable output

A pipeline step that runs a chosen language model over every row of a table and returns typed fields that match a schema the user defines.

Partner
An international professional services firm
Period
October 2024 to January 2026
Photo: Ian Simmonds on Unsplash
Model providers
5
OpenAI, Anthropic, Google, Groq and xAI
Output fields
Typed
Validated against a user-defined schema
Retries
Per row
Failures handled without losing the batch

The challenge

Teams wanted to classify, extract and summarize information across thousands of records. Free-text model answers could not be fed reliably into reports, databases or other systems.

What we built

An analysis step that sends each row to a selected model with a templated prompt and requires a structured answer. Rows are processed in rate-limited batches, failures are retried per row, and generated fields are added to the original data as typed columns.

What was delivered

  • Structured output defined by the user and checked before a run starts
  • Batching and rate limiting tied to a single batch size setting
  • A choice between stopping on failure or continuing with default values

Partner background

Our partner is an international professional services firm whose teams work with large volumes of documents, spreadsheets, databases and email. It wanted one internal platform where those teams could build data pipelines and AI agents themselves, instead of commissioning a new application for each need. The platform had to run inside the partner's Microsoft 365 and Azure environment, keep each team's work separate, and move work from experiment to production through controlled environments.

The challenge

Free text does not fit a table

Model responses vary in shape and wording. Without a fixed structure, results need manual cleanup before anyone can sort, filter or load them.

Volume meets rate limits

Running a model over a large table means many calls in a short time. Providers enforce rate limits, and individual calls fail for transient reasons.

Mistakes surface late

A prompt that references a missing field, or output fields that clash with existing columns, would otherwise fail halfway through a long and costly run.

Objectives

  • Return model results as typed fields that downstream steps can rely on
  • Process large tables within provider rate limits
  • Contain failures to the rows they affect
  • Catch configuration errors before any model call is made
  • Allow the model and provider to be chosen per step

Our role

CharCentric provided technical leadership and architecture within a multidisciplinary engineering team, and contributed directly to implementation. The platform was built over 16 months, from October 2024 to January 2026, as a Python and FastAPI backend on Azure.

Scope and timeline

The analysis step was developed between November 2024 and October 2025 as part of the platform's pipeline library, alongside its configuration model, prompt validation and logging.

Delivery timeline from the project history, against the overall platform build
Delivery timeline from the project history, against the overall platform build

Approach

Schema first

The user defines the output as a JSON Schema. The step converts it into a typed model and asks the model provider for structured output, so every answer either matches the schema or is treated as a failure.

Validate before spending

Before the first call, the step checks that prompt variables exist in the data, that templates are valid, and that generated field names do not collide with source columns.

One setting for pace

Batch size controls both concurrency and the request rate through a token bucket limiter. Teams tune one number instead of several interacting settings.

Implementation

Architecture overview
Architecture overview

Prompt templates

Prompts combine fixed values with fields from each row, including nested fields. The same template runs across the whole table, which keeps results comparable.

Batch processing

Rows are processed in asynchronous batches with a pause between batches. Each row is retried a configurable number of times, and every retry is logged.

Failure policy

With tolerate failure on, a row that still fails receives default values and the run continues, with the count of failed rows logged. With it off, the run stops with a clear error.

Output

Generated fields are merged with the original row into a new typed table that writers, joins and other steps can use directly.

How it works, step by step
How it works, step by step

Tools and technologies

ToolPurpose
LangChainModel initialization and structured output
PydanticTyped models generated from JSON Schema
OpenAI, Anthropic, Google, Groq, xAISelectable model providers
asyncioConcurrent batch processing

What was delivered

Providers
5
Output contract
JSON Schema
Validation
Pre-run
Retry and fallback
Per row
  • A reusable analysis step available to every pipeline on the platform
  • Typed, schema-checked output instead of free text
  • Rate-limited batch execution with per-row retries
  • Configuration errors caught before any model cost is incurred

Why it matters

Language models become useful in operations when their output can be trusted by the next system in the chain. Treating structure, pacing and failure as design decisions turns a model call into a dependable data step.

If your organization is planning a platform of this kind, or needs a specific part of one designed and delivered, we would be glad to discuss it.