Data & AI engineering consultancy · London

Data platforms and AI systems you can prove are right.

We design, build and migrate the data and machine-learning platforms that financial services and other regulated businesses run on — and hand them over with the evidence that every number reconciles.

Engineering delivered for
  • CIBC Capital Markets
  • Marshall Wace
  • Barclays
  • NHS Property Services
What we do

Three practices, one standard of engineering.

From the pipelines that land your data, to the models and agents that act on it, to the way your engineers build with AI — systems that hold up under audit, scale and change.

Data platform engineering

Modern data platforms on Snowflake and Databricks — designed for the volume, auditability and change control that regulated businesses need.

  • Greenfield platform builds and legacy platform migrations
  • Batch, streaming and event-driven ingestion
  • Analytics engineering and data modelling with dbt
  • Data quality frameworks and reconciliation
  • Self-service “paved road” provisioning: projects, identities and warehouse access on demand
SnowflakeDatabricksdbtSparkAirflowTerraformAWSAzureGCP

ML & AI engineering

Machine learning and AI that runs in production, not in a notebook — from model pipelines to agents that take real work off your teams’ plates.

  • ML platform migrations, e.g. Dataiku to Vertex AI, proven with parity reconciliation
  • Training, inference and monitoring pipelines, including GPU-accelerated training
  • Retrieval-augmented generation over company knowledge
  • AI agents that automate back-office and finance operations
  • LLM integration, evaluation and monitoring
Vertex AIKubeflowBigQueryDataikuLangChainClaudeGemini

AI-native engineering

Put coding agents to work inside your engineering team — grounded in your codebase, tickets and domain, with humans firmly in control.

  • Knowledge bases built automatically from code, Jira, Confluence, Slack and meeting notes
  • Ticket-to-pull-request agents that plan, implement, test and document changes
  • Domain-aware automated code review
  • Regression test generation and data-diff validation
  • Guardrails: scoped permissions, hooks and human approval gates
Claude CodeMCPLangGraphJiraConfluenceSlackGitHubGitLab
How we work

Built to be trusted, not just shipped.

Thirteen years of delivery in banks, hedge funds and the public sector taught us that the hard part of data work isn’t building the pipeline — it’s proving it’s right.

01

Proven, not assumed

Every migration ships with automated reconciliation against the system it replaces. Differences are traced to a named cause — never waved through as noise.

02

Senior engineers, hands on

The people who scope your project are the people who build it. No hand-off to a junior bench, no learning on your budget.

03

Faster and cheaper to run

Run time and compute cost are requirements, not afterthoughts. Rebuilt pipelines routinely run in a fraction of their legacy time.

04

Yours on day one

Tests, documentation and runbooks come as standard, so your team owns the platform the moment we step back.

Phase 1

Discover

Map the data, the systems and the decisions that depend on them. Agree what “correct” means before anything is built.

Phase 2

Architect

A target design sized to your team and budget — pragmatic choices over fashionable ones.

Phase 3

Build in slices

Working increments in production early, each with tests and data-quality assertions from the first commit.

Phase 4

Reconcile & hand over

Parity evidence against the legacy output, then documentation and a clean handover to your team.

Selected work

Where the numbers have to be right.

Projects where correctness wasn’t optional — what we built, and what it changed.

Asset management

A regulatory-grade performance reporting engine

Led the build of the return, risk and exposure engine behind monthly investor reporting — from greenfield to production — replacing a decade-old system that had survived several failed rewrites. Position data modelled with full derivative look-through across hundreds of billions of rows.

OutcomeMonthly numbers released without per-fund manual sign-off, backed by a reconciliation strategy owned by Risk and Fund Accounting.

Capital markets

Event-driven trade and risk data platform

Streaming ingestion on Databricks for high-volume trade and risk feeds — metadata-driven, schema-evolving and built to absorb thousands of files per trigger — feeding a dbt risk layer on a medallion architecture.

OutcomeDownstream validation rules encoded as automated tests, so outbound submissions passed receiving systems’ gates first time.

Machine learning

Production ML moved off a legacy platform

Migrated production training, scoring and monitoring pipelines from a legacy ML platform to Vertex AI, with the data layer rebuilt in dbt on Snowflake. Legacy model artefacts were frozen as versioned seeds so the new pipelines could prove parity in scoring-only mode before retraining.

OutcomePipelines running several times faster than the platform they replaced, with byte-identical reconciliation as the handover acceptance test.

Public sector

Finance reporting moved to the cloud

Migrated on-premise business data to Azure through fault-tolerant pipelines, and replaced a legacy finance reporting solution with a cloud data warehouse.

OutcomeLess manual reconciliation and more accurate monthly finance reporting.

Sectors we know well Capital marketsAsset managementBankingInsurance & retirementRetailPublic sector
In-house

We run on what we build.

The best test of an automation practice is its own back office. These systems run Nephorium day to day — the same patterns we build for clients, proven on ourselves first.

Nephorium operations

An AI-run back office

Our own finance operations run on agents. Supplier invoices, statements and statutory documents are picked up from email within a minute and filed; bank-feed transactions are categorised with the right VAT treatment and posted to the ledger; invoices are reconciled against payments. Deterministic code handles the routine, an agent with full business context handles the ambiguous, and only genuine judgement calls reach a person.

ClaudePythonPrefectPostgresQuickBooks API

ResultDay-to-day bookkeeping without manual data entry, with an audit trail behind every figure.

Infrastructure

Production on our own platform

Around forty services — pipelines, orchestration, databases and internal apps — run on our own hardware, deployed from git, with edge-triggered alerting, health checks that test real dependencies rather than liveness, and nightly verified backups.

DockerPrometheusGrafanaTailscaleCloudflare

ResultThe operational discipline we bring to client platforms, applied to our own.

Contact

Got a data problem worth solving?

Tell us about the platform, the migration or the process you want automated. We’ll come back within two working days with honest thoughts on whether — and how — we can help.

Or email us directly [email protected]

We only use your details to reply. Privacy notice.