Skip to content
Rubra Digital

AI & RAG Consulting for Financial Services

LLM and retrieval systems for banks, insurers and asset managers, built to satisfy DORA, model risk governance and the evidence your regulators expect.

Frameworks I design against

  • EU AI Act
  • DORA
  • GDPR
  • MiFID II
  • BaFin BAIT / MaRisk
  • FCA Consumer Duty
  • OSFI E-23
  • SEC / FINRA guidance

Short answer

Rubra builds LLM and RAG systems for banks, insurers and asset managers across Europe and North America, designed from the start to meet model risk governance, DORA operational resilience, and supervisory expectations on explainability and audit trails.

Last reviewed

Where I see this working

Regulatory change monitoring

Retrieval across consultation papers, final rules and internal policy, surfacing exactly which internal documents a new requirement touches, with citations a compliance officer can verify.

Credit memo and underwriting support

Drafting assistance grounded in financial statements, covenant documents and prior decisions. Almost always an EU AI Act high-risk use case, so it is designed with the human decision point and audit trail built in.

KYC and onboarding document review

Extraction and cross-checking across identity documents, corporate structures and adverse media, with confidence scoring and mandatory human review on anything uncertain.

Client-facing advisory support

Answers grounded in current product terms rather than a model’s recollection of them. That is the difference between a helpful assistant and a mis-selling incident.

Internal policy and procedure assistant

The lowest-risk, fastest-payback use case in most banks: staff finding the right procedure in seconds instead of asking a colleague.

Financial services was the first sector to take LLM governance seriously, largely because it already had the machinery. Model risk management, validation cycles, three lines of defence. The frameworks exist. What they were not built for is a model you did not train, whose behaviour can change when a vendor ships an update.

What changes with foundation models

You do not own the model. Traditional validation assumes you can inspect training data and reproduce results. With a hosted foundation model you can do neither. Validation shifts to behavioural evidence: a fixed evaluation set, run continuously, with declared thresholds and alerting when performance moves.

The failure mode is fluent and wrong. A traditional model fails visibly, producing a number outside plausible bounds. An LLM fails by producing an articulate, well-formatted answer that is incorrect. Controls have to assume plausibility is not evidence.

Behaviour drifts without a release. Your provider can change model behaviour without your deployment changing at all. Continuous evaluation is not a nicety here; it is the only way you would find out.

How I work in this sector

I build the evaluation harness before the application, because it is the evidence your second line will ask for and the thing that lets you keep shipping afterwards. I keep a human decision point on anything that affects a customer outcome. I instrument retrieval logging so a supervisor can be shown exactly which documents informed a given output. And I design for provider substitutability, because a single hard-coded endpoint is a DORA problem as much as an engineering one.

Most banking engagements start with the internal policy assistant: low risk, fast payback, and it builds the governance muscle on a use case where a mistake is inexpensive. The high-risk use cases come second, once the organisation has learned what evidence it actually needs to produce.

Frequently asked questions

How do you handle model risk governance for LLM systems?

I produce the artefacts your model risk function already expects, mapped to LLM specifics: documented purpose and limitations, data lineage, performance metrics with declared thresholds, ongoing monitoring, and a defined validation cadence. The main difference from traditional model validation is that the performance evidence comes from an evaluation harness rather than a backtest, so I build that harness as a first-class deliverable rather than an afterthought.

Can these systems run entirely within our own infrastructure?

Yes. I deploy into your own cloud tenancy, and for institutions with strict egress rules I design around in-region or in-VPC inference. For Swiss private banking and some EU public-sector-adjacent clients, that has meant fully in-country deployment with no data leaving the jurisdiction at any point.

Does DORA apply to our AI vendors?

If the AI system supports a critical or important function, then yes. The ICT third-party risk provisions apply to the model provider as they do to any other ICT service provider. Practically this means contractual requirements on subcontracting, exit strategies, and register entries. I design for provider substitutability from the start, which is what makes a credible exit plan possible rather than theoretical.

Get a straight answer on your AI roadmap

A 30-minute call with the engineer who would do the work, not a salesperson. You will get an honest read on what is worth building, what is not, and roughly what it costs.

No NDA needed to talk. EU and UK hours in full, with afternoons overlapping US Eastern and Central.