Skip to content
Rubra Digital

Approach

Evaluation first, then everything else

Most AI projects fail in the same order for the same reasons. My process is arranged to hit those failure points as early and as cheaply as possible.

Short answer

A Rubra engagement runs in five phases: a free 30-minute scoping call, an optional three to five week discovery, a build that starts by constructing the labelled evaluation set, the production build itself, and a two-week handover. The distinguishing choice is building evaluation before retrieval, which turns every later decision into a measurement instead of an argument.

  1. Phase 0

    Free

    A 30-minute call

    You describe the problem. I tell you whether it is the kind of thing that works, roughly what it costs, and whether I am the right people. Sometimes the answer is that you do not need a consultancy, and I say so.

    • An honest read on feasibility
    • A rough cost band
    • A recommendation on next step
  2. Phase 1

    3–5 weeks

    Discovery

    Structured interviews, a hard look at your actual documents, and feasibility probes against the riskiest assumptions. Nothing gets built yet. This phase establishes what is worth building.

    • Ranked use case portfolio
    • Feasibility findings on real data
    • Cost and effort estimates
    • A written no list, with reasons
  3. Phase 2

    Weeks 1–2 of the build

    Evaluation first

    Before retrieval code is written, I build the labelled evaluation set with your subject-matter experts. It is unglamorous, and it is the highest-leverage work in the project, because every later decision stops being a matter of opinion.

    • 150–300 labelled questions
    • Retrieval and generation metrics defined
    • CI harness scaffolded
  4. Phase 3

    8–14 weeks

    Build

    Ingestion, retrieval tuning as measured experiments, the generation layer with citations and refusal behaviour, permissions enforced at retrieval, observability from the first commit. You see working software every week.

    • Production system in your infrastructure
    • Architecture decision records
    • Tracing, cost attribution and alerting
  5. Phase 4

    Final 2 weeks

    Handover

    Working sessions with your engineers, a runbook covering the failure modes I actually found, and a documented set of things I would do next. The measure of success is that you do not need me afterwards.

    • Runbook and failure-mode guide
    • Handover sessions
    • A prioritised next-steps list

Working with me

Can we skip discovery and go straight to a build?

Often, yes. If the use case is well defined, the data is understood and someone internally has already validated feasibility, discovery adds delay without adding information. I will tell you which situation you are in during the first call rather than selling you a phase you do not need.

What do you need from our team?

Realistically: a subject-matter expert for roughly two days a week during evaluation set construction, an engineer who knows your data platform, and someone empowered to make decisions. Access to source systems is the usual long pole, so I ask for those requests to start in week one.

How do you handle scope changes mid-engagement?

I re-quote in writing before doing the work, and I am explicit about what comes out if something goes in. The most common change is discovering that document quality is worse than the sample suggested, which is why I assess the worst documents rather than the best ones in week one.

What happens if the project is not working?

I tell you early, in writing, with the evidence. I have stopped engagements at the halfway point and refunded the remainder when the data turned out not to support the use case. It is a better outcome than delivering something that will not be used.

Get a straight answer on your AI roadmap

A 30-minute call with the engineer who would do the work, not a salesperson. You will get an honest read on what is worth building, what is not, and roughly what it costs.

No NDA needed to talk. EU and UK hours in full, with afternoons overlapping US Eastern and Central.