The engineers who stay until it ships
We put senior engineers inside your organisation — your repository, your standups, your constraints — and take one real workflow from prototype to production. You keep the code, the evaluation suite and the runbook.
- Engagement
- Six-week production sprint
- Fee
- Fixed, agreed before week one
- Code ownership
- Yours, from the first commit
- First pull request
- Within seven days
- Starting at
- $12,000
- Palantir — the original forward deployed engineer, from 2003
- OpenAI — Forward Deployed Engineering
- Anthropic — Applied AI
- Scale AI — forward deployed and frontier agent engineering
- Cohere — forward deployed engineers on private deployments
- Sierra — the same role, named agent engineering
The model was never the hard part
Capability stopped being the constraint some time ago. Delivery did not. Every figure below is somebody else's research rather than ours, and they all measure the same gap — the distance between a system that works in a demonstration and one a business can actually run on.
of enterprise generative-AI pilots produce no measurable return
of companies cannot scale value from AI beyond a first success
of AI proofs of concept never reach production at all
of AI projects never move from pilot into production
What a forward deployed engineer actually is
Not a consultant with a laptop, and not a contractor on your headcount. A senior engineer who works inside your organisation on one named outcome, with commit rights, and who is accountable for that outcome reaching production rather than for the hours spent approaching it.
One outcome, scoped in writing
We agree a single production outcome and the criteria that decide whether it was met, and both sides sign them before any code is written. What is out of scope is written down too, which is the half that usually goes unrecorded.
Commit rights in week one
Repository, CI, deployment pipeline, Slack, standups, on-call rota. Our engineers appear on your board like anyone else on it, and their work is reviewed like anyone else's. The first pull request lands inside seven days.
Built against an evaluation set, not a demonstration
A golden set of real cases with known answers, assembled with your domain experts in the first fortnight. Every candidate is scored against it before it goes anywhere near a user, and the score is the argument for shipping.
Production, or it did not happen
Real traffic, real volumes, real failure modes, with rollback armed and an audit trail live. A prototype that never carried load taught us something; it did not deliver anything, and we do not invoice it as though it had.
You are already holding three other options
A consultancy, a staffing firm, and the solutions engineer your software vendor sends before the contract is signed. Each is the right answer to some question. Here is which questions, scored honestly enough to be worth reading.
| Criterion | Forward deployed engineer | Management consultancy | Staff augmentation | Vendor solutions engineer |
|---|---|---|---|---|
| Primary outputProduction code in your repository, running on real traffic | Fully | Not what it is for | Partly | Not what it is for |
| Accountable for the outcomeAgainst written acceptance criteria, not against hours billed | Fully | Partly | Not what it is for | Not what it is for |
| Works inside your systemsYour repository, your pipeline, your data, under your access controls | Fully | Not what it is for | Fully | Partly |
| Time to a merged pull requestSeven days from access being granted | Fully | Not what it is for | Partly | Not what it is for |
| Domain fluency in the actual workflowSits with the people doing the work, not the people describing it | Fully | Partly | Not what it is for | Partly |
| Still there after go-liveThrough handover and on call for it, then deliberately less | Fully | Not what it is for | Partly | Partly |
| Leaves your team able to run itRunbook, evaluation suite and a dry run your engineers complete unaided | Fully | Partly | Not what it is for | Not what it is for |
| Priced before you commitFixed fee against a milestone rubric, quoted on the first call | Fully | Not what it is for | Partly | Not what it is for |
Fully
Partly
Not what it is for
Six weeks, and what exists at the end of each one
A process whose steps are called discover, design and deliver commits nobody to anything. Every step below names the artefact that exists when it closes, because that is the only part of a plan you can hold anyone to.
Scope
We pick one outcome with the people who own it and the people who do it — usually not the same people — and write down what would have to be true for it to count as delivered. Access provisioning starts here, because it is what actually causes slippage.
A signed milestone rubric, and provisioned access
Embed
Repository, CI, deployment pipeline, Slack, standups. Our engineer joins your rituals rather than running a parallel set, and is reviewed by your reviewers. Nothing about the working arrangement is bespoke.
A merged pull request, inside seven days
Prototype
Something that runs on your stack against your data, behind a feature flag, in front of the people whose work it changes. In parallel, a golden set of real cases with known answers, built with your domain experts.
A working prototype and a scored golden set
Build
Integration into the systems of record, the evaluation harness, the human review queue, the audit trail, and whatever your risk and compliance functions need in order to sign. This is the part that separates a demonstration from a system, and it is most of the work.
A production candidate scored against the incumbent
Cut over
Real traffic, real volumes, real failure modes, with rollback armed and monitoring in place. We are on call for it. Nothing here is thrown over a wall on a Friday.
The system live, and its service levels met
Transfer
Your engineers run a release, a retrain and a rollback with us in the room but not on the keyboard. If that dry run does not go cleanly, the engagement is not finished, whatever the calendar says.
Runbook, evaluation suite and pipeline in your repository — and a dry run your team completed unaided
What you keep when we leave
The measure of this work is not what we built. It is what your team can change six months later without calling us. Everything below exists so that the answer to that is all of it.
Your repository, from the first commit
Code lives in your repository from day one, not in ours pending a transfer at the end. Intellectual property is assigned in writing before week one starts. There is no platform to license, no runtime to rent, and nothing to migrate off later.
The evaluation suite ships with the system
The golden set, the scoring harness and the regression run go into your pipeline and fail your build when quality drops. It is the artefact that lets someone who was not there change the system safely, which makes it the one that matters most.
A runbook your team can actually run
Deploy, rotate credentials, retrain, roll back, escalate. Written as procedure rather than architecture, and validated by your engineers executing it while we watch — not by us executing it while they watch.
Dependency that goes down, not up
A good engagement ends with you needing us less. We hand work back on a deliberate schedule and say so when the remaining scope no longer justifies an outside engineer, because a retainer that outlives its reason is how this model turns into the thing it was meant to replace.
Senior by default, and named before you sign
In this model the engagement is a person rather than a methodology, so the bar is the product. You meet the engineer who will do the work during scoping — not a partner who introduces someone else once the contract is signed.
Engineers who have shipped and then run it
- Eight years or more writing production software.
- Has carried a system through an incident, not only through a launch.
- No junior rotation, and nobody learning on your budget.
Fluent in the room, not just the repository
- Runs discovery with the people doing the work.
- Can hold an architecture review and an executive briefing in the same day.
- Says no to scope that will not survive the handover.
Production AI rather than demonstration AI
- Has taken language-model systems past a proof of concept.
- Writes the evaluation harness before tuning the prompt.
- Treats non-determinism as an engineering problem with known controls.
Named, and yours for the engagement
- You meet them during scoping, before anything is signed.
- One engagement at a time, not four in parallel.
- The same person through build, cut-over and handover.
We meet your stack rather than importing ours
Nothing here is a partnership badge or a preference. It is the set of tools our engineers have taken into production and can be useful in from the first week, which is the only claim a list like this can honestly make.
- Claude
- GPT
- Gemini
- Llama
- Mistral
- Qwen
- Amazon Bedrock
- Vertex AI
- Azure OpenAI
- MCP
- LangGraph
- Temporal
- Pydantic AI
- DSPy
- Celery
- Step Functions
- Postgres
- pgvector
- Snowflake
- Databricks
- Elasticsearch
- Kafka
- dbt
- Airflow
- TypeScript
- Python
- Go
- React
- Next.js
- FastAPI
- GraphQL
- AWS
- Azure
- Google Cloud
- Kubernetes
- Terraform
- GitHub Actions
- On-premises
- Golden sets
- LLM-as-judge
- OpenTelemetry
- Langfuse
- Datadog
- Grafana
- Sentry
- Salesforce
- SAP
- Dynamics
- ServiceNow
- Workday
- Epic
- Guidewire
- NetSuite
- Okta
- Entra ID
- SAML
- SCIM
- Vault
- Row-level security
- Audit logging
- Your VPC
- Your tenancy
- On premises
- Air-gapped
- Sovereign region
Depth in a few places rather than presence in all of them
Every sector below is one where the hard part is not the model. The constraint named under each is the thing that actually decides whether a system reaches production, and it is what an engagement in that sector is mostly spent on.
Financial services
Credit operations, onboarding and KYC review, client reporting, and the middle-office work that sits between a system of record and a spreadsheet nobody will admit to.
SOC 2 and SOX, plus a model risk function that has to sign every decision the system makes on its own.
Healthcare and life sciences
Referral and intake triage, prior authorisation, clinical documentation, and the coordination work that consumes a clinician's day without appearing in any record of it.
Protected health information under HIPAA, and a clinician who will not tolerate a second screen to check the first one.
Insurance
First notice of loss, claims triage, subrogation review, and underwriting support where the file is longer than anyone has time to read.
Regulated decisioning: every declination needs a reason a person can read and a regulator can follow.
Legal and professional services
Contract review, diligence, disclosure, and the research work that is currently absorbed by whoever is most junior in the room.
Privilege boundaries that cannot leak between matters, and citations that survive being checked.
Manufacturing and logistics
Exception handling across consignments and orders, supplier correspondence, quality documentation, and the reconciliation between two systems that were never meant to speak.
Systems of record older than the team maintaining them, and a line that does not stop for a deployment.
Public sector
Case intake and routing, correspondence handling, eligibility review, and the backlog work that grows faster than headcount can be approved.
Procurement timelines, data residency, and an audit trail that has to outlive the contract that produced it.
The same numbers we would give you on a call
Almost nobody in this category publishes a price, which asks a reader to spend a meeting finding out whether they can afford the first conversation. These are the bands. If you are not in them, you will know before you have booked anything.
Readiness sprint
$12,000 – $20,000
For teams holding a shortlist of candidate workflows and no agreement on which one to build first.
- Workflow, data and access assessment
- One outcome selected and scoped
- A milestone rubric you can hand to anyone
- A fixed price for building it, or a recommendation not to
Production sprint
$45,000 – $85,000
The core engagement. One workflow taken from prototype to production inside your own systems.
- One senior engineer embedded full time
- A merged pull request within seven days
- Golden set and evaluation harness
- Live on real traffic, with rollback armed
- Runbook, handover and a dry run your team completes
Regulated deployment
$95,000 – $180,000
Where the constraint is the regime rather than the model — protected data, model risk, privilege, residency.
- Everything in the production sprint
- Compliance architecture document
- Security review package and threat model
- Audit trail and human review queue
- Deployment inside your VPC, tenancy or data centre
Embedded retainer
$18,000 – $40,000 per month
After the first system is live, for teams building the second and third before they have hired for it.
- Continuing capacity against a named backlog
- Quarterly commitment, monthly notice
- Shared on-call for what we built
- Scaled down as your own team scales up
These are bands, not quotes — scope moves them. We will tell you where in the band you land on the first call rather than in a proposal three weeks later. Every engagement is a fixed fee against written acceptance criteria: we do not bill hourly, we do not run open-ended retainers, and we do not invoice for a milestone that was not met.
When this is the right call, and when it is not
This model is being sold at the moment as though it suited everything. It does not, and the cases below are the ones where we would rather say so now than three weeks into an engagement.
This works when
- You have one workflow that matters and a date attached to it
- You can grant real access to real systems and real data inside a fortnight
- Someone senior owns the outcome and can clear a blocker in a day
- Your constraint is delivery capacity rather than knowing what to build
- You want your own team running the system afterwards, and will make room for that
- You would rather have one workflow in production than six in pilot
This is the wrong call when
- What you actually need is a strategy document, a vendor evaluation or a board paper
- Access to the systems and the data will take a quarter to arrange
- The problem is a genuine research problem and no model does it yet
- You want someone to own the software indefinitely — that is a managed service, and we will point you at one
- Nobody internally has the capacity to take the handover
- The outcome cannot be written down in a way both sides would sign
Three engagements, described as far as we are permitted to
A health system's referral intake stopped being read by hand
Inbound referrals arrived as scanned faxes, portal submissions and email attachments, and were triaged by a team reading each one. The work was extraction, routing and a review queue for anything the system was not confident about — built inside their environment, against a golden set their own coordinators labelled.
- Six weeks from kickoff to live on real referral volume
- 94% routed without a human read, measured on a 480-case golden set
- Runbook dry run completed by their team, unaided, on day 41
A specialty insurer made its declination letters defensible
Claims decisions were consistent but the reasoning behind them lived in adjusters' heads, which made regulatory review expensive. We built decision support that cites the policy clause and the evidence for every recommendation, with an adjuster approving each one before it leaves.
- First merged pull request on day five
- Every decision carries a cited reason a regulator can follow
- Ran entirely inside their VPC; no policyholder data left it
A logistics operator replaced a spreadsheet nobody would admit to
Exceptions between a transport management system and an order system were reconciled manually by three people, using a spreadsheet that had no owner and no backup. The engagement ran long because the integration had never been documented, and finding out why took most of a fortnight.
- Eleven weeks, two systems of record, one undocumented integration
- Exceptions requiring manual reconciliation down from 40% to 6%
- Their engineers shipped the next two releases without us
FAQs
The deliverable and the incentive. A consultancy is accountable for a recommendation; a forward deployed engineer has commit rights on your stack and is accountable for a system reaching production. It is a different output, a different fee structure, and a different hiring bar — we are hiring people who ship software, not people who can describe shipping it.
A contractor takes your ticket queue and works it. We take one outcome and are measured on whether it reaches production, which means scoping it, arguing about it, and occasionally telling you the ticket you wrote is not the thing that will move the number. If what you need is capacity against an existing backlog, staff augmentation is cheaper and you should use it.
You do, entirely. Code lands in your repository from the first commit rather than in ours pending a transfer at the end, and assignment is signed before week one starts. There is no platform to license, no runtime to rent, and nothing to migrate off if you never speak to us again.
Two to four weeks from a signed scope, and the gating item is almost never us. It is access — repository permissions, a staging environment, and whatever your security review requires before an outside engineer can see production data. We start that conversation in week zero for exactly this reason.
Both, and the split is decided by where the domain knowledge is rather than by preference. Discovery is on site, because the useful information is what people do rather than what the process document says they do. Build is mostly remote, in your tooling, at your standups. Cut-over is on site again.
The two-week readiness sprint. It is deliberately small enough to be a decision rather than a commitment, and it ends in a scoped outcome and a fixed price — including, sometimes, a recommendation that you do not need us for it.
Yours. We have no platform to sell and no reference architecture to defend, so the question is only ever what your team can maintain after we leave. Where we do have a strong opinion — evaluation harnesses, audit trails, human review queues — it is about practice rather than product, and it applies whatever you are running.
Yes, and it is most of what we do. Deployments run inside your VPC, your tenancy, your data centre or an air-gapped environment. The regulated engagement carries a compliance architecture document, a threat model and a security review package, because in those environments the review is the timeline.
No — it is the normal starting position, and the readiness sprint exists for it. What would be a problem is not being able to grant access to the systems and the people, because then nobody can tell which candidate workflow is worth building and the choice ends up being made from a slide.
Your engineers run a release, a retrain and a rollback with us in the room but not on the keyboard, and the engagement is not finished until that goes cleanly. After that, some teams take a retainer for the next system and some do not; we will say plainly which we think applies, including when the answer is that you no longer need us.
Pick one workflow. Six weeks. Then decide.
Thirty minutes is enough to know whether this is a fit. No pitch and no deck — we will map one workflow with you, give you the band it falls in, and tell you plainly if we are the wrong people for it.











