Field guide / primary domain

Forward Deployed Engineering

9

sources in this field

Updated September 20, 2026

Current thesis

The shortest path to orientation.

This synthesis records claims and practices from the cited sources; reported outcomes and product capabilities have not been independently verified.

FDE anchors AI production value in a specialized integration layer bridging raw model capability to enterprise workflows through process reengineering, context aggregation, human-in-the-loop design, change management, domain-specific evals, and governance. Greater model capability amplifies rather than reduces this layer's importance—more powerful models enable more complex tasks, raising integration costs. Palantir's expand-stage accounts improved from -43% to +35% contribution margin, and scale-stage accounts reached 55% (top quartile 87%), demonstrating that year-one losses convert into durable relationships. Yet the UK Dept for Business and Trade's Microsoft 365 Copilot trial (1,000 licenses, 3 months) achieved only 1.14 actions/user/day despite 72% satisfaction, confirming that companies automate broken processes rather than redesign them. Success requires consolidating process definition before deployment via process mining and 10–20 domain expert interviews; skipping either is the biggest discovery failure mode. A critical risk: letting model providers route enterprise tokens creates a conflict of interest. Current production workaround deploys small LLM classifiers for routing, citations, tool use, and escalation—described as 'hacky' pending better calibration methods like RLCD. Application roadmaps for law firms target matter selection, associate staffing, and billing dispute prediction.

Evidence board

Claims worth carrying forward
01

LLM softmax probabilities are not necessarily calibrated confidence estimates — a known limitation of using LLMs as classifiers for routing, citation selection, tool use, and escalation decisions.

02

Jev (new frontier model from RLCD training, per @CompleteSkeptic) does not generate text — it produces categorical decisions only, so its 'hallucination-free' framing doesn't fully solve hallucination in systems that still need text generation elsewhere.

03

Current production workaround for calibrated categorical decisions: small LLM classifiers deployed for routing, citations, parts of a knowledge vault, tool use, and user escalation — described as 'hacky' pending better calibration methods like RLCD.

04

Better routing, citation selection, and escalation (via calibrated models like Jev) could reduce hallucinations across a broader system even though the core generation model isn't hallucination-free.

05

Longer-term application roadmap for law firms using calibrated categorical AI: matter selection, associate staffing decisions, and predicting billing disputes.

06

Open-source implementation of RLCD (the training method behind Jev) is anticipated so teams can post-train their own calibrated decision models rather than relying on a closed frontier model.

07

Applied AI value lives in a layer that connects raw model intelligence to enterprise workflows: reengineering processes, aggregating context/data, designing human-in-the-loop experiences, driving change management, running domain-specific evals, and managing security/governance.

08

As models get more capable, the applied/workflow layer becomes MORE important, not less—greater capability enables more complex tasks, which amplifies the cost of not building the integration layer well.

Adjacent fields

Key voices

Latest evidence

Recent additions

All synthesized insights →

Tanay Jaipuria

Ramp's Agentic Integration Factory

Ramp built an agentic system where customer requests for integrations are autonomously converted into working integrations, and that same demand feeds an internal 'integration factory' that productizes them as first-party features—shipping 75 integrations this way and cutting build time from weeks/months to hours.

ethan ding 📊

Why You Can't Copy Palantir — Deploy Co Tier List Analysis

Ethan Ding's essay argues Palantir is a vertically-integrated outcome-based software provider whose moat comes from three compounding inputs: 12 years of patient capital during a commercial PMF desert, government deployments lacking modern cloud primitives, and an outcome-based (not usage or hours-based) pricing model. This combination produced a differentiated FDE talent pool, an unusually broad product (Foundry/Ontology), and account-level economics competitors structurally cannot replicate, making Palantir near-impossible to clone in the current AI/VC funding environment that demands 3-month results.

Aaron Levie

Forward Deployed Engineers: Discovery Tool, Not Business Model

Aaron Levie and Jesse Zhang analyze why FDEs are resurging in AI go-to-market: AI agents require discovering entirely new, non-deterministic workflows that neither vendor nor customer has ever seen, unlike deterministic traditional software. The Palantir playbook shows FDE pain should convert into product primitives, but many AI startups risk turning FDEs into a permanent services crutch rather than a temporary discovery phase.

Lunar

Agent Harness vs Loop vs Graph Engineering — A Practical Stack

A framework distinguishing three separate layers in production agent systems: harness engineering (environment/tools/state), loop engineering (feedback/retry cycles), and graph engineering (explicit workflow topology). Includes diagnostic rules, common mistakes, and a production checklist, plus a reference to OpenAI's Forward Deployed Engineer (FDE) talk on production AI deployment.