Inference Founder // Field Course
0% completeCourse brief

Lesson 06 / 07 · III · Company

Turn pilots into a company

Price the wedge, run design partnerships, hire the nucleus, and raise against evidence.

Founder decision
Raise to cross a measured technical or commercial gate, not to fund an undefined platform.
Target time
6 min briefing + 60–120 min fieldwork
Evidence
5 primary / first-party records
Before the answer: Write the one metric that would make a design partner expand from pilot to production.

NARRATED · CAPTIONED · PRECISE ENGINEER

Narrated visual briefing

Turn pilots into a company

01 · Paid learningA pilot is a contract to learn
02 · Pricing ladderAlign price with the control point
03 · Team nucleusFour owners, no orphaned surface
04 · Pilot exampleOne paid pilot can finance the roadmap
05 · Milestone financingRaise to retire the next named risk
Paid learning

Turn your first deployments into paid design partnerships. Define one workload, a frozen baseline, representative traffic, a four-to-eight-week window, weekly operator access, security boundaries, success metrics, a fee, and a production conversion decision. The customer receives priority and influence. You receive data, truth, and a reference if the result holds. Free open-ended pilots hide urgency, encourage custom work without commitment, and make it impossible to distinguish learning from unpaid support.

Pricing ladder

Choose pricing that mirrors what you control. Per-token pricing fits a shared API, but you absorb utilization risk. GPU-hour plus a platform fee fits dedicated deployments, but can look like commodity resale. Reserved commitments match predictable enterprise capacity. Outcome pricing works only when your wedge maps tightly to business value. Baseten combines usage pricing with enterprise deployment and support. Your early pilot fee should cover focused attention and test willingness to pay. It does not need to optimize the final margin model.

Team nucleus

Build a small nucleus around four ownership zones. The founder owns discovery, sales, priority, capital, and partnerships. An inference engineer owns profiling, runtimes, kernels, and quality-performance tradeoffs. A distributed-systems engineer owns the gateway, orchestration, observability, and incidents. A customer engineer owns integration, evaluation, traffic replay, and production handoff. These may initially be three people wearing four hats. Add functions only when repeated work, not imagined scale, demands them.

Pilot example

Consider an illustrative four-week design partnership priced at fifteen thousand dollars, not as a universal recommendation but as a structure to test. The partner supplies a frozen model, two traffic traces, an evaluation set, and a weekly operator. You commit to one primary result: cut p95 time to first token by forty percent at the target concurrency, while answer quality and error rate do not regress. Week one reproduces the baseline. Week two deploys the control loop. Week three runs shadow traffic and a game day. Week four produces the final benchmark, security packet, economics, and a production decision. The fee tests urgency and pays for focused learning. If the target holds, conversion includes a reserved-capacity commitment and platform fee. If one named gap remains, extend once with an explicit decision date. Otherwise stop. Three pilots with the same architecture and buying motion are evidence for repeatability. Three unrelated consulting projects are evidence that the wedge is still too broad.

Milestone financing

Tell the financing story as risk retirement. Show repeated customer incidents, committed design partners, a reproducible benchmark, unit economics by regime, and reliability evidence. Then name the next gate: perhaps three paid pilots converted to production, a measured cost advantage at target concurrency, or a pre-silicon architecture proof. Ask for enough runway to cross that gate with buffer. Do not finance an undefined horizontal platform, and do not present a GPU reseller as if capacity alone creates software margins.

WORKING MODEL

Four ideas to carry

Paid learning

A design partnership buys access, priority, and a defined result. Free pilots hide urgency and create orphan integrations.

Price the control point

Per token fits APIs; per GPU-hour fits dedicated capacity; platform or minimum commitments fund reliability and support.

The first team is a nucleus

Founder-led product, inference/runtime, distributed systems/SRE, and a customer-owning engineer. Add functions after repeatability.

Milestones finance risk retirement

Seed proves repeated pain and a working wedge. Series A should prove repeatable deployment and expansion, not promise them.

REFERENCE

The design-partner contract

Define one workload, a frozen baseline, representative traffic, a four-to-eight-week window, weekly operator access, security boundaries, success metrics, a fee, and a production conversion decision. The customer receives priority and influence; you receive data, truth, and a reference if successful.

Baseline: current p95 TTFT, quality, cost, and deployment time.

Target: one primary outcome plus non-regression constraints.

Inputs: traces, model artifacts, test corpus, owner, environment access.

Decision: expand, extend for one named gap, or stop.

Pricing ladders

ModelWhen it fitsRisk
Per token/requestShared API, elastic trafficYou absorb utilization risk
Per GPU-hour + platform feeDedicated custom deploymentCan look like low-margin resale
Reserved capacity commitmentPredictable enterprise loadCapacity forecasting error
Value or outcome componentWedge tied tightly to business outputMeasurement and procurement complexity

Baseten itself mixes usage pricing with enterprise deployment and support capabilities.E4, E7 Your pilot price should cover attention and test willingness to pay; it does not need to maximize early gross margin.

The first four roles

  1. CEO/founder-product: discovery, sales, priority, capital, partnerships.
  2. Inference/compiler engineer: model runtime, kernels, profiling, quality-performance tradeoffs.
  3. Distributed systems/SRE: gateway, orchestration, observability, incident behavior.
  4. Customer engineer: integration, evaluation, traffic replay, production handoff.

For a hardware path, add architecture, verification, physical design, compiler, board/system, and manufacturing leadership before tapeout planning. Taalas's long-collaborating team is a warning against assembling that capability casually.E3

Fundraising narrative

Show: repeated incidents; design partners; a reproducible benchmark; unit economics by regime; reliability evidence; and the next risk-retiring milestone. Ask for enough runway to cross that milestone with buffer. Do not pitch a GPU reseller as a software-margin company or a research project as a near-term infrastructure business.

SHIP EVIDENCE

Write the pilot and financing packet

Complete the design-partner one-pager, pricing model, 18-month hiring plan, use-of-funds table, and five milestone charts. Ask one target customer to redline the pilot offer.

Course artifact: labs/design-partner-one-pager.md

DIAGNOSTIC CHECKS

Can you use the model?

CHECK 1 · application

What makes a design partnership diagnostic?

CHECK 2 · recall

Which pricing model best matches dedicated predictable capacity?

CHECK 3 · transfer

What should a seed round primarily finance?

EVIDENCE LEDGER

Sources and limits

Vendor performance and product claims are labeled as first-party. Prices and product surfaces can change; follow the live links before making a purchasing or fundraising decision.

  1. ev-taalas-team The path to ubiquitous AI · Taalas
    Taalas says its first product took about two and a half years, a 24-person team, and 30 million dollars spent from more than 200 million dollars raised. It describes many team members as long-time collaborators and depends on experienced external partners.
  2. ev-baseten-product Baseten overview · Baseten
    Baseten presents a managed path from an open, fine-tuned, or custom model to a production API with containerization, GPU scheduling, autoscaling, observability, model-specific runtime optimization, and multi-cloud placement.
  3. ev-baseten-pricing Cloud pricing · Baseten
    Baseten lists pay-as-you-go dedicated GPU deployment prices by minute and offers scale-to-zero. Its July 2026 page lists H100 and B200 instances alongside smaller GPUs, while enterprise plans add self-hosting, custom SLAs, security controls, and use of existing cloud commitments. Prices are volatile; consult the live page.
  4. ev-yc-launch YC's essential startup advice · Y Combinator
    Y Combinator advises founders to launch a useful early product, talk directly to users, iterate from observed needs, keep the team small before product-market fit, and care about unit economics rather than scaling an unprofitable product.
  5. ev-blank-search Search versus execution · Steve Blank
    Steve Blank distinguishes a startup's search for a repeatable business model from the later execution of a known model. Customer development turns business-model assumptions into experiments that can be falsified and revised.