Dacard.ai · a working synthesis
Building software got cheap, so the scarce work moved. This is the whole model in one place: how the work runs, how you build it, how you read whether it landed, and how you tell where you are.
Every claim in here carries one of those three. A team that cannot tell a measured result from a projected one is steering on hope, however confident the dashboard looks. This document is an illustrative model, not a report on any one company. What is measured, and where it came from, is listed at the end.
00 · The through-line
The expensive part of software was writing it. Every estimate, plan, and handoff existed to manage that one cost.
Take that cost close to zero and the machine does not get lighter. It points at the wrong thing. A model built to ration capacity now optimizes for a constraint that moved.
When an agent can build almost anything, the edge is no longer building. It is choosing what to build, knowing whether it landed, and steering the next investment from what you learn.
Strategy sets direction. Operations reads the result. At this speed they are one motion, not two functions that meet quarterly.
01 · The approach
Five moves. People decide and judge, the agent builds, and two things run underneath every move rather than between them: what it costs, and how it gets checked.
A person decides and judges. The agent builds. Ship is the one step that runs on its own, and only behind a flag it can be pulled back through.
| Move | What happens |
|---|---|
| Shape | A thin brief, not a specification document: the problem, the definition of good, the non-goals, the cost budget, and the line between what is fixed code and what is a model call. Short and load-bearing. |
| Prototype | The working version arrives in hours, in front of real people, internally first. The prototype is the specification until it cannot be. |
| Test it | The written test of good, agreed before anyone builds anything. A demo is one happy path; behaviour is a distribution. You do not ship judgment on vibes. |
| Ship | Behind a flag, to a slice of traffic, watched, reversible in one step. Deploy is not release. |
| Learn | Read what happened and decide what it is allowed to mean. Real usage is the research library. That feeds the next Shape. |
Shape, judge, learn. That is the job now, and it is why titles stopped predicting who is valuable.
An agent hits the target you give it, not the one you meant. A vague shape no longer stalls. It ships five confident wrong builds before lunch.
The tokens and inference baked into what you ship, paid again on every use. Meter it at design time, the way you track latency.
What your team burns using agents to move faster. A one-time bet. Leave it uncapped through the window where velocity compounds.
Mix them and you either starve the build or ship a product that loses money at scale. Volume discounts are safe at near-zero marginal cost and a trap when every action is metered: your heaviest users are your most expensive. An account that grows into negative margin is one you are paying to keep.
Each check goes to the cheapest thing that actually catches it. Only the calls that set a new standard reach a person.
In 2015 you could run a product off a laptop, the way you read a car’s diagnostics. Agents move at racing speed, and you cannot steer that blind. So you build a pit wall: one live surface across strategy, design, the build, and how it sells.
The job did not change. One person still makes the call. What multiplied around them is the machine: the screens, the data, and the agents that read it.
A racing team watches telemetry to decide, mid-race, where the next lap of effort goes. More into what is working, less into what is not, while the work is still moving. That is where strategy and operations meet.
Reading against strategy means reading against where the money is meant to go. Four pillars hold the allocation.
New value that opens accounts and markets.
The work that keeps the value you already won.
Bets on fit and a durable edge.
Efficiency and faster delivery that widen the spread.
02 · The build
The loop is the rhythm. Seven other layers sit under it, and the order matters more than the tooling. Building agents before the pipeline is the common early mistake.
| Layer | Phase | What it is |
|---|---|---|
| Developer experience | Phase 0 | The precondition for everything. One-command stack, continuous integration under five minutes, native toolchain. |
| The context agents run on | Phase 0 to 1 | Three specifications, not one. Covered below. |
| The loop | Phase 1 | The five moves, run on one workstream first. |
| Validation and review | Phase 1 to 2 | Two gates: validation, then review and auto-merge. |
| Where checks happen | Ongoing | Runs under every other layer, not between them. |
| What it costs | Phase 2 | What each feature costs to run, visible to the people deciding what to build. |
| Running features that can be wrong | Phase 2 to 3 | Reliability. Lives inside Ship. |
| Adoption and rhythm | Continuous | Culture. The actual limiter. |
Each phase opens the next one. Skip ahead and the gate you opened has nothing solid under it.
A pit wall is only as good as what feeds it. Better judgment is not a smarter model, it is better data, kept fresh and traceable. It needs a team that writes things down: context nobody wrote down does not exist for an agent, any more than for a new hire.
The judgment layer, made reusable: problem, definition of good, non-goals, cost budget, and the line between fixed code and model call.
Open, composable, token-based, published to a registry both people and tools read. The rules carry the reasoning and the evidence, not just the shipped result.
Generated from a single source so the reference cannot drift. A vague description makes an agent call the wrong endpoint, so operation names are user-facing copy.
Every model call records what it did. Without that, none of it is checkable a month later.
Naming vendors here dates within a quarter. The tiering is the durable part.
| Tier | What it sees |
|---|---|
| Tier 1 | Source code or customer data. The short list. Every one needs a signed agreement, no-training confirmed in writing, and minimum retention. |
| Tier 2 | Usage signals and metadata, with personal data possible in the logs. Strip it on the way in. |
| Tier 3 | Nothing sensitive, or self-hostable. Move things here where you can. |
03 · The read
Almost everyone measures how much they shipped. That number misleads now, because shipping got easy for your competitors too. A faster factory cannot tell you which shipments were worth making.
So every launch carries a standing verdict, set from the measured numbers rather than the mood of the room.
| Verdict | What it means |
|---|---|
| Landed | The number you said would move, moved. |
| Watch | It is early and the signal is thin. |
| Stalled | It shipped and nothing happened. |
Read from data that exists.
Compared to a known bar.
A forward bet, not yet proven.
How much you know is a different axis from how good the number is, so one mark never carries both. Where something spans two labels, split it rather than average it into nothing.
In practice: a fixed set of hand-picked cases, each with a rubric. A model scores them, the results are cached so the build can read them without paying for a fresh run, and a nightly sample of real traffic raises an alert when quality slips.
The rule nobody wants to hold: never lower the pass bar to hide a drop in quality. Fix the prompt, or accept the truth.
04 · The ladder
A model you cannot locate yourself on is a lecture. Three views, each with its own ladder, because the functions they measure mature differently.
A team that scores well on how it builds and poorly on what it shipped has a translation problem, not a capability problem. The fix is nothing like the fix for a team that is weak on both.
| Framework | The question | Units |
|---|---|---|
| Team Operations | People: where you are | 27 dimensions across 6 team functions (strategy, design, development, intelligence, operations, go-to-market), each scored 1 to 5. |
| Development Lifecycle | Process: how you build | 34 tasks, plus three concerns that cut across all of them: token economics, role fluidity, cognitive debt. |
| Product Assessment | Product: what you shipped | 27 dimensions across 6 product attributes (architecture, adaptive experience, learning systems, economics, trust and reliability, compound mechanics). |
Collapsed into one read, the stages are verbs: what a team does, not what it is.
Foundation → Building → Scaling → Leading → Compounding
Specify → Context → Orchestrate → Validate → Ship → Compound
Wrapper → Augmented → Integrated → Native → Compounding
05 · Sources
An illustrative model, not a report on any one company. The loop, the two costs, the pit wall, the read, and the ladder are a synthesis of work done across several product organizations and a corpus of published primary sources. Nothing here is a claim about a named company's internal practice unless it is attributed below.
The Linear Method. The Resend handbook (company, people, engineering, design, success, marketing and sales) and its philosophy page. Anthropic's prototype-first process, via Aakash Gupta. Vercel on teaching agents product design, and on prototyping with design systems. Moesif's teardown of Stripe's developer experience. Ricardo Cardoso on organizations and AI.
Partial: Stripe's own writing on API change flow, summary only. Not pulled: several mid-value entries from the Resend handbook.
Amplitude published its own before-and-after from rebuilding its software factory: pull request cycle time from 5.2 hours to 44 minutes, front-end continuous integration from 30 minutes to 3 or 4, bug reports from 715 to 319 a month, and non-engineer pull requests from none to roughly five per cent. That is a company reporting on itself. It is the shape of the change, not independent validation, and it is the only measured before-and-after in this document.
Section 02 names no vendors, on purpose: any list dates within a quarter. The tiering came out of a vendor review done in July 2026, and two things from it are worth checking yourself before you rely on them. Free and personal tiers of coding assistants often train on your data unless you opt out, where the business tiers do not. And several vertical software vendors publish no data processing agreement at all, so you have to ask.
The ninety-day sequence and the stage bands are design judgments, not measured outcomes. A starting point to push on, not a result to copy.
Darren Card. Fractional product and technology leadership, and hands-on building. dacard.ai/consulting