Business & Strategy

Before You Scope It, Test the Stack’s Hidden Costs

A product manager choosing a framework for a user-facing workflow is also choosing how failures will appear to customers and who will fix them. I would not shortlist vendors by throughput charts or polished demos, because neither reveals what happens when one dependency times out after a customer clicks Pay. Buy the smallest implementation that can prove a recoverable user journey and an affordable on-call contract.

The unit of evaluation is a failed journey, not a successful request

Ship faster or sleep better Choosing microservices vs monoliths frames a useful architectural tension, but a product manager evaluating a backend-for-frontend (BFF) framework needs a narrower test: can a customer finish, safely retry, or understand why a task stopped? A BFF is the server endpoint between the interface and its dependencies. Its value is not the number of services it can call; it is its ability to give the interface an honest answer when those calls disagree or fail.

Pick one revenue- or trust-sensitive journey before inviting vendors to demonstrate anything. For a checkout, draw the path from the Pay button through the BFF, payment provider, order record, confirmation screen, and support view. Mark the point at which a payment becomes irreversible. A framework trial that stops at an HTTP 200 response cannot establish whether the confirmation screen tells the truth, because a successful response from one dependency may precede a failed write to another.

Ask each candidate to handle three deliberately injected conditions: a dependency that takes longer than the interface can wait, a duplicate submission after a customer retries, and an accepted operation whose confirmation response is lost. The expected outputs should be written before implementation. For example, a timeout may produce a pending state with a lookup route; it must not invite an immediate second charge when the first attempt might have succeeded. That distinction belongs in the scope, since it needs product copy, endpoint behavior, and support guidance rather than a server-side retry alone.

The 7 UX mistakes that page DevOps at 2 a.m. and their cost puts operational consequences beside interface choices, but a vendor evaluation should demand evidence for one specific chain: what the customer saw, what the system recorded, and what alerted an operator. Use WCAG 2.2 as a check on the recovery interface: an error message that relies only on color or strands keyboard users is still a failed recovery, even if the backend returned the intended status.

Make acceptance criteria observable. Give the journey a provisional 800 ms p95 server-response budget, a threshold to tune against real traffic rather than a claim that every customer expects that number. Also specify the maximum wait before the interface changes from “processing” to “check status.” That second limit prevents a fast BFF response from disguising an indefinitely unresolved task.

Contract tests reveal costs that demonstrations conceal

Require a small executable slice instead of a slide deck: one endpoint, one interface state for success, one for uncertainty, and one for a definite failure. OpenAPI 3.1 should describe the request and response shapes so the interface and server teams can review the same contract. RFC 9457 Problem Details provides a standard error envelope, but require an application-specific error code as well, because “service unavailable” alone cannot tell the interface whether to offer a retry or a status lookup.

This Fastify 5 mock is deliberately too small for production; it runs and gives a team a repeatable 503 response with which to start the interface test:

npm init -y && npm install fastify@5
node --input-type=module <<'EOF'
import Fastify from 'fastify';
const app = Fastify({ logger: true });
app.get('/api/checkout-summary', async (req, reply) => {
  if (req.query.slow === '1') {
    return reply.code(503).type('application/problem+json')
      .send({ type: 'about:blank', title: 'Summary unavailable', status: 503 });
  }
  return { totalCents: 1299, currency: 'USD' };
});
await app.listen({ port: 3000, host: '127.0.0.1' });
EOF

Calling /api/checkout-summary?slow=1 exercises the error response, but it does not prove retry safety, authentication, or recovery from a real timeout; those require separate trial tasks. Have the team add a documented idempotency-key rule for the irreversible operation and test what happens when the same key arrives twice. A vendor claim that retries are “built in” is insufficient because only the application can define which operations may be repeated without duplicating an outcome.

Playwright can check what the customer actually sees after each injected response, while Pact can detect a contract change between a BFF and a dependency before release. k6 can apply a controlled load to the slice, but its p95 result should be paired with browser observations because a quick error response can meet a latency target while leaving a customer stuck. In an illustrative trial report, a measured p95 of 920 ms would be evidence against the provisional 800 ms budget, not proof that the framework is inherently slow; the team would first inspect dependency time, hosting, and the test workload.

Scope these as deliverables, not as vague “testing.” The estimate needs time to define failure states with design, implement the endpoint and interface states, build fixtures for duplicate and delayed responses, automate assertions, and fix what the tests expose. I would not approve a full migration on the strength of this slice, because its purpose is to price the remaining work and find a reason to stop early.

Next.js 15 and Fastify 5 win different contracts

Compare named options against the journey rather than asking which framework is more modern. Next.js 15 wins when the team already owns a Next.js interface and needs a small set of Route Handlers close to its rendering code, because one repository and deployment path can reduce coordination for interface changes. Its cost is coupling: server changes may share release timing, hosting constraints, and incident ownership with the interface. The trial should verify the deployed cache behavior and runtime of each route rather than assume local development matches production.

Fastify 5 wins when the BFF needs an independently deployed API with explicit route schemas and a separate owner, because its narrower server role makes the interface-to-API boundary visible. Its cost is another deployable service, with its own authentication integration, deployment pipeline, dashboards, and rollback procedure. Fastify’s published Node.js 20-or-newer requirement is a procurement check: an organization pinned to an older runtime must include an upgrade in the estimate, rather than treating installation as the whole adoption cost.

Neither choice should be credited with business behavior supplied by surrounding systems. OAuth 2.0 with PKCE may protect the customer sign-in flow, but the trial must show how the chosen deployment validates credentials and propagates identity to dependencies. W3C Trace Context can carry a trace across calls, but the team must test whether its gateway and dependencies preserve the headers. OpenTelemetry can produce spans for those calls; it cannot decide which failed payment requires a person to wake up. These distinctions keep a framework feature checklist from swallowing integration work.

Ask engineering for two estimates for each option: the effort to ship the slice and the recurring effort to operate it. As a planning assumption, reserve 10 working days for a time-boxed trial with an engineer, a designer available for error states, and an operations reviewer; revise that allowance if authentication or a payment sandbox must be built first. Record separately the effort for production hardening, including secret handling, rate limits, accessibility review, and a rollback rehearsal. A cheap proof of concept is not a cheap launch when those items remain unowned.

The on-call contract belongs in the purchase price

Before selecting a vendor, require a live incident exercise in a non-production environment. Break one dependency, have a tester attempt the journey, and ask an operator to find the affected request without being given its timestamp. The operator should be able to move from a customer-visible reference to a trace, identify whether the operation completed, and tell support which action is safe. If that lookup requires searching raw logs for personal data, the product has an incident-design task as well as a tooling task.

Prometheus can measure request outcomes and Alertmanager can route alerts; Grafana can show the trend, and Sentry can connect interface errors to affected sessions. Buy or configure only what the exercise uses, because another dashboard is a recurring maintenance cost when nobody has defined an action for its signals. A starting availability objective of 99.9% over 30 days permits about 43.2 minutes of failure by arithmetic; treat that figure as a proposed service target, not as evidence that customers can complete checkout. Track a journey-level completion measure alongside server availability so a healthy endpoint cannot conceal a broken confirmation screen.

Put ownership and exit terms in the vendor review. Ask who updates the library after a security release, whether exported traces and logs use portable OpenTelemetry formats, what happens when a hosted component is unavailable, and how long a rollback takes in a rehearsal. Request the vendor’s published support window in writing, then compare it with the team’s release cadence; a support promise that expires before the planned migration finishes creates work the initial quote omits.

The decision record should name the chosen option, the failed-journey evidence, open risks, and a stop condition. For example, stop the purchase if the team cannot demonstrate a safe answer to a lost confirmation response within the trial, because adding servers or dashboards will not settle whether the customer should retry. That is a position a vendor may contest, but it gives a product manager a bounded decision instead of an open-ended platform commitment.

A one-page failure brief should come before a vendor call

Start by writing one page for the journey most likely to cause customer harm: its irreversible step, three failure injections, the messages customers should see, and the evidence support needs to resolve uncertainty. Give that page to engineering, design, and the prospective vendor before requesting estimates. If they cannot agree on the safe retry behavior, scope that decision first; choosing a framework will only make the unresolved behavior faster to deploy.