Skip to content
All articles
Engineering

Your Backend Is a Kitchen With No Takeout Counter

A customer asked to connect their own tools to our product. We had hundreds of working endpoints and nothing to hand them. This is the build log for the public API — including the incidents that taught us what to log.

Rushit Jivani8 min read

The short version

  • A customer wanted their data flowing into their own spreadsheet, automatically. Our backend was hundreds of Firebase callable functions that only our own app could call — a working kitchen with no takeout counter.
  • The fix wasn't a new backend. A thin REST gateway sits in front of the same functions the app already calls. API keys are the membership card, scopes are what the card can order, and the request id on every response is the order number.
  • The war stories shaped the rest: a feature flag that silently 404'd every request taught us to log rejections, a write that succeeded while returning a 500 taught us why idempotency keys exist, and an API that could delete things the UI couldn't taught us where parity rules must live.

The email that started it

A customer wrote in with a simple ask: every time someone new signs up on one of my records, add a row to my Google Sheet. They weren't asking for a feature. They had a spreadsheet, a Tuesday-morning routine, and a product that held their data hostage behind a login screen.

Without an API, people solve this in exactly two ways, and both are bad. They export by hand every week and forget every other week. Or — and this is the one that should scare you — they hand their password to some third-party tool so it can log in as them and scrape the data out. Full account access, forever, to move one column of one table.

We couldn't say yes. The backend had hundreds of endpoints doing exactly this work all day — but every one was a Firebase callable function, shaped for our own frontend, authenticated by our own app's session. Useful to exactly one caller: us.

A kitchen with no takeout counter

Here's the picture that unlocked the project for me.

Our backend is a restaurant kitchen. It's genuinely good — it cooks thousands of orders a day. But every order arrives from our own dining room, on our own waiters' notepads, in our own shorthand. There is no window where someone from outside can walk up and order.

The customer with the spreadsheet isn't asking us to cook anything new. The dish they want — "list the signups on my record" — gets cooked hundreds of times an hour for our own app. They're asking for a takeout counter: a window on the side of the building where they can show some ID, order off a menu, and get the same food in a box.

That reframing tells you what not to build. You don't build a second kitchen. You build:

  • a membership card, so the counter knows who's ordering — the API key
  • a menu posted by the window — the OpenAPI doc
  • an order number on every receipt, so "my order was wrong" is solvable — the request id
  • a line that moves fairly, so one person can't order 500 meals — the rate limit
  • and standard packaging — the response envelope

One kitchen, two windows

the kitchen

hundreds of shared pure functions — one recipe per dish, cooked for every caller

the dining room

Our own app. Orders arrive on our own waiters' notepads: callable functions, session auth, UI-shaped payloads.

the takeout counter

The REST gateway. Show some ID, order off a posted menu, get the same food in a box.

  • membership cardthe API key
  • menu by the windowthe OpenAPI doc
  • order numberthe request id
  • a line that moves fairlythe rate limit
  • standard packagingthe response envelope
No second kitchen. The public API is a window on the side of the same building, plus the five things every counter needs.

Every technical decision below is one of those five things.

The build, as the commit history tells it

I went back through the git log while writing this, and the honest shape of the project is right there.

The gateway landed as nine chained PRs. We'd estimated the whole thing at two or three weeks; it became nine pull requests instead, each small enough to actually review — which is what you want for the code that will hold credentials. And every one of them left main deployable: anything that wasn't ready yet simply wasn't mounted, rather than shipped half-working and hidden.

Then the writes came, and with them the middleware you don't think about on day one. Reads are forgiving — the worst failed read is a retry. Writes needed three new layers in front of the pure functions:

  • Idempotency for writes. A client that times out will retry, and a retried "approve this person" must not approve twice. Write routes accept an idempotency key; the middleware caches the response body and replays it on retry instead of re-running the operation.
  • Write-path guards and an audit trail. Every mutating call leaves an audit record of who did what through which key.
  • Rate limiting split by read and write, with much tighter write budgets.

Each of those exists because of a specific way an external caller behaves that our own frontend never did. Our app doesn't retry blindly; a cron job on someone's server absolutely does.

A request walks through the gateway

  1. 1The key arrivesSHA-256 the bearer token, look up the hash. Revoked or expired? Rejected — with a log line saying why.
  2. 2Scope checkrequireScope("records:read"). A read key cannot write, no matter what the request says.
  3. 3Rate limitSeparate read and write budgets. The remaining budget is printed on every response header.
  4. 4Idempotencywrites onlyWrite routes only: a retried "approve this person" replays the cached response instead of approving twice.
  5. 5The same pure function the app callsThe counter never cooks. No permission rule or calculation is re-implemented at the route.
  6. 6One envelope out{ data, meta: { request_id } } — success or failure, the order number is on the receipt.
Each layer exists because of something an external caller does that our own frontend never did.

Three incidents, three rules

If this were an ad, it worked on the first try. The commit log says otherwise, and the three best lessons each came with a timestamp.

Three incidents, three rules

  1. incident 1The flag that 404'd everything

    A deploy set one feature flag but not the other. The gateway rejected every request with a blanket 404 — indistinguishable from the API not existing.

    the ruleEvery deliberate rejection says why, somewhere we can see.

  2. incident 2The 500 that lied

    A success body with data: undefined blew up the idempotency cache write inside the response hook. Completed operations returned errors that invited retries.

    the ruleSafety middleware is critical-path code. Test the empty-body success case.

  3. incident 3The API that could do more than the UI

    Hard-delete of a live record, a required name blanked after create, a system-only status settable from the payload. No errors — just drift.

    the ruleParity fixes go into the shared functions, not the route.

Each rule below came with a timestamp in the commit log, not from a design discussion.

The flag that 404'd everything — why we log rejections

The gateway shipped behind a feature flag, as it should. Then staging went up, and every single request — any route, valid key or not — came back as a blanket 404.

The cause was almost embarrassing: the deploy workflow set one feature flag but not the other, so the gateway saw its kill switch as "off" and rejected everything. The painful part wasn't the bug. It was that the rejection was invisible. An unset flag and a genuinely missing API look identical from outside: 404, no body worth reading, nothing in the logs. We burned real debugging time proving the API existed at all.

The fix was two lines of env config — and one logger.warn that fires whenever the gateway rejects a request because of the flag, so the kill switch shows up in Cloud Logging as "rejected: flag off" instead of silence.

That incident set our logging policy more than any design discussion:

  • Every deliberate rejection says why, somewhere we can see. Flag off, bad key, missing scope, rate limited — each has a log line with context. The caller gets the safe public error; the log gets the real reason.
  • Every 5xx logs structured context: request id, route, key id (never the key), identity, error code. The error handler writes one structured entry per failure, which is what makes "send me the order number" actually work — the request id on the customer's receipt matches one line in our logs.
  • Every fire-and-forget write logs its own failure. Usage counters, audit records, idempotency cache writes, and the last-used timestamp all happen off the hot path, which means they're allowed to fail without breaking the request — but each one has a .catch that warns. Bookkeeping may fail silently by design; it must never fail invisibly.

The 500 that lied — why idempotency is harder than it looks

A few weeks after writes shipped, approve and delete calls started returning 500s… while actually succeeding. The operation ran, the data changed, and the caller got an error telling them to try again. For an API, that's the most dangerous bug class there is: a retry on a "failed" write that actually succeeded is how things get approved twice and deleted twice. Idempotency middleware exists to prevent exactly this — and it was the thing causing it.

The mechanics were subtle. Some success responses legitimately have no data payload — { success, message, data: undefined }. The idempotency middleware caches response bodies to Firestore for replay, and Firestore rejects undefined values synchronously. So the cache write blew up inside the response hook itself, before its own error handling could catch it, and turned a completed operation into a 500.

Two takeaways we kept. First, the middleware you add for safety is itself code on the critical path — our "replay protection" sat inside the response serializer, where its failure became the caller's failure. Second, test write routes whose success body is empty; every example and test had a payload, and the empty case is the one that shipped broken.

The API that could do more than the UI — why rules live in the kitchen

The scariest class of bug had no error at all. An audit a few weeks after the write routes landed found three places where the counter accepted orders the dining room never would:

  • A live record could be hard-deleted through the API. The UI only offers "close" once something is live; the API checked the wrong status and allowed real deletion.
  • A record's name could be blanked between create and publish — the UI validates at publish time, and the API's update route happily emptied a required field afterward.
  • A status field was settable directly through create and update payloads — classic mass-assignment, where a client writes a field that only the system should transition.

The fixes are the whole point of the architecture. The first two guards went into the shared pure functions — the kitchen — so the app and the API are incapable of disagreeing again. Only the third was a route-level guard, because it's about what the payload may contain, not what the operation may do. When you find drift between your app and your API, the fix belongs in the shared layer; patching the route is how you get to disagree twice.

The membership card

A credential leaks — through logs, screenshots, paste bins — so the card's design is security design.

  • It names itself. Keys start with a self-identifying prefix, like Stripe's sk_live_. When one shows up in a log, nobody guesses what leaked, and a scanner can match it.
  • We keep a photo of the card, not the card. Keys are stored as SHA-256 hashes; the raw key is shown exactly once, next to a copy button. After that: name, prefix, scopes, created, last used.
  • The card lists what you may order. Scopes are checkboxes at creation, named resource:verb. A read key cannot write, no matter what the request says.
  • Other people's contact details need a special card. The people signed up on a customer's record have emails and phone numbers — the most sensitive thing this API returns. That lives behind its own :pii scope nobody gets by default; the counter hands out that dish redacted unless the card explicitly says otherwise.

The :pii scope — try it

The card

sk_live_redacted

records:read← click to grant or revoke

{

"name": "Alex P.",

"email": "[redacted]",

"phone": "[redacted]"

}

no :pii scope — contact fields redacted

Contact details are the most sensitive thing this API returns. The dish is handed out redacted unless the card explicitly says otherwise.

The route, in the shipped shape

// middleware upstream: SHA-256 the bearer token, look up the key,
// reject revoked/expired, attach { identityId, scopes } to the request.
 
router.get(
  "/v1/records/:recordId",
  requireScope("records:read"),
  async (req, res) => {
    const record = await loadRecord(req.params.recordId) // the one read
 
    // "not yours" looks identical to "doesn't exist"
    if (!record || !canView(record, req.identity)) {
      return res
        .status(404)
        .json(errorBody("not_found", "Record not found", req.id))
    }
 
    // the same function our own app calls — one kitchen, one recipe
    const data = await getRecordDetails(record, req.identity)
 
    const body = req.scopes.includes("records:people:pii")
      ? data
      : redactContactFields(data)
 
    res.json({ data: body, meta: { request_id: req.id } })
  },
)

The deliberate choices: one read, with the permission check running on the already-loaded object (early handlers fetched once to authorize and again to serve). 404, never 403, for anything the caller can't see — a 403 on someone else's record confirms it exists, which lets anyone enumerate ids. And the order number on everything, success or failure.

The menu is posted outside the door too: the OpenAPI document is served at its own URL with no auth. Code generators, API tools, and — increasingly — AI agents read the menu before they order.

The line is printed on your receipt

The naive rate limit is a silent 429; the caller's retry loop turns one overage into a hundred. Instead, the queue status rides on every receipt — X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset — and a 429 carries Retry-After with an actual number.

This paid off in a way we didn't plan: when AI assistants became API callers, a model told "14 requests left" slows down; a model handed a bare 429 retries forever.

The counter paid for itself

The quiet reason to build all this: the gateway became the only door into the building.

When we built an MCP server so AI assistants could use the product, it was a thin adapter over this API — same keys, same scopes, same rate limits, zero new security decisions. When OAuth arrived, its tokens flow through the same middleware as keys; routes can't tell the difference. A future SDK is code generated from the menu.

And the customer with the spreadsheet? Their nightly sync is a small script now: one key with two read scopes, one GET per night, one new row per signup. No password shared, revocable in one click — and if it ever breaks, they send us the order number, and the order number finds the log line.

Steal my workflow

  1. Start from a real external request — someone who wants their own data out. Build the counter for them, not an imagined marketplace.
  2. Ship in small chained PRs that each leave main deployable. Don't mount what isn't ready.
  3. Hash keys, show them once, prefix them so leaks identify themselves. PII behind its own scope.
  4. Mount a thin gateway over the functions you already have. When the API and app drift, fix the shared function, not the route.
  5. One envelope, a request id on everything, 404 for anything the caller can't see.
  6. Log every deliberate rejection with its real reason, every 5xx with context, and every fire-and-forget failure with a warning. Silent is worse than broken.
  7. Add idempotency to writes before external retries find them — and test the success case with an empty body.
  8. Print the rate budget on every response. Split read and write pools.

FAQ

Those endpoints are shaped by our UI and authenticated by our session. Every frontend refactor would break someone's integration, and a session credential is all-or-nothing — the opposite of a scoped card.

  • #RESTAPI
  • #APIDesign
  • #Firebase
  • #Nodejs
  • #Security
  • #PlatformEngineering
Rushit Jivani headshot

Written by

Rushit Jivani

Senior Software Architect & AI Application Engineer. Architecture that shows up in crash rate, performance and how fast the team ships.