Practice Explaining an Idempotency Bug
Trace an idempotency bug through lost responses, durable operation identity, concurrent retries and uncertain external effects.
TL;DR
- In this fictional shipment incident, trace how a retried request produced a duplicate before proposing a key.
- Define operation identity and coordinate deduplication with the durable business effect, treating external work as another boundary.
- Test replay, expiration, concurrency, changed parameters, and crashes after commit.
Trace the duplicate before proposing a key
Practice this fictional incident: a client submits a request to create a shipment. The server commits the shipment, but the response is lost. The client retries, and a second shipment appears. Explain the failure and design a retry path that does not repeat the same business operation.
Begin with the timeline. The client observed a timeout, not proof that the server did nothing. That uncertainty is the reason a retry can duplicate an effect. Saying “retry only on failure” does not solve the problem when the client cannot distinguish a failed operation from a successful operation with a missing response.
Define the operation identity
An idempotency key should identify one intended operation, not every network attempt. If the client generates a fresh key on each retry, the server cannot recognize the repeated intent. If it reuses one key for unrelated shipments, the server may incorrectly collapse distinct operations.
For this exercise, scope the key to the authenticated tenant and the create-shipment operation. Store a representation of the request parameters or a suitable comparison value so the same key cannot silently be reused with a different destination or package. Explain the conflict response rather than accepting whichever payload arrives last.
Stripe's idempotent-request documentation provides a concrete primary example of replaying a stored result and checking request consistency. Its retention and endpoint behavior are provider-specific; your fictional shipment service needs its own defined contract.
Place deduplication beside the durable effect
An in-memory set on one server is insufficient when another instance receives the retry or the first process restarts. Use a durable record with a uniqueness boundary appropriate to the operation. Then reason about atomicity between claiming the key and creating the shipment.
| Failure point | Question the design must answer |
|---|---|
| Before any durable write | Can a retry safely begin? |
| After claiming the key | How is an abandoned in-progress operation recovered? |
| After shipment commit | Can the existing result be returned? |
| During two concurrent attempts | Which attempt owns execution? |
| After a response is lost | Does retry find the same business result? |
If the shipment row and idempotency record are in one database, a transaction may let you coordinate them. Explain the exact state transition rather than assuming a unique key alone makes every side effect atomic.
Handle external work as a separate boundary
Suppose creating a shipment also calls a carrier API. A database transaction cannot automatically roll back an external carrier operation. You now need a plan for the uncertainty between local state and the remote effect.
One possible design records an intended task durably and lets a worker execute it with a stable operation identity, using the carrier's own idempotency support if available. If the carrier has no such support, explain how you would reconcile an uncertain result before repeating it. Do not promise exactly-once effects merely because the local queue deduplicates messages.
Our system design frameworks guide can help separate these boundaries. The interview answer should identify where the guarantee holds and where a different system's behavior limits it.
Define replay and expiration behavior
Decide what a retry receives while the first attempt is still running and after it completes. It might receive a defined in-progress response or the completed result. The client needs enough information to avoid starting a new logical operation unnecessarily.
Retention also matters. If the deduplication record expires before a late retry, the service may treat the request as new. Choose a retention policy based on the expected retry and business lifecycle, and document the limit. Do not copy a provider's retention period without checking whether it fits shipment creation.
A permanent business identifier can sometimes provide another protection against duplicates, but it solves a different question from response replay. Explain whether the shipment itself has a unique external reference and how that interacts with the idempotency record.
Discuss access to the stored result as well as uniqueness. A request must not obtain another tenant’s shipment merely by guessing or reusing a key. Bind lookup and replay to the authorized operation scope, and avoid recording unnecessary sensitive payload data in logs used to investigate duplicate attempts.
Rehearse concurrent and uncertain attempts
Test the design with two simultaneous requests using the same key, a retry with changed parameters and a process crash after the business commit. For each case, state the durable records and the response the client should see. A sequence diagram is useful only if those states are explicit.
Then change the prompt: two different keys contain identical shipment details. Should they be deduplicated? Not necessarily; they may represent two legitimate shipments. Operation identity is a product contract, not a guess based on similar payloads.
Use the mock interview strategy guide to repeat the explanation without relying on the word idempotent as shorthand. A strong answer traces the lost response, defines stable intent, coordinates durable state and acknowledges external uncertainty. That is much more convincing than adding a key header without explaining what it protects.