Free lesson · about 15 minutes
“Try again” needs an expiry date
An app shows “Something went wrong. Try again.” The user taps it. Sometimes that saves their note. Sometimes it saves it twice, because the first attempt had already worked. This lesson shows which retry rules prevent that, and the one cleanup job that quietly breaks most of them.
No sign-up, no tracking, nothing stored. Every number on this page is an invented teaching constant, not a measurement of any real product.
The misconception
“An error means nothing was saved.” A timeout only tells the app that no reply arrived. The save may have failed before it reached the database, or it may have committed a millisecond before the connection dropped. From the phone, those two look identical.
The usual fix is an idempotency key: the app sends the same random key with every attempt at the same save, and the server remembers which keys it has already handled. That works until the server forgets. Every record of a handled key gets deleted eventually, and a retry that arrives after that point looks brand new.
Try it on a deliberately small model
One operation: save a note. The app promises it's safe to retry for 10 minutes. A cleanup job runs at minute 12 and deletes saved replies older than 10 minutes. Pick what goes wrong, then pick the server's retry rule.
What happened, minute by minute
All cases, all rules
A rule fails a case if it saves the same note twice or tells the user something was saved when it wasn't. Click a cell to load it above.
Scroll the table sideways to see all four rules.
| Case | Naive retry | Stored response | Tombstone | Bounded age |
|---|
What the cases show
Retrying isn't the problem. When the server crashes before saving, every rule ends with exactly one note. The trouble starts when the save worked and only the reply was lost: with no key, “Try again” writes a second note.
A stored reply fixes the common cases. Keep the key and the reply, and a lost reply, a double click or a retry at the very edge of the window all end with one note.
Then the cleanup job runs. An offline phone sends its queued save at minute 15. The stored reply was deleted at minute 12, so the server has never heard of this key, and it saves the note again. The rule was correct for every retry inside the window and wrong for the one that arrived after it.
Two ways to close the gap. A tombstone keeps the key and a hash of the content after the reply is deleted, so the late request is refused as “already saved”. It costs a little storage, forever. A bounded age rule refuses any request created more than 10 minutes ago. It keeps nothing, but the user has to go and check whether their note exists. Both are safe. Neither is free.
Keep the content hash. Without it, a buggy client that reuses a key for a different note gets back the first note's “saved” reply, and the user believes their new note exists.
But “forever” is a promise too. The tombstone above is never deleted. The next section deletes it, and asks what a late request should be allowed to do then.
When the tombstone expires too
A tombstone can remember a save. Who remembers when the tombstone is gone? Same note, same lost reply, same cleanup at minute 12. Now the tombstone itself is removed once the server's clock passes minute 20, and three rules decide what a late replay may do:
- Indefinite tombstone. The existing baseline: the tombstone is never removed.
- Finite tombstone, client clock. Tombstone removed after minute 20; request age measured from the creation time the client sends.
- Finite tombstone + verified deadline. Tombstone removed after minute 20; the server checks a deadline fixed when the operation was issued.
The cases and expected results were written down before the model ran. The “verified deadline” is checked by a simulated issuer and verifier: a toy digest standing in for a server signature, not real cryptography.
Scroll the table sideways to see all three rules.
| Case | Indefinite tombstone | Finite tombstone, client clock | Finite tombstone + verified deadline |
|---|---|---|---|
| Replay at 19 Unchanged replay one minute before the tombstone may go. | Safe refused: already processed, reply no longer kept | Safe refused: already processed, reply no longer kept | Safe refused: already processed, reply no longer kept |
| Replay at 20, cleanup first The cleanup job runs at minute 20, then the replay arrives at 20. | Safe refused: already processed, reply no longer kept | Safe refused: already processed, reply no longer kept | Safe refused: already processed, reply no longer kept |
| Replay at 20, request first The replay arrives at minute 20, then the cleanup job runs at 20. | Safe refused: already processed, reply no longer kept | Safe refused: already processed, reply no longer kept | Safe refused: already processed, reply no longer kept |
| Replay at 21 Unchanged replay after the finite tombstone is removed. | Safe refused: already processed, reply no longer kept | Safe · user checks refused: too old to retry safely | Safe · user checks refused: past the operation deadline |
| Replay at 21, clock reset Same late replay, but a faulty client resets its creation time to 21. | Safe refused: already processed, reply no longer kept | Duplicate saved | Safe · user checks refused: past the operation deadline |
| New content, same key A different note under the same key at minute 15, carrying the original token. | Safe refused: key reused with different content | Safe refused: key reused with different content | Safe · user checks refused: missing or invalid operation metadata |
| No metadata An unchanged replay at minute 15 that lost its operation token. | Safe refused: already processed, reply no longer kept | Safe refused: already processed, reply no longer kept | Safe · user checks refused: missing or invalid operation metadata |
| Edited deadline A replay at minute 22 whose token deadline was edited from 20 to 40. | Safe refused: already processed, reply no longer kept | Safe · user checks refused: too old to retry safely | Safe · user checks refused: missing or invalid operation metadata |
| New operation at 21 A genuinely new note at minute 21: new key, new content, its own deadline. | Safe saved | Safe saved | Safe saved |
| Fails before commit The server fails before saving at 0; the user retries at minute 5. | Safe saved | Safe saved | Safe saved |
Client clock, clock reset at minute 21
- 0 Save; commits, reply lost (created 0, received 0)
- 0 Age from the client's creation time: 0 ≤ 20
- 0 Decision: note saved, but the reply is lost · notes saved: 1
- 12 Cleanup job: removes the stored reply for k1, keeps a tombstone (digest 947639ee)
- 20 Cleanup job: nothing to remove
- 21 Cleanup job: removes the tombstone for k1 (minute 21 > 20)
- 21 Replay at minute 21; the client resets its creation time to 21 (created 21, received 21)
- 21 Age from the client's creation time: 0 ≤ 20
- 21 Decision: saved · notes saved: 2
Fails: Saved the same note 2 times.
Verified deadline, same reset
- 0 Save; commits, reply lost (created 0, received 0)
- 0 Verifier: valid; deadline 20 ≥ minute 0
- 0 Decision: note saved, but the reply is lost · notes saved: 1
- 12 Cleanup job: removes the stored reply for k1, keeps a tombstone (digest 947639ee)
- 20 Cleanup job: nothing to remove
- 21 Cleanup job: removes the tombstone for k1 (minute 21 > 20)
- 21 Replay at minute 21; the client resets its creation time to 21 (created 21, received 21)
- 21 Verifier: valid; deadline 20 < minute 21
- 21 Decision: refused: past the operation deadline · notes saved: 1
Safe: No duplicate, but the user has to check whether the note exists.
Forgetting is fine; trusting the client's clock is not. With the client clock, an unchanged late replay is refused as too old. But when a faulty client resets its creation time to minute 21, the age check passes, the tombstone is already gone, and the note is saved a second time. That was the predicted failure, and it happened.
A deadline the server fixed at the start can't be refreshed. The verified-deadline rule ignores the client's timestamp and checks a deadline bound to the key, the content and the caller when the operation began. It wrote no duplicate in any of its 10 cases, refused missing and edited metadata, and still saved a genuinely new note at minute 21.
The price is honesty, not storage. The indefinite tombstone never left a user unsure, but it holds its record forever. The verified deadline frees the record, and in 5 of 10 cases the user is left to check whether their note was saved. In 2 the request came after the deadline: “expired, check your notes”. In 3 its operation metadata was missing, edited or didn't match the content, sometimes while the tombstone still existed, so the server refused rather than guess. Both answers need their own message in the UI, different from “saved” and from “already saved”. These are ten hand-picked boundary cases, not a measured rate of anything.
The boundary is the rule, not the timing. At exactly minute 20 the tombstone is kept whether the cleanup job runs before or after the replay, because the rule removes it only when the clock is later than 20.
Check your reasoning
Three questions. Each answer explains itself, right or wrong.
What to do on a real system
- Write the retry window down in the user's words, for example: “If you see ‘Try again’, it's safe to tap for 10 minutes. After that, check your notes first.”
- Find the oldest retry a client can send, including offline queues and background sync, and make record retention outlive it. Or refuse requests older than the window.
- Store a content hash with every key, and refuse a reused key with different content.
- Count, per operation: notes written, false “saved” messages, and users left unsure. Report each case with its denominator, not one blended reliability score.
- If records expire, put the operation's deadline in metadata the server issues and verifies, bound to the key, the content and the user. Never let a client timestamp extend it.
- Give “expired, check your notes” its own message. It is not “saved”, and it is not “error”.
- Test the cases above on purpose, including a run with deduplication switched off, to prove your test can see a duplicate at all.
What this lesson does not show
- No distributed system. One server handles one request at a time, so the double click can't race. Real databases need a uniqueness constraint on the key to get the same result.
- No clocks that disagree. The first age rule trusts the client's creation time, and the expiry section shows how that fails. The verified deadline moves that decision to the server, whose single clock is assumed correct here.
- No security. The verifier is a simulation with a toy digest, and the first section's keys aren't scoped to a user. On a real system, use real signatures or server-side records, scoped per user.
- No claim about any product. The cases are synthetic and the numbers are invented.