Issue #3618198: Write a session's tariff and allotment anchors when it is created, so a rush does not deadlock writing them
Two of the four lock-anchor scopes were never bookkept.
LockAnchors::ensure() says in its own docblock that it is "called when the thing it anchors is created". That contract is met for SCOPE_SLOT (yoyaku_slot_insert) and SCOPE_SECTION (the same hook in yoyaku_placement). Nothing anywhere called it for SCOPE_TARIFF or SCOPE_ALLOTMENT: the only code that ever created those rows was the repair inside lock() itself.
On a session with no capacity of its own, the tariff's quota is the number that bounds a booker, so the tariff anchor is the row that booker waits at. Every session therefore reached its first booker with it missing, and locking a row that does not exist takes a gap lock rather than a row lock.
What is in it
LockAnchorHookswrites an anchor for every tariff and allotment the resource already carries when a session is created, andyoyaku_slot_tariff_insert/yoyaku_slot_allotment_insertwrite one for anything given to a session afterwards. Both reads are guarded by a table check, since a session can be created before everything that might bound it exists and a kernel test installs only the entity schemas it is about. Without that guard, every test that creates a session without the allotment schema fatals, which is how it was caught.LockAnchors::lock()sorts before repairing. The read is ordered for the express purpose of stopping two holds taking the same rows in opposite orders; the repair then wrote in whatever order the caller's lines arrived.
Measured, on the load harness, over HTTP against MySQL
Forty bookers arriving together, each claiming a party of three at a session whose quota leaves two, so the last two units are contested by many parties at once. Reconciliation on every round: what the bookers were told they hold, summed from the endpoint's own answers, against the held lines in the database.
| condition | rounds | failures |
|---|---|---|
| anchors present | 20 | 0 of ~800 claims |
| tariff anchor stripped each round | 10 | 13 of 400, in 8 rounds of 10 |
| stripped, with the ordering fix only | 10 | 9 of 400 |
| session born with its anchors, under this MR | 10 | 0 of 400 |
Every failure is SQLSTATE[40001] Serialization failure: 1213 Deadlock found on the anchor read. The gap was zero in all 50 rounds: nothing was overbooked and no unit was ever committed without being reported, so this costs a click rather than a seat.
The ordering row is reported honestly: on its own it is noise, and the MR does not claim otherwise. It is in because a cycle this method invents is one it can stop inventing.
Tests
Three in LockAnchorsTest, each seen to fail against 1.x first: 1 anchor where 3 were wanted, 1 where 2, and the two repair orders disagreeing.
Left for its own issue
Two repairs can still meet over the gap where a row is about to be. Retrying is not valid there, because the rollback a deadlock causes is of the whole transaction, so the fix is to write missing anchors before the hold opens its transaction. That is a change to where the repair is called from rather than to what it does.