Fix PersistentQueue deadlock and add operational visibility - #2742
Conversation
Transaction cost differencesNo cost or size differences found |
End-to-end benchmark differencesComparing this PR ( Baseline Scenario
Three local nodes
|
Transaction costsSizes and execution budgets for Hydra protocol transactions. Note that unlisted parameters are currently using
Script summary
|
| Parties | Tx size | % max Mem | % max CPU | Min fee ₳ |
|---|---|---|---|---|
| 1 | 5355 | 9.03 | 2.97 | 0.48 |
| 2 | 5447 | 9.31 | 3.05 | 0.49 |
| 3 | 5543 | 10.29 | 3.39 | 0.50 |
| 5 | 5737 | 10.87 | 3.55 | 0.52 |
| 10 | 6217 | 13.13 | 4.27 | 0.56 |
| 50 | 10058 | 34.71 | 11.02 | 0.95 |
| 100 | 14859 | 61.69 | 19.47 | 1.44 |
| 115 | 16296 | 69.67 | 21.97 | 1.59 |
Cost of Increment Transaction
| Parties | Tx size | % max Mem | % max CPU | Min fee ₳ |
|---|---|---|---|---|
| 1 | 2278 | 17.76 | 6.40 | 0.44 |
| 2 | 2409 | 18.65 | 7.30 | 0.46 |
| 3 | 2540 | 20.24 | 8.43 | 0.49 |
| 5 | 2801 | 21.50 | 10.05 | 0.52 |
| 10 | 3456 | 26.01 | 14.56 | 0.62 |
| 50 | 8697 | 66.62 | 51.84 | 1.44 |
| 75 | 11975 | 91.37 | 74.90 | 1.96 |
Cost of Decrement Transaction
| Parties | Tx size | % max Mem | % max CPU | Min fee ₳ |
|---|---|---|---|---|
| 1 | 603 | 16.03 | 5.75 | 0.35 |
| 2 | 734 | 16.94 | 6.65 | 0.37 |
| 3 | 860 | 17.83 | 7.55 | 0.39 |
| 5 | 1122 | 19.57 | 9.33 | 0.43 |
| 10 | 1778 | 24.17 | 13.86 | 0.52 |
| 50 | 7019 | 62.20 | 50.32 | 1.32 |
| 75 | 10296 | 86.26 | 73.18 | 1.83 |
Close transaction costs
| Parties | Tx size | % max Mem | % max CPU | Min fee ₳ |
|---|---|---|---|---|
| 1 | 597 | 15.49 | 10.37 | 0.38 |
| 2 | 724 | 16.37 | 11.27 | 0.40 |
| 3 | 854 | 17.27 | 12.17 | 0.42 |
| 5 | 1117 | 19.16 | 14.00 | 0.46 |
| 10 | 1772 | 23.70 | 18.52 | 0.55 |
| 50 | 7009 | 61.72 | 55.07 | 1.35 |
| 75 | 10288 | 86.06 | 78.07 | 1.86 |
Contest transaction costs
| Parties | Tx size | % max Mem | % max CPU | Min fee ₳ |
|---|---|---|---|---|
| 1 | 628 | 18.79 | 13.48 | 0.43 |
| 2 | 755 | 19.89 | 14.44 | 0.45 |
| 3 | 885 | 20.97 | 15.40 | 0.47 |
| 5 | 1149 | 23.10 | 17.30 | 0.51 |
| 10 | 1804 | 28.56 | 22.10 | 0.62 |
| 50 | 7044 | 74.22 | 60.90 | 1.50 |
| 72 | 9927 | 98.41 | 81.98 | 1.97 |
FanOut transaction costs
Involves spending head output and burning head tokens. Uses ada-only UTXO for better comparability.
| Parties | UTxO | UTxO (bytes) | Tx size | % max Mem | % max CPU | Min fee ₳ |
|---|---|---|---|---|---|---|
| 10 | 0 | 0 | 5529 | 22.68 | 41.73 | 0.88 |
| 10 | 1 | 57 | 5563 | 25.03 | 44.25 | 0.92 |
| 10 | 5 | 284 | 5698 | 35.28 | 54.57 | 1.08 |
| 10 | 10 | 569 | 5869 | 49.29 | 67.85 | 1.30 |
| 10 | 20 | 1140 | 6209 | 81.82 | 95.83 | 1.78 |
| 10 | 20 | 1137 | 6206 | 81.82 | 95.83 | 1.78 |
PartialFanOut transaction costs
Largest chunk of ada-only outputs that can be distributed in one partial fanout step, computed dynamically. The last row is the maximum total UTxO count where at least one output can still be distributed.
| Distributed | UTxO (bytes) | Tx size | % max Mem | % max CPU | Min fee ₳ |
|---|---|---|---|---|---|
| 11 | 569 | 986 | 34.31 | 65.19 | 0.94 |
| 25 | 1310 | 1429 | 67.64 | 98.26 | 1.47 |
| 30 | 1309 | 1428 | 67.64 | 98.26 | 1.47 |
| 40 | 1310 | 1429 | 67.64 | 98.26 | 1.47 |
| 50 | 1310 | 1425 | 67.64 | 98.26 | 1.47 |
| 100 | 1309 | 1428 | 67.64 | 98.26 | 1.47 |
| 150 | 1309 | 1428 | 67.64 | 98.26 | 1.47 |
| 200 | 1306 | 1425 | 67.64 | 98.26 | 1.47 |
| 200 | 1311 | 1430 | 67.64 | 98.26 | 1.47 |
PartialFanOut transaction costs (with native tokens)
Largest chunk of native-token outputs that can be distributed in one partial fanout step, computed dynamically. The last row is the maximum total UTxO count where at least one output can still be distributed.
| Distributed | UTxO (bytes) | Tx size | % max Mem | % max CPU | Min fee ₳ |
|---|---|---|---|---|---|
| 11 | 1120 | 1609 | 41.40 | 67.74 | 1.04 |
| 25 | 2478 | 2742 | 75.99 | 98.08 | 1.59 |
| 30 | 2499 | 2764 | 75.97 | 98.12 | 1.59 |
| 40 | 2205 | 2456 | 75.99 | 98.03 | 1.58 |
| 50 | 2079 | 2321 | 75.97 | 97.98 | 1.57 |
| 100 | 2247 | 2501 | 75.99 | 98.03 | 1.58 |
| 150 | 2100 | 2347 | 75.99 | 97.98 | 1.57 |
| 200 | 2016 | 2259 | 75.97 | 97.97 | 1.57 |
| 200 | 2520 | 2787 | 75.99 | 98.13 | 1.59 |
FinalPartialFanOut transaction costs (with native tokens)
Terminal partial fanout step (FanoutProgress → Final) with outputs carrying a native token. Burns all head tokens and proves accumulator exhaustion via BLS proof.
| Distributed | UTxO (bytes) | Tx size | % max Mem | % max CPU | Min fee ₳ |
|---|---|---|---|---|---|
| 1 | 112 | 5417 | 21.54 | 43.18 | 0.87 |
| 5 | 590 | 5811 | 35.15 | 54.72 | 1.08 |
| 10 | 1190 | 6305 | 53.23 | 69.51 | 1.36 |
| 10 | 990 | 6105 | 53.35 | 69.48 | 1.35 |
End-to-end benchmark results
This page is intended to collect the latest end-to-end benchmark results produced by Hydra's continuous integration (CI) system from the latest master code.
Please note that these results are approximate as they are currently produced from limited cloud VMs and not controlled hardware. Rather than focusing on the absolute results, the emphasis should be on relative results, such as how the timings for a scenario evolve as the code changes.
Generated at 2026-07-08 08:39:27.775290624 UTC
Baseline Scenario
| Number of nodes | 1 |
|---|---|
| Number of txs | 300 |
| Avg. Confirmation Time (ms) | 514.4 |
| P99 | 527.9ms |
| P95 | 527.6ms |
| P50 | 518.6ms |
| End-to-end TPS | 563.06 tx/s |
| Snapshots observed | 4 |
| Per-snapshot TPS P50 | 2999.88 tx/s |
| Per-snapshot TPS P95 | 7686.48 tx/s |
| Per-snapshot TPS max | 8072.98 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Three local nodes
| Number of nodes | 3 |
|---|---|
| Number of txs | 900 |
| Avg. Confirmation Time (ms) | 2618.9 |
| P99 | 2897.8ms |
| P95 | 2880.0ms |
| P50 | 2665.1ms |
| End-to-end TPS | 309.69 tx/s |
| Snapshots observed | 10 |
| Per-snapshot TPS P50 | 1081.86 tx/s |
| Per-snapshot TPS P95 | 2354.89 tx/s |
| Per-snapshot TPS max | 2421.04 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Scenario benchmark results
This page collects results from the scenario matrix: every combination of cluster size, UTxO shape, and incremental-ops mode is exercised by CI from the latest master code and reported below.
Numbers are approximate. They come from cloud VMs rather than controlled hardware, so the useful signal is the relative change between cells and between commits, not the absolute throughput.
Generated at 2026-07-08 08:49:56.312987086 UTC
Summary across cells
TPS columns are rates (transactions per second); Wall clock (s) is the measured elapsed time from the first tx submission to the last confirmation. Times are rounded to one decimal.
| Scenario | Txs | Wall clock (s) | End-to-end TPS (tx/s) | Per-snapshot p50 TPS (tx/s) | Avg conf (ms) | P95 conf (ms) |
|---|---|---|---|---|---|---|
| Nodes=1, Constant, incremental ops off, fire and forget | 30 | 0.1 | 549.49 | 2505.87 | 53.8 | 54.3 |
| Nodes=1, Constant, incremental ops off, wait for tx valid | 30 | 0.2 | 171.16 | 179.17 | 5.8 | 7.6 |
| Nodes=1, Growing, incremental ops off, fire and forget | 30 | 0.1 | 412.60 | 1035.92 | 71.6 | 72.4 |
| Nodes=1, Growing, incremental ops off, wait for tx valid | 30 | 0.3 | 113.02 | 116.30 | 8.7 | 11.7 |
| Nodes=1, Mixed, incremental ops off, fire and forget | 30 | 0.1 | 514.32 | 2316.40 | 57.6 | 58.1 |
| Nodes=1, Mixed, incremental ops off, wait for tx valid | 30 | 0.2 | 140.81 | 141.73 | 7.0 | 8.8 |
| Nodes=2, Constant, incremental ops off, fire and forget | 60 | 0.2 | 394.44 | 1729.09 | 149.8 | 151.6 |
| Nodes=2, Constant, incremental ops off, wait for tx valid | 60 | 0.6 | 105.61 | 119.60 | 18.6 | 27.6 |
| Nodes=2, Growing, incremental ops off, fire and forget | 60 | 0.2 | 323.90 | 846.61 | 183.2 | 184.4 |
| Nodes=2, Growing, incremental ops off, wait for tx valid | 60 | 1.0 | 62.05 | 62.32 | 31.7 | 44.2 |
| Nodes=2, Mixed, incremental ops off, fire and forget | 60 | 0.2 | 336.33 | 1327.81 | 175.9 | 177.1 |
| Nodes=2, Mixed, incremental ops off, wait for tx valid | 60 | 0.8 | 79.06 | 80.32 | 25.0 | 34.6 |
| Nodes=3, Constant, incremental ops off, fire and forget | 90 | 0.3 | 332.04 | 1649.90 | 266.7 | 269.3 |
| Nodes=3, Constant, incremental ops off, wait for tx valid | 90 | 0.9 | 95.24 | 89.19 | 31.3 | 39.0 |
| Nodes=3, Growing, incremental ops off, fire and forget | 90 | 0.4 | 252.41 | 415.28 | 350.7 | 354.4 |
| Nodes=3, Growing, incremental ops off, wait for tx valid | 90 | 1.9 | 46.37 | 43.23 | 61.9 | 92.8 |
| Nodes=3, Mixed, incremental ops off, fire and forget | 90 | 0.3 | 310.03 | 1137.20 | 286.1 | 289.8 |
| Nodes=3, Mixed, incremental ops off, wait for tx valid | 90 | 1.5 | 61.65 | 55.93 | 47.5 | 65.5 |
Nodes=1, Constant, incremental ops off, fire and forget
| Number of nodes | 1 |
|---|---|
| Number of txs | 30 |
| Avg. Confirmation Time (ms) | 53.8 |
| P99 | 54.4ms |
| P95 | 54.3ms |
| P50 | 54.0ms |
| End-to-end TPS | 549.49 tx/s |
| Snapshots observed | 2 |
| Per-snapshot TPS P50 | 2505.87 tx/s |
| Per-snapshot TPS P95 | 4742.71 tx/s |
| Per-snapshot TPS max | 4941.54 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=1, Constant, incremental ops off, wait for tx valid
| Number of nodes | 1 |
|---|---|
| Number of txs | 30 |
| Avg. Confirmation Time (ms) | 5.8 |
| P99 | 7.8ms |
| P95 | 7.6ms |
| P50 | 5.5ms |
| End-to-end TPS | 171.16 tx/s |
| Snapshots observed | 30 |
| Per-snapshot TPS P50 | 179.17 tx/s |
| Per-snapshot TPS P95 | 192.62 tx/s |
| Per-snapshot TPS max | 193.75 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=1, Growing, incremental ops off, fire and forget
| Number of nodes | 1 |
|---|---|
| Number of txs | 30 |
| Avg. Confirmation Time (ms) | 71.6 |
| P99 | 72.5ms |
| P95 | 72.4ms |
| P50 | 72.1ms |
| End-to-end TPS | 412.60 tx/s |
| Snapshots observed | 2 |
| Per-snapshot TPS P50 | 1035.92 tx/s |
| Per-snapshot TPS P95 | 1952.87 tx/s |
| Per-snapshot TPS max | 2034.38 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=1, Growing, incremental ops off, wait for tx valid
| Number of nodes | 1 |
|---|---|
| Number of txs | 30 |
| Avg. Confirmation Time (ms) | 8.7 |
| P99 | 12.4ms |
| P95 | 11.7ms |
| P50 | 8.5ms |
| End-to-end TPS | 113.02 tx/s |
| Snapshots observed | 30 |
| Per-snapshot TPS P50 | 116.30 tx/s |
| Per-snapshot TPS P95 | 160.44 tx/s |
| Per-snapshot TPS max | 174.38 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=1, Mixed, incremental ops off, fire and forget
Each client first grows its UTxO set (1-in to 2-out) for half of its tx budget, then contracts it back (2-in to 1-out) for the remainder.
| Number of nodes | 1 |
|---|---|
| Number of txs | 30 |
| Avg. Confirmation Time (ms) | 57.6 |
| P99 | 58.1ms |
| P95 | 58.1ms |
| P50 | 57.8ms |
| End-to-end TPS | 514.32 tx/s |
| Snapshots observed | 2 |
| Per-snapshot TPS P50 | 2316.40 tx/s |
| Per-snapshot TPS P95 | 4383.86 tx/s |
| Per-snapshot TPS max | 4567.63 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=1, Mixed, incremental ops off, wait for tx valid
Each client first grows its UTxO set (1-in to 2-out) for half of its tx budget, then contracts it back (2-in to 1-out) for the remainder.
| Number of nodes | 1 |
|---|---|
| Number of txs | 30 |
| Avg. Confirmation Time (ms) | 7.0 |
| P99 | 11.3ms |
| P95 | 8.8ms |
| P50 | 7.0ms |
| End-to-end TPS | 140.81 tx/s |
| Snapshots observed | 30 |
| Per-snapshot TPS P50 | 141.73 tx/s |
| Per-snapshot TPS P95 | 178.99 tx/s |
| Per-snapshot TPS max | 187.83 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=2, Constant, incremental ops off, fire and forget
| Number of nodes | 2 |
|---|---|
| Number of txs | 60 |
| Avg. Confirmation Time (ms) | 149.8 |
| P99 | 151.7ms |
| P95 | 151.6ms |
| P50 | 150.1ms |
| End-to-end TPS | 394.44 tx/s |
| Snapshots observed | 2 |
| Per-snapshot TPS P50 | 1729.09 tx/s |
| Per-snapshot TPS P95 | 3278.57 tx/s |
| Per-snapshot TPS max | 3416.30 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=2, Constant, incremental ops off, wait for tx valid
| Number of nodes | 2 |
|---|---|
| Number of txs | 60 |
| Avg. Confirmation Time (ms) | 18.6 |
| P99 | 28.3ms |
| P95 | 27.6ms |
| P50 | 18.7ms |
| End-to-end TPS | 105.61 tx/s |
| Snapshots observed | 60 |
| Per-snapshot TPS P50 | 119.60 tx/s |
| Per-snapshot TPS P95 | 141.94 tx/s |
| Per-snapshot TPS max | 148.95 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=2, Growing, incremental ops off, fire and forget
| Number of nodes | 2 |
|---|---|
| Number of txs | 60 |
| Avg. Confirmation Time (ms) | 183.2 |
| P99 | 184.5ms |
| P95 | 184.4ms |
| P50 | 183.8ms |
| End-to-end TPS | 323.90 tx/s |
| Snapshots observed | 2 |
| Per-snapshot TPS P50 | 846.61 tx/s |
| Per-snapshot TPS P95 | 1602.53 tx/s |
| Per-snapshot TPS max | 1669.73 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=2, Growing, incremental ops off, wait for tx valid
| Number of nodes | 2 |
|---|---|
| Number of txs | 60 |
| Avg. Confirmation Time (ms) | 31.7 |
| P99 | 48.8ms |
| P95 | 44.2ms |
| P50 | 32.1ms |
| End-to-end TPS | 62.05 tx/s |
| Snapshots observed | 60 |
| Per-snapshot TPS P50 | 62.32 tx/s |
| Per-snapshot TPS P95 | 108.42 tx/s |
| Per-snapshot TPS max | 115.34 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=2, Mixed, incremental ops off, fire and forget
Each client first grows its UTxO set (1-in to 2-out) for half of its tx budget, then contracts it back (2-in to 1-out) for the remainder.
| Number of nodes | 2 |
|---|---|
| Number of txs | 60 |
| Avg. Confirmation Time (ms) | 175.9 |
| P99 | 177.7ms |
| P95 | 177.1ms |
| P50 | 176.2ms |
| End-to-end TPS | 336.33 tx/s |
| Snapshots observed | 2 |
| Per-snapshot TPS P50 | 1327.81 tx/s |
| Per-snapshot TPS P95 | 2516.99 tx/s |
| Per-snapshot TPS max | 2622.70 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=2, Mixed, incremental ops off, wait for tx valid
Each client first grows its UTxO set (1-in to 2-out) for half of its tx budget, then contracts it back (2-in to 1-out) for the remainder.
| Number of nodes | 2 |
|---|---|
| Number of txs | 60 |
| Avg. Confirmation Time (ms) | 25.0 |
| P99 | 36.2ms |
| P95 | 34.6ms |
| P50 | 25.6ms |
| End-to-end TPS | 79.06 tx/s |
| Snapshots observed | 60 |
| Per-snapshot TPS P50 | 80.32 tx/s |
| Per-snapshot TPS P95 | 120.21 tx/s |
| Per-snapshot TPS max | 138.94 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=3, Constant, incremental ops off, fire and forget
| Number of nodes | 3 |
|---|---|
| Number of txs | 90 |
| Avg. Confirmation Time (ms) | 266.7 |
| P99 | 269.4ms |
| P95 | 269.3ms |
| P50 | 268.6ms |
| End-to-end TPS | 332.04 tx/s |
| Snapshots observed | 2 |
| Per-snapshot TPS P50 | 1649.90 tx/s |
| Per-snapshot TPS P95 | 3131.09 tx/s |
| Per-snapshot TPS max | 3262.75 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=3, Constant, incremental ops off, wait for tx valid
| Number of nodes | 3 |
|---|---|
| Number of txs | 90 |
| Avg. Confirmation Time (ms) | 31.3 |
| P99 | 39.9ms |
| P95 | 39.0ms |
| P50 | 30.5ms |
| End-to-end TPS | 95.24 tx/s |
| Snapshots observed | 60 |
| Per-snapshot TPS P50 | 89.19 tx/s |
| Per-snapshot TPS P95 | 173.00 tx/s |
| Per-snapshot TPS max | 186.35 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=3, Growing, incremental ops off, fire and forget
| Number of nodes | 3 |
|---|---|
| Number of txs | 90 |
| Avg. Confirmation Time (ms) | 350.7 |
| P99 | 354.7ms |
| P95 | 354.4ms |
| P50 | 352.4ms |
| End-to-end TPS | 252.41 tx/s |
| Snapshots observed | 2 |
| Per-snapshot TPS P50 | 415.28 tx/s |
| Per-snapshot TPS P95 | 785.33 tx/s |
| Per-snapshot TPS max | 818.23 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=3, Growing, incremental ops off, wait for tx valid
| Number of nodes | 3 |
|---|---|
| Number of txs | 90 |
| Avg. Confirmation Time (ms) | 61.9 |
| P99 | 112.1ms |
| P95 | 92.8ms |
| P50 | 59.9ms |
| End-to-end TPS | 46.37 tx/s |
| Snapshots observed | 62 |
| Per-snapshot TPS P50 | 43.23 tx/s |
| Per-snapshot TPS P95 | 109.99 tx/s |
| Per-snapshot TPS max | 160.78 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=3, Mixed, incremental ops off, fire and forget
Each client first grows its UTxO set (1-in to 2-out) for half of its tx budget, then contracts it back (2-in to 1-out) for the remainder.
| Number of nodes | 3 |
|---|---|
| Number of txs | 90 |
| Avg. Confirmation Time (ms) | 286.1 |
| P99 | 290.2ms |
| P95 | 289.8ms |
| P50 | 287.5ms |
| End-to-end TPS | 310.03 tx/s |
| Snapshots observed | 2 |
| Per-snapshot TPS P50 | 1137.20 tx/s |
| Per-snapshot TPS P95 | 2157.06 tx/s |
| Per-snapshot TPS max | 2247.72 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
Nodes=3, Mixed, incremental ops off, wait for tx valid
Each client first grows its UTxO set (1-in to 2-out) for half of its tx budget, then contracts it back (2-in to 1-out) for the remainder.
| Number of nodes | 3 |
|---|---|
| Number of txs | 90 |
| Avg. Confirmation Time (ms) | 47.5 |
| P99 | 84.9ms |
| P95 | 65.5ms |
| P50 | 46.2ms |
| End-to-end TPS | 61.65 tx/s |
| Snapshots observed | 63 |
| Per-snapshot TPS P50 | 55.93 tx/s |
| Per-snapshot TPS P95 | 120.80 tx/s |
| Per-snapshot TPS max | 153.12 tx/s |
| Number of Invalid txs | 0 |
| Fanout outputs | 0 |
vrom911
left a comment
There was a problem hiding this comment.
I have a couple of comments 🙌🏼
d31a30f to
ec6d0f7
Compare
popPersistentQueue previously matched the dequeued item by value equality, which could silently no-op if Eq returned False, leaving the head item stuck and blocking all subsequent writes once capacity was reached. It now pops unconditionally by index — safe because broadcastMessages is the sole consumer. Also: guard against removeFile throwing isDoesNotExistError; roll back nextIx if writeFileBS fails; emit PersistentQueueFull and PersistentQueueLoadFailed traces so operators can see when the queue is under pressure or failed to load. Drops the Eq msg constraint from withEtcdNetwork and broadcastMessages. Signed-off-by: Sasha Bogicevic <sasha.bogicevic@iohk.io>
The onException guard was unnecessary: index gaps are harmless (files are loaded by sorted listing, not assumed contiguous), and the rollback has no effect on restart since nextIx is always derived fresh from disk. Signed-off-by: Sasha Bogicevic <sasha.bogicevic@iohk.io>
A removeFile error other than "does not exist" previously propagated out
of broadcastMessages' forever loop (its catch only handles GrpcException)
and tore down the whole network component. By the time the delete runs the
message was already broadcast and dequeued, so the failure is now traced
as PersistentQueueDeleteFailed and the loop keeps going; the leftover file
only means the message may be re-broadcast after a restart, same as the
existing crash-recovery path. The handler stays pinned to IOException so
async cancellation still propagates. Genuinely broken disks still fail
fast via the unguarded write in writePersistentQueue.
Also add the previously missing popPersistentQueue coverage, pinning the
liveness contract behind the deadlock fixed in 553cd0b: a writer blocked
at capacity must be released by a pop (using the PersistentQueueFull trace
as the block signal), a 100-item soak through a capacity-10 queue asserts
FIFO exactly-once delivery and full disk drainage, popped items are not
reloaded on restart, a missing backing file is tolerated, a failing
deletion (EISDIR) is traced without crashing, and corrupt items on load
are traced as PersistentQueueLoadFailed.
Signed-off-by: Sasha Bogicevic <sasha.bogicevic@iohk.io>
7f19373 to
ca08cb7
Compare
popPersistentQueuewas using value equality todecide whether to remove the head item. A false
Eqresult silentlyno-oped, leaving the head item stuck. With the queue at capacity (100),
all callers of
broadcastwould block indefinitely with no log output.Now pops unconditionally by index — safe because
broadcastMessagesisthe only consumer.
removeFilesafety: catchesisDoesNotExistErrorin pop so a missingfile (e.g. from a prior partial crash) doesn't kill the broadcast loop.
PersistentQueueFullbefore blocking ona full queue, and
PersistentQueueLoadFailedinstead of silently discardinga startup load error.
Eq msgdropped fromwithEtcdNetworkandbroadcastMessages.