首页 > AI前沿 > Deterministic Simulation Testing in Celld

Deterministic Simulation Testing in Celld

Hacker News 2026-10-10 02:31 4 阅读 查看原文
Deterministic simulation testing in celld: exploring execution orders and reproducing bugs Yusuke Tanaka·October 2, 2026 We are building celld, a runtime for Cloudflare Workers and Durable Objects applications on your own machines. It can run with an S3-compatible object store as its only external service dependency. celld is a distributed system. Making distributed systems reliable is hard, partly because we can unintentionally rely on assumptions that don’t hold in practice, even when we know the pitfalls. Peter Deutsch’s The Eight Fallacies of Distributed Computing lists eight such assumptions, including “The network is reliable” and “Latency is zero.” A bug can depend on a particular sequence of delayed messages, failed writes, and node restarts. Those events can happen in a different order on the next test run, making the failure difficult to reproduce. We need to be able to repeat the failing run so we can investigate the cause and check whether a proposed fix really resolves the problem. That is why we use deterministic simulation testing (DST). Our simulator is still under development and is not included in celld’s public repository, but it has already found previously unknown bugs. In this article, we’ll walk through how DST works in celld and how it helped us find, reproduce, and fix one of those bugs. How DST works in celld DST runs celld’s production code in an environment controlled by a simulator. A cell in celld runs application code and has its own SQLite database. As cells do their work, celld handles events such as incoming requests, completed storage operations, and timer firings. The code that selects the next event is separate from the code that handles it. This lets the simulator control the order of events while running the same event-handling code as in production. The simulator also controls when asynchronous tasks run and how object storage responds. It can delay a write or make it fail. It can advance simulated time without waiting for real time to pass. We use this control to test how time-dependent behavior, such as retries and alarms, interacts with other events. We configure which requests and faults the simulator can explore. Within those limits, it uses random choices to generate requests, select what runs next, decide whether storage operations succeed or fail, and choose which faults occur and when. A seed is the starting value for its random-number generator. With the same code, settings, and seed, we get the same sequence, intermediate states, and result. We can explore different runs by changing the seed, then repeat a failing run to trace where things went wrong. After each action, a checker uses observed responses and stored data to test invariants: conditions we define that the system must uphold throughout a run. The test initializes celld with the simulated environment, then explores actions up to a configured limit. In pseudocode: const simulation = createSimulation({ settings, seed }); for (let step = 0; step < settings.maxActions; step++) { const candidates = simulation.availableActions(); if (candidates.length === 0) { checkCompletion(simulation); break; } const action = simulation.choose(candidates); simulation.run(action); checkInvariants(simulation); } Here, an action is something the simulator can advance: delivering a request, letting a task run, advancing the simulated clock, or completing a storage operation. If no actions are available, the test checks its completion conditions, such as whether all planned requests have finished. The simplified replay below starts after Cell A has inserted a row into SQLite. Watch how the data is saved to object storage before celld receives confirmation. Other requests or background tasks can run in that gap, while celld is still waiting to return success to the client. Bugs can emerge when these operations interleave in unexpected ways. The simulator controls these steps separately so we can explore different execution orders and replay those that expose a bug. Press Next to step through choosing an action, animating it, and revealing its result. Send the SQLite changes; let object storage receive them. 1Choose 2Run 3Result The row is in SQLite. The client is still waiting for success. Enable JavaScript to step through the recording. The complete sequence is also available as text. This excerpt starts after Cell A has inserted a row into SQLite. It follows the changes to object storage, then the write result back to celld and the success response to the client. The seed determines which available action the simulator selects next, including request generation and delivery. Only the candidates relevant to this illustration are shown. Clicking Next advances through choosing an action, running it, and showing its result. These are presentation phases, not simulator ticks. The ▶ marker identifies the components selected to advance; ◆ marks the resulting changes. celld handles the events with its production code. The code that inserts the row is written for the test. The client and object storage also use test implementations. A simplified excerpt from a recorded development run. Cell A is the cell handling this request; the other cells illustrate that each cell has its own SQLite database. Concurrent work and intermediate actions are omitted. Upload the database changes Chosen: Send the SQLite changes; let object storage receive them. Candidates: Send the SQLite changes; let object storage receive them. celld sends a batch of SQLite changes to object storage. The simulated object store receives the data, but the write has not completed yet. Chosen: Send the SQLite changes; let object storage receive them. Candidates: Send the SQLite changes; let object storage receive them. celld sends a batch of SQLite changes to object storage. The simulated object store receives the data, but the write has not completed yet. Complete the write to object storage Chosen: Complete the storage write. Candidates: Complete the storage write. The simulator lets the write take effect. The data is saved, but the success result has not yet been delivered to celld. Chosen: Complete the storage write. Candidates: Complete the storage write. The simulator lets the write take effect. The data is saved, but the success result has not yet been delivered to celld. Deliver the result and resume waiting work Chosen: Deliver the storage write result to celld. Candidates: Deliver the storage write result to celld. After the storage result arrives, celld resumes and confirms that the database changes are durable. Chosen: Deliver the storage write result to celld. Candidates: Deliver the storage write result to celld. After the storage result arrives, celld resumes and confirms that the database changes are durable. Return a successful response Chosen: Let celld check again whether it can reply. Candidates: Let celld check again whether it can reply. After celld completes its checks, it releases the response. The client receives a success response for the row insertion. Chosen: Let celld check again whether it can reply. Candidates: Let celld check again whether it can reply. After celld completes its checks, it releases the response. The client receives a success response for the row insertion. The replay above shows a successful request. Next, we’ll look at an execution order that exposed a bug in celld. An alarm race the simulator found An alarm schedules a cell to run application code at a specified time. The cell need not stay in memory until then: celld can unload an idle cell to make room for others and load it again in time to run its alarm. The alarm's scheduled time is stored in the cell's SQLite database. If the cell has been unloaded, finding its alarm from SQLite alone would mean opening its database again. Doing that for every cell would be expensive, so celld instead scans wake entries in object storage. Each entry identifies a cell and when to reactivate it. Once the cell is active, celld reads the scheduled time from SQLite to determine when to run the alarm. When we found this bug, celld reused wake entries to reduce the number of writes to object storage. For example, an application could delete a 10:00 alarm and then set a new one for 10:05. If the 10:00 wake entry was still present, celld could reuse it to reactivate the cell at 10:00, early enough for the new alarm. Once the cell was active, celld would read 10:05 from SQLite and wait until then to run the alarm.1 Deleting the 10:00 alarm clears its scheduled time from SQLite. Once that change has been saved to object storage, celld returns success to the client. A separate cleanup task removes the 10:00 wake entry later, so the client does not have to wait for that extra storage operation. The client can therefore set the 10:05 alarm while the old cleanup is still waiting to run. After celld confirms that the 10:05 alarm is set, a wake entry must remain that can reactivate the cell by 10:05. The simulator checks that this remains true until the alarm runs or is canceled. The figure below shows how this invariant is violated when the old cleanup runs after the 10:05 alarm has been set. Delete Cell A's 10:00 alarm. 1Choose 2Run 3Result 4Check Cell A has a 10:00 alarm in SQLite and a 10:00 wake entry in object storage. Enable JavaScript to step through the recording. The complete sequence is also available as text. This is an illustrated reconstruction of a race found during development, rather than an exact trace replay. The simulator selects which requests and pending operations advance. Only the candidates relevant to this illustration are shown. Clicking Next advances through choosing an action, running it, showing its result, and checking the invariant when applicable. These are presentation phases, not simulator ticks. The ▶ marker identifies the selected component; ◆ marks the resulting changes. At the key interleaving, two candidates show which operation advances and which waits. celld’s code decides how to process the events. The test checks the wake information after an alarm is set successfully. This check is fixed, not chosen by the seed. The sequence describes the implementation at discovery. Each displayed step groups several lower-level actions. All states and the check result are prepared for playback; the browser does not run celld or its simulator. Delete the 10:00 alarm Chosen: Delete Cell A's 10:00 alarm. Candidates: Delete Cell A's 10:00 alarm. The alarm is gone from SQLite. The 10:00 wake entry remains in storage; its deletion task has not run yet. Chosen: Delete Cell A's 10:00 alarm. Candidates: Delete Cell A's 10:00 alarm. The alarm is gone from SQLite. The 10:00 wake entry remains in storage; its deletion task has not run yet. Acknowledge the deletion Chosen: Let celld reply to the deletion request. Candidates: Let celld reply to the deletion request. Run cleanup of the 10:00 wake entry. The client receives a success response for deleting the alarm. Cleanup of the 10:00 wake entry still waits. Chosen: Let celld reply to the deletion request. Candidates: Let celld reply to the deletion request. Run cleanup of the 10:00 wake entry. The client receives a success response for deleting the alarm. Cleanup of the 10:00 wake entry still waits. Set a 10:05 alarm first Chosen: Set Cell A's alarm for 10:05. Candidates: Set Cell A's alarm for 10:05. Run cleanup of the 10:00 wake entry. The 10:05 alarm reuses the 10:00 wake entry: the cell can wake early and wait until 10:05. Chosen: Set Cell A's alarm for 10:05. Candidates: Set Cell A's alarm for 10:05. Run cleanup of the 10:00 wake entry. The 10:05 alarm reuses the 10:00 wake entry: the cell can wake early and wait until 10:05. Acknowledge the new setting Chosen: Let celld reply to the 10:05 alarm request. Candidates: Let celld reply to the 10:05 alarm request. Run cleanup of the 10:00 wake entry. Setting the 10:05 alarm succeeds. The reused 10:00 wake entry must remain available. Chosen: Let celld reply to the 10:05 alarm request. Candidates: Let celld reply to the 10:05 alarm request. Run cleanup of the 10:00 wake entry. Setting the 10:05 alarm succeeds. The reused 10:00 wake entry must remain available. Run the delayed cleanup Chosen: Run cleanup of the 10:00 wake entry. Candidates: Run cleanup of the 10:00 wake entry. Acting on the earlier deletion, the delayed cleanup task requests deletion of the 10:00 wake entry. Chosen: Run cleanup of the 10:00 wake entry. Candidates: Run cleanup of the 10:00 wake entry. Acting on the earlier deletion, the delayed cleanup task requests deletion of the 10:00 wake entry. The check catches lost wake information Chosen: Let object storage apply the deletion. Candidates: Let object storage apply the deletion. The 10:00 wake entry is gone, even though the 10:05 alarm was set successfully. Chosen: Let object storage apply the deletion. Candidates: Let object storage apply the deletion. The 10:00 wake entry is gone, even though the 10:05 alarm was set successfully. In this run, both deleting the 10:00 alarm and setting the new 10:05 alarm returned success while the old cleanup waited. When the simulator finally advanced that cleanup, it removed the 10:00 wake entry that the new alarm still needed. The checker reported an invariant violation: the 10:05 alarm was still set in SQLite, but no wake entry remained to reactivate the cell. If the node restarted before 10:05, the missing wake entry could prevent the alarm from running, despite the earlier success response. The simulator found this failing order by exploring requests and pending work. We had not written a test that prescribed this sequence. Reproducing and fixing the race We replayed the failure with the same code, settings, and seed. Each run let us inspect the moment the old cleanup deleted the needed wake entry, without searching for the failing order again. Since then, we have changed the wake-entry design. Each new alarm setting gets its own entry in object storage. A delayed deletion for an older alarm can only remove that older entry, leaving the new one intact. Old entries are removed once celld has confirmed they are no longer needed. We turned the failing run into a regression test that repeats the requests and cleanup order shown above. It checks that the new alarm keeps a usable wake entry. The regression test checks the invariant during the run, not just at the end. If the wake entry disappeared and was later restored, checking only the final state would miss the violation. What's next The simulator is still under development, but it has already helped us find and fix previously unknown bugs in celld. Next, we will test more combinations of requests, storage failures, and node restarts. For example, a node could restart during a storage outage and receive new requests before storage recovers. We will also add checks for more of celld’s guarantees. A write reported as successful, for example, must survive a node restart. When these tests find a bug, we can replay the failure, fix it, and keep the sequence as a regression test. If celld needed to evict the cell while waiting for 10:05, it would first ensure that a wake entry for 10:05 was in place. ↩