30-SECOND SUMMARY
What to take away
- The starting question was who owned a job at any moment. If reading a waiting row and marking it running are separate operations, two workers can select the same job; if a worker disappears, a bare running flag cannot explain whether another worker may retry. We treated duplicate claims, stuck running state, and late completion as failures the starter must reproduce and prevent, not as measured incidents across a finished fleet.
- Candidate selection, eligibility checks, lease creation, and the transition from `waiting` to `running` occur inside one `BEGIN IMMEDIATE` transaction. Exceptions roll the transaction back, and model calls stay outside this short write section so they do not hold the database lock. Initialization is followed by `PRAGMA quick_check`; a copy created with SQLite `.backup` is then opened and checked separately as a manual recovery exercise.
- SQLite fits the current assumption of one coordinator and a modest write rate because ownership rules remain inspectable and easy to back up. The accepted result body still lives directly in `result_json`, so large artifacts need external storage, hashes, lifecycle state, and coordinated backup. Multiple coordinators, high write contention, runtime adapters, and repeated real-device recovery measurements remain outside the verified boundary.
The operating problem
The starting question was who owned a job at any moment. If reading a waiting row and marking it running are separate operations, two workers can select the same job; if a worker disappears, a bare running flag cannot explain whether another worker may retry. We treated duplicate claims, stuck running state, and late completion as failures the starter must reproduce and prevent, not as measured incidents across a finished fleet.
The working rule for “The operating problem” is: The starting question was who owned a job at any moment. If reading a waiting row and marking it running are separate operations, two workers can select the same job; if a worker disappears, a bare running flag cannot explain whether another worker may retry. We treated duplicate claims, stuck running state, and late completion as failures the starter must reproduce and prevent, not as measured incidents across a finished fleet. Preserve timestamped state transitions, reason codes, and retry outcomes so that an operational decision can be reconstructed later.
For verification, save the model and runtime versions, source input, relevant settings, and observed output together. Repeat the step while changing only one factor, and record unexpected results and untested limits as carefully as successes before applying the guidance to private or production data.
The design decision
The starter separates workers, jobs, and leases so each failure has a specific place to inspect. Workers preserve the last heartbeat, jobs preserve input, priority, status, and the accepted `result_json`, while leases preserve the current owner and expiry. Artifact hashes, external result storage, and validation stages belong to the expansion design and are not tables already shipped in the basic schema.
The working rule for “The design decision” is: Candidate selection, eligibility checks, lease creation, and the transition from `waiting` to `running` occur inside one `BEGIN IMMEDIATE` transaction. Exceptions roll the transaction back, and model calls stay outside this short write section so they do not hold the database lock. Initialization is followed by `PRAGMA quick_check`; a copy created with SQLite `.backup` is then opened and checked separately as a manual recovery exercise. Preserve timestamped state transitions, reason codes, and retry outcomes so that an operational decision can be reconstructed later.
Preserve the before-and-after state and the time of the check so that another run can reproduce the result. Include at least one failure condition—such as empty input, constrained resources, or a restart—to reveal the boundary of the step rather than documenting only the happy path.
How the mechanism works
Candidate selection, eligibility checks, lease creation, and the transition from `waiting` to `running` occur inside one `BEGIN IMMEDIATE` transaction. Exceptions roll the transaction back, and model calls stay outside this short write section so they do not hold the database lock. Initialization is followed by `PRAGMA quick_check`; a copy created with SQLite `.backup` is then opened and checked separately as a manual recovery exercise.
The working rule for “How the mechanism works” is: SQLite fits the current assumption of one coordinator and a modest write rate because ownership rules remain inspectable and easy to back up. The accepted result body still lives directly in `result_json`, so large artifacts need external storage, hashes, lifecycle state, and coordinated backup. Multiple coordinators, high write contention, runtime adapters, and repeated real-device recovery measurements remain outside the verified boundary. Preserve timestamped state transitions, reason codes, and retry outcomes so that an operational decision can be reconstructed later.
Define completion with an observable result instead of a general impression. Repeat the same input, and if the output changes, isolate whether the model, runtime settings, or source data changed before moving to the next stage.
What we verified
Five tests cover P0 priority, rejection at 90% RAM against the 85% policy gate, recovery after a 45-second lease expires, duplicate completion, and completion submitted after expiry. In the fixed-time recovery case, a lease granted at t=102 expires at t=147, so completion at t=148 must fail before the reaper returns the job to waiting. These tests verify coordinator rules, not operating-system sensors, model-process cancellation, or fleet throughput.
The working rule for “What we verified” is: The starting question was who owned a job at any moment. If reading a waiting row and marking it running are separate operations, two workers can select the same job; if a worker disappears, a bare running flag cannot explain whether another worker may retry. We treated duplicate claims, stuck running state, and late completion as failures the starter must reproduce and prevent, not as measured incidents across a finished fleet. Preserve timestamped state transitions, reason codes, and retry outcomes so that an operational decision can be reconstructed later.
For verification, save the model and runtime versions, source input, relevant settings, and observed output together. Repeat the step while changing only one factor, and record unexpected results and untested limits as carefully as successes before applying the guidance to private or production data.
Boundary and completion criteria
SQLite fits the current assumption of one coordinator and a modest write rate because ownership rules remain inspectable and easy to back up. The accepted result body still lives directly in `result_json`, so large artifacts need external storage, hashes, lifecycle state, and coordinated backup. Multiple coordinators, high write contention, runtime adapters, and repeated real-device recovery measurements remain outside the verified boundary.
The working rule for “Boundary and completion criteria” is: Candidate selection, eligibility checks, lease creation, and the transition from `waiting` to `running` occur inside one `BEGIN IMMEDIATE` transaction. Exceptions roll the transaction back, and model calls stay outside this short write section so they do not hold the database lock. Initialization is followed by `PRAGMA quick_check`; a copy created with SQLite `.backup` is then opened and checked separately as a manual recovery exercise. Preserve timestamped state transitions, reason codes, and retry outcomes so that an operational decision can be reconstructed later.
Preserve the before-and-after state and the time of the check so that another run can reproduce the result. Include at least one failure condition—such as empty input, constrained resources, or a restart—to reveal the boundary of the step rather than documenting only the happy path.
Frequently asked questions
Is this a universal finished cluster product?
No. The coordinator rules are tested, while runtime, sensor, and channel adapters depend on the actual environment. For a practical check, follow the “The operating problem” section, change one condition at a time, and record the result.
Why publish the unfinished boundary?
It distinguishes reproduced behavior from planned integration and makes the field note auditable. Candidate selection, eligibility checks, lease creation, and the transition from `waiting` to `running` occur inside one `BEGIN IMMEDIATE` transaction. Exceptions roll the transaction back, and model calls stay outside this short write section so they do not hold the database lock. Initialization is followed by `PRAGMA quick_check`; a copy created with SQLite `.backup` is then opened and checked separately as a manual recovery exercise. For a practical check, follow the “The design decision” section, change one condition at a time, and record the result.
Primary sources
Check the original documentation for version-specific details.
SQLite Transactions Ollama API