I wanted a job scheduler that still has an audit trail, but the usual shape falls over once several workers hit the same table. A broker scales, and then the retries, the schedules, and the history are gone.
I came up with a bucket in front of the database. Every job is written into a per-worker bucket first, so scheduling does not wait on the audit table. The Master DB is only the long-term record. A job due soon stays in the bucket. A job due next month is parked in the Master DB and moved back in a batch when it is close, instead of one row at a time. Each worker owns its own buckets, so two workers never grab the same job. Failures go back to the Master DB and are sent out again up to the retry limit.
On a 50k burst with 20 workers that did 23.8k jobs/sec with NATS in front of RavenDB, against 6.5k when one database did both.
Does that split make sense, or is the Master DB still going to be the bottleneck? The experiment is here: https://github.com/hugoj0s3/jobmaster-net
https://jobmaster-doc.hugo86jose.workers.dev/docs/architecture-under-the-hood/architecture-overview
Replies
Know the answer? Post it — somebody with the same question will find it here.