Previously, it was possible for the spool in tasks to end up
being spawned on just a subset of the available spool in threads
if there was some idleness in the task processing startup, for
example, if the dns for the first few messages returned from
spool enumeration is slow to resolve.
In that situation we can end up with no effective concurrency
during spool enumeration, leading to a very slow startup.
What I'd like to see to resolve this wholistically is adopting
the main tokio work stealing task runner, but we are prevented
from doing this until mlua 0.10 is released.
What this commit does is refactor the core of the Runtime
code to extract the function that sets up the thread pool so
that we can directly spawn the spool in thread logic into
each of the worker threads, guaranteeing that they are spread
out one to a thread.
With this change in place, I always observe 100% utilization
of spoolin on startup where I previously would see only around
60 or 70%.