There are a lot of metrics these days, we need to scroll through them!
Use the arrow keys, page up/down and home/end for this purpose.
closes: https://github.com/KumoCorp/kumomta/issues/372
This was probably the casualty of some earlier refactoring
that has gone unnoticed until now.
Since we don't have explicit context on which key format to
parse in this helper function that is used in multiple places,
let's just make it try to parse both rsa and ed25519.
closes: #368
Occasionally we'll have someone report that systemd timed out
and sigkill'd their kumo on shutdown.
One possible scenario for this is a lua delivery handler that
is taking too long, presumably because the other end of it
(eg: webhook or other custom endpoint) is not responding in
a timely fashion.
The way that we handle shutdown is that we compute a maximum
theoretical timeout value by summing up all of the smtp client
timeout values. Some of those can be several minutes in
duration because the are using default values derived from
a very conservative set of values suggested by the SMTP
RFCs from the '70s.
Those obviously should not apply to a custom delivery handler,
but also, in the context of an established SMTP session, we
should not add in the connection-establishment-specific values
when we're just waiting for a per-message send.
This commit addresses this situation on two fronts:
* Introduce a new system_shutdown_timeout value that allows the
user to conveniently express their desired timeout value
in a single option. This is *not* set by default!
* The default value for system_shutdown_timeout is computed by
summing the per-message-delivery smtp timeout options, which
is a much more reasonable, and more importantly, shorter than
our 300s TimeoutStopSec value in kumomta.service
Previously, we'd use the qmaint pool to spawn both the scheduled
and ready queue maintenance tasks.
This commit splits them apart in order to avoid the potential for
cross-task contention on the same threads if a scheduled queue
and ready queue pair are communicating with each other.
We recently observed a system where tsa-daemon run out of file
descriptors because it hit a systemd default of 1024 on the host system,
which had 192 cores. We spawn a thread per core for internal processing
in tsa-daemon, and that has associated with it a number of kernel
objects that each have an fd.
1024 is an unreasonably small limit for the number of fds, so let's
just try to raise to match the normally much higher hard limit, just as
we do in kumod.
We recently observed a system running on an over-committed VM that
reported 4x the actually available parallelism.
Since we scale our thread pool sizes from this value, it resulted
in an extra-over-committed configuration for kumod.
You may now set KUMO_AVAILABLE_PARALLELISM in the environment to
override the value that we see both interally and expose via
the lua function with the same name.
There are a few breaking changes in the API that we need to tackle
here, but they're generally fine.
Since our libunbound crate uses hickory-proto, I had to upgrade
that dep over there and reference the updated git rev here in
this commit; we don't publish that to crates.io.
The main thing that stood out in the upgrade is that the zone
file parser used by the TestResolver now seems to create
records with varying FQDN-ness. This may just be that it is
now passing data through from the underlying zone data, but
it caused a number of the SPF and DKIM test cases to fail
without canonicalizing the names to FQDN. I chose to do that
in the TestResolver rather than reviewing all the input zone
data, with the rationale being that it is least surprising
to have the test resolver fix that up than to puzzle over
records not resolving due to a missing trailing dot when
more tests are written in the future.
This commit doesn't try to take advantage of the improvements
to DNSSEC that are available in this version of hickory,
it's just upgrading to the API changes.
closes: https://github.com/KumoCorp/kumomta/pull/361
This allows pre-defining connection metadata values. When coupled with
`peer` and/or `via`, these can be done based on the corresponding
addresses associated with the session.
closes: https://github.com/KumoCorp/kumomta/issues/355
The motivation here is to remove tls_config from EsmtpListenerParams
to make some future configuration changes easier, so this commit
moves that simple cache out to an explicit lru ttl cache.
This has the welcome side effect of enabling periodic reloading
of the tls parameters, which in turn makes it a hands-off process
for updating certificates: we no longer require the service to
be restarted for that.
These are hooked up only for memoize at this time. No default
behavior is changed by this commit, but you can optionally
specify these parameters in order to change the behavior.
The introduction of the
`opportunistic_tls_reconnect_on_failed_handshake` option resulted in
this regression, which is because I misread the `match` statement
for this case as being only for the opportunistic case, but it
also encompasses the required case.
The issue is:
* A site has an MTA-STS policy enforcing Required tls
* The handshake with that site fails (for reasons unknown and
irrelevant)
* We would unconditionally (wrt. Required vs. Opportunistic) respect
opportunistic_tls_reconnect_on_failed_handshake and re-queue the
current address for the next connection attempt
* Ordinarily, opportunistic_tls_reconnect_on_failed_handshake +
the remembered broken state would cause that next attempt to
downgrade to clear text, but MTA-STS forces the policy to
Require
* Goto step 2 (modulated by connection rate throttling)
The fix is simply to only apply
opportunistic_tls_reconnect_on_failed_handshake when the policy
is actually opportunistic.
We were using a fairly tight limit of 16 messages in the channel
that buffers the effects of changing bounces/suspensions from
any websocket-connected-clients.
A busy server could hit that limit fairly easily, resulting
in a `channel lagged by NUMBER` error that drops the websocket,
causing the client to need to reconnect and resync.
This commit resolves that by making the buffer a much more healthy size.
We were deduping just by rule_hash, but each of these tables has
additional required fields as part of the primary key.
The result was that, for sites with a lot of bounces/suspensions
triggered by the same rules across a related set of sources,
the full set of bounces and suspensions would not be correctly
reported as part of a websocket push.
sqlite doesn't have a native async interface, and instead will
use traditional OS-level mutexes to ensure thread safety.
Using those when under contention in a tokio scheduler thread
can lead to blocking of the tokio scheduler threads, which can
prevent timely delivery of data via websockets, or timely
processing of incoming log records.
This commit fixes up the sqlite access points to use tokio's
spawn_blocking function to move that style of mutex acquisition to a
more suitable context.
We'll wait up to 3s at a time for however many mesages are available
to extract from the tsa daemon websocket, then process the results
in batches.
This avoids the potential for geometric complexity if there is a run of
subscription updates happening around the same time.
I'm not totally sure why this isn't universally broken when using
openssl (instead of rustls), but in the specific case we were
investigating, the destination was configured via a routing_domain
and the resulting mx_host name had the trailing FQDN dot on it.
Removing that dot allows the certificate to verify, so let's
ensure that we strip it here in the client.
The issue here is that when an rfc2047 encoded display name is split
across multiple lines, the whitespace between them is not recognized
at the right time, which results in the second encoded word being
passed through as-is, without being decoded.
This commit fixes the precedence of whitespace parsing in that
case.
This commit allows setting a per-message `expires` timestamp
via msg:set_scheduling (and thus msg:import_scheduling_header).
The expiration takes precedence over max_age; max_age will be
ignored for messages that have configured and expiration time.
The expiration time is independent of the other scheduling
restrictions.
This resolves an issue where the default behavior for serde is to
silently swallow issues with this struct, because we use a flattened
optional structure for those restrictions.
Previously, we'd pick a source whether it had room for the new
message or not, then generate a TransientFailure when we subsequent
figure out that it is full.
This commit will try to deliver through one of the other possible
sources instead of delaying the message.
While auditing Answer::as_txt usage as a follow up from the recent
SPF fix, I noticed a TODO in the dkim code (which we forked from
another implementation) to support processing multiple TXT
records.
This commit implements the necessary tweaks to extract multiple
signatures and attempt to verify them against the incoming message.
The issue here was essentially a data fidelity issue around
TXT record representation.
A DNS TXT record can be composed from multiple strings, and a domain can
return multiple TXT records, so there is some nesting.
The SPF RFC says:
```
3.3. Multiple Strings in a Single DNS Record
As defined in [RFC1035], Sections 3.3 and 3.3.14, a single text DNS
record can be composed of more than one string. If a published
record contains multiple character-strings, then the record MUST be
treated as if those strings are concatenated together without adding
spaces. For example:
IN TXT "v=spf1 .... first" "second string..."
is equivalent to:
IN TXT "v=spf1 .... firstsecond string..."
TXT records containing multiple strings are useful in constructing
records that would exceed the 255-octet maximum length of a
character-string within a single TXT record.
```
so the SPF logic was dutifully joining records together around the
empty string.
Howerver, if you look at `dig yahoo.com txt` you'll see a bunch
of non-SPF records:
```
yahoo.com. 1800 IN TXT "google-site-verification=Z3-Vh6zqUMgybVH4wQl1GxKSKN7JE13kyCyeZ3TZZ-I"
yahoo.com. 1800 IN TXT "v=spf1 redirect=_spf.mail.yahoo.com"
yahoo.com. 1800 IN TXT "Zoom=13284637"
yahoo.com. 1800 IN TXT "edb3bff2c0d64622a9b2250438277a59"
```
these were getting joined together and producing a bogus input.
Obviously we should not join the results from the txt lookup
together like that, but then why would the RFC make a point
of talking about joining stuff together, and where should
that logic live?
Our `Answer::as_txt` implementation was doing some joining
of its own and it turned out that it was concatenating across
the outer layer of the aforementioned TXT record nesting.
This commit fixes that up and adds some test coverage.