The prior commit focused on avoiding a panic when we encountered an invalid
boundary string. This commit makes thing degrade more gracefully: we'll
treat the bad MIME part in the same way that we deal with overly deep nesting
and handle it as an opqaque leaf part.
The new_multipart constructor now rejects an invalid boundary rather than
building an unserializable part, so a flawed policy can't accidentally create
broken messages.
to_message_bytes now propagates errors instead of panicking. check_fix_conformance
and the HTTP injection API now report this error directly, and DSN generation
omits the original message when it cannot be serialized.
/0 prefixes now correctly match all addresses. Invalid overlong prefix lengths
are rejected at parse time instead of silently producing wrong masks (or, in
debug builds, a panic).
A malformed or forged `id` in an xfer payload could panic kumod, since
spool ids are assumed to always be v1 UUIDs with an embedded
timestamp. Every construction path now rejects anything that isn't
v1.
Deeply nested multipart content is now retained as an opaque part with
its headers parsed and raw body preserved once nesting passes 100
levels, instead of recursing until the process stack overflows.
Those messages are tagged with the new MIME_NESTING_LIMIT_EXCEEDED
conformance flag.
The RFC 5322 comment parser recursed once per nesting level, so a header
with thousands of nested parentheses could exhaust the stack and crash
the process. It now tracks nesting depth with an explicit loop instead.
When `invalid_line_endings` is set to Fix or Allow, correctly measure line
lengths even when messages use bare LF or CR instead of CRLF. Previously
such messages were incorrectly rejected as having overlong lines.
In most cases this error context was discarded, but if your policy
script triggered an explicit parse, you might see the error context in
the lua error that it would trigger in that case. Since it was not very
useful (and was wasteful) we now simply report the input message size
for the context instead. The actual parse error is still the primary
reason in the error chain, so this is not a loss in information.
Bounds memory amplification from crafted messages with many minimal short
headers, which previously consumed far more resident memory than their
wire size would suggest.
Fix O(N^2) CPU usage when rebuilding messages with many MIME parameters.
Decoding the parameter map, re-emitting the header, and merging original
parameters back into the rebuilt header were each quadratic. Duplicate
parameters now use consistent last-wins semantics.
Avoid quadratic work for messages with many malformed line endings when using
`invalid_line_endings="Fix"`. Remove the in-place implementation entirely to
prevent future accidental use.
Strip control characters from all values when encoding Authentication-Results
and ARC headers. This prevents sender-controlled values from embedding raw
CR/LF to forge headers or push trusted results into the message body.
closes: https://github.com/KumoCorp/kumomta/pull/523
Avoid process abort when kumo.fs.glob is passed an absolute pattern
containing a recursive wildcard. This was an issue in the upstream
filenamegen crate, resolved in the latest release.
closes: https://github.com/KumoCorp/kumomta/issues/578
This fixes two vulnerabilities reachable via msg:dkim_verify() on inbound
mail: an out-of-bounds read panic on very short b= tags, and the anti-DoS
cap now correctly bounds all parsed signatures, not just parse failures.
Address headers in the HTTP inject API's generic `headers` map (`Cc`, `To`,
etc.) were emitted as unstructured text. When they contained non-ASCII
content, the entire value (as opposed to just the display name) would
be qp-encoded, producing an invalid header.
closes: https://github.com/KumoCorp/kumomta/pull/598
normalize() previously encoded the from/reply_to display name into a
Mailbox before the template engine ran. That allowed a substituted value to
splice a second mailbox into the header, since the templated name was
concatenated into an already-encoded string rather than escaped as a unit.
The raw name is now templated first and encoded afterward, over the
assembled Mailbox.
refs: https://github.com/KumoCorp/kumomta/pull/598
quote_string passed a CR or LF straight through into a quoted display name.
that would terminate the header line early, causing the subsequent bytes
to appear as a separate injected header.
This has been resolved, taking care to preserve legitimate folding.
Mailbox::encode_value now folds between the display name and `<addr>`,
treating both as atomic, instead of leaving callers to whitespace-wrap
the whole encoded mailbox. That whitespace wrap could fold inside a
quoted display name and corrupt it.
Route the Lua `header.value` field getter through
`Header::structured()`, eliminating the duplicate name-to-parser
`NAME_GETTER` table.
`Date` continues to return the raw header string for backwards
compatibility.
Headers were built and rebuilt as "Mime-Version" rather than
"MIME-Version". Both are valid per RFC 2045, but some spam filters,
such as rspamd, score the mixed-case form as a deliverability signal.
Fix the header name in the accessor macro and in ParsedHeader's
grammar table, the places that spell it, and update the snapshot
tests across mailparsing, message, kumo-log-types, mod-mimepart, and
kumod that captured the old spelling, along with the reference docs
that showed the old casing in their examples. #564
Add `ParsedHeader`, which parses a header's raw bytes using the grammar
implied by its name (mailbox list, address list, date, MIME parameters,
and so on) and re-encodes it back to canonical form. `Header::rebuild()`
now goes through this lookup instead of a hand-written table, which
also fixes `Authentication-Results`: it was falling through the old
table's unstructured fallback and being left as free text instead of
being parsed and re-encoded like the other structured headers.
wrap() hard-wrapped an over-long word byte by byte, which could slice a
multi-byte UTF-8 sequence in half and panic when the result was
validated as UTF-8. Walk the word in UTF-8 chunks instead, splitting
only between whole characters and passing invalid byte runs through
unchanged.
When a record exceeds the configured max_line_size, discard just that record
and continue reading the rest of the segment instead of aborting the entire
segment.
The zstd decompression buffer now grows on demand to read records up to
a configurable 128 MiB cap. Oversized records are rejected explicitly
both on read and write, with matching limits so all written records are
guaranteed readable.
The gate could latch permanently when only one of the two error signals
(foreground or background) ever moved, requiring an operator restart. It now
auto-reopens for a retry after `error_unlatch_duration`, regardless of which
signal latched it.
Also serializes foreground error reporting with gate transitions so a
concurrent fatal error cannot be undone by a reopen, and rejects a zero
`error_unlatch_duration` when automatic reopening is enabled. The underlying
cause is always logged. Adds a reusable LD_PRELOAD fault-injection crate and
integration tests that reproduce a full disk and a slow disk to drive the
latch-and-reopen cycle end to end.
Fixes#597
A literal underscore in a header value that needed RFC 2047 Q encoding was
emitted unescaped, so a conforming mail client decoded it back to a space
and silently corrupted the header. Underscores are now escaped as =5F, and
the Q encoder passes through only the punctuation RFC 2047 permits unencoded
in a phrase.
Iterating a response's headers with pairs() looped forever when a header name
repeated, which could happen for example with multiple Set-Cookie headers.
We have mailparsing wrappers that format and parse rfc2822 dates without
panicking on out-of-range values or rejecting the obsolete timezones that
real senders emit. This moves every existing caller onto those wrappers and
adds a clippy lint that requires them in place of chrono's versions.