mirror of
https://github.com/mailscope/kumomta.git
synced 2026-09-07 02:58:56 +00:00
We had a user report problems with an incoming OOB message. The issue was that `msg:from_header()` would raise an error because one of the MIME parts in the incoming message had 8-bit data and didn't apply transfer encoding on the offending part. It's actually a bit deeper than just missing transfer encoding; the issue was really that Spam Assassin was employed on the remote system and it generated an `X-Ham-Report` header that embedded 8-bit data (looks mostly UTF-8, but had an invalid byte) in that header without applying RFC2047 header encoding. ``` X-Ham-Report: Spam detection software, running on the system "example.com", has NOT identified this incoming email as spam. The original message has been attached to this so you can view it or label similar future email. If you have any questions, see root\@localhost for details. Content preview: <unencoded binary data here> ``` In the resultant rfc3464 delivery status report, the original message payload is included as a message/rfc822 part in the body. So to our parser this looks like a badly encoded body (which it is), but only because the header in that part was badly encoded. So this is a double-fail; the Spam Assassin header content preview logic is generating a non-conforming header, and Exim's DSN generation at the offending site is generating a non-conforming message payload. What this commit does is: * Adds an `IntoSharedString` trait to encapsulate the conversion from String, str or bytes into SharedString. Previously we used a fallible conversion for this, but now this conversion is infallible but returns a MessageConformance value. Internally, if the data is not UTF-8, we fall back to the lossy conversion and flag the part as needing transfer encoding. * That will allow the parser to return something, even if it is slightly mangled by the unicode replacement character. This mangling is not a bug: it's a case of "garbage-in, garbage-out". * This change allows check_fix_conformance to report NEEDS_TRANSFER_ENCODING when run in check mode. When run in fix mode, the offending part, *including the replacement character* will have transfer encoding applied to it. We can't "do better" here because the input message is bogus and is missing proper transfer encoding. * Fixes an issue where whitespace from this fixed part was stripped out. I'm not sure why whitespace was being stripped; no unit tests fail as a result of this change, so it must be good?