Opened 70 minutes ago
#37413 new Bug
MultiPartParser treats boundary tokens not preceded by CRLF as delimiters
| Reported by: | Natalia Bidart | Owned by: | |
|---|---|---|---|
| Component: | HTTP handling | Version: | dev |
| Severity: | Normal | Keywords: | MultiPartParser boundary |
| Cc: | Triage Stage: | Unreviewed | |
| Has patch: | no | Needs documentation: | no |
| Needs tests: | no | Patch needs improvement: | no |
| Easy pickings: | no | UI/UX: | no |
Description
BoundaryIter._find_boundary() in django/http/multipartparser.py looks for the bare --<boundary> string and only strips a preceding CRLF if present. Per RFC 2046 section 5.1.1, a delimiter is CRLF "--" boundary, so a boundary token appearing inside part data without a leading CRLF should be treated as content, not as the end of the part:
The boundary delimiter MUST occur at the beginning of a line, i.e., following a CRLF, and the initial CRLF is considered to be attached to the boundary delimiter line rather than part of the preceding part.
For example, with boundary XYZ, this body (CRLF line endings):
--XYZ Content-Disposition: form-data; name="a" foo--XYZ Content-Disposition: form-data; name="b" bar --XYZ--
is parsed as two fields (a="foo", b="bar"), while a strict parser would return a single field a whose value includes the embedded token.
Requiring the CRLF prefix (other than for the first delimiter at the start of the body) would align Django with the RFC. Lenient clients that rely on the current behavior could be affected, so this may need a deprecation path or at least a release note.