Opened 10 hours ago
Last modified 4 hours ago
#37223 closed New feature
Signing's JSONSerializer uses latin-1, causing silent mojibake with UTF-8-emitting serializers — at Version 1
| Reported by: | Paolo Melchiorre | Owned by: | Paolo Melchiorre |
|---|---|---|---|
| Component: | Core (Other) | Version: | dev |
| Severity: | Normal | Keywords: | signing latin1 utf8 json |
| Cc: | Andy Miller | Triage Stage: | Unreviewed |
| Has patch: | yes | Needs documentation: | no |
| Needs tests: | no | Patch needs improvement: | no |
| Easy pickings: | no | UI/UX: | no |
Description (last modified by )
Summary
django.core.signing.JSONSerializer encodes its JSON payload with latin-1 and decodes it with latin-1. With the default ensure_ascii=True serializer the output is pure ASCII, so latin-1 and UTF-8 are indistinguishable and the bug is invisible. But any serializer that emits raw UTF-8 (orjson, msgspec, json.dumps(..., ensure_ascii=False)) produces non-ASCII bytes for non-ASCII data, and JSONSerializer.loads silently mojibakes them when decoding as latin-1. No BadSignature and no UnicodeDecodeError is raised — the data is returned corrupted as if it were valid.
Both SESSION_SERIALIZER and signing.dumps(serializer=...) are documented extension points, so a user adopting a UTF-8-emitting serializer hits this. It was documented from the third-party-package side by Adam Johnson in django-orjson PR #15 (https://github.com/adamchainz/django-orjson/pull/15).
How to reproduce
JSONSerializer currently does:
# django/core/signing.py class JSONSerializer: def dumps(self, obj): return json.dumps(obj, separators=(",", ":")).encode("latin-1") def loads(self, data): return json.loads(data.decode("latin-1"))
A payload containing héllo, written by a UTF-8-emitting serializer as the bytes h\xc3\xa9llo, is decoded by loads as latin-1 into héllo. The HMAC is over the bytes (which are intact), so the signature still verifies and the corrupted value is returned to the caller as if it were valid. A minimal reproduction:
from django.core.signing import JSONSerializer # What a UTF-8-emitting serializer (e.g. orjson) would produce for {"text": "héllo"} external_bytes = b'{"text":"h\xc3\xa9llo"}' # Django's loads mojibakes it assert JSONSerializer().loads(external_bytes) == {"text": "héllo"} # silently corrupted assert JSONSerializer().loads(external_bytes) != {"text": "héllo"} # the real value is lost
Diagnosis
The latin-1 step was introduced in commit [58a086acfb] (Oct 2012, "Required serializer to use bytes in loads/dumps") as part of the Python 3 bytes-interface porting, chosen for consistency with the bytes interface rather than because latin-1 was required. With the default ensure_ascii=True serializer the output is always pure ASCII, so latin-1 and UTF-8 produce identical bytes and the bug is invisible — which is why it has gone unfixed since 2012. It surfaces the moment any non-default serializer is used.
The bug was surfaced while analyzing the pluggable JSON backend proposal (#187), which had to carve core/signing.py out of the swappable backend because of this latin-1 step.
Scope
Only django/core/signing.py is in scope. Two contrib consumers are transitively covered because they use signing through the public API:
django/contrib/sessions/serializers.py—JSONSerializeris a re-export ofsigning.JSONSerializer.django/contrib/sessions/backends/base.pyanddjango/contrib/sessions/backends/signed_cookies.py— usesigning.dumps/signing.loads.
django/contrib/messages/storage/cookie.py is not covered by this fix: it has its own MessagePartGatherSerializer.dumps and MessageSerializer.loads that call json.dumps(...).encode("latin-1") and json.loads(data.decode("latin-1")) directly (lines 68 and 73), independent of signing.JSONSerializer. That is the same bug pattern on a separate code path and should be filed as a follow-up ticket. django/contrib/admin does not use signing directly.
Proposed direction
The natural fix is to switch the encode/decode step from latin-1 to utf-8 in both dumps and loads, leaving the separators=(",", ":") argument untouched so determinism is preserved. The exact shape of the change is left to the pull request.
Backwards compatibility
Fully backward-compatible. Existing signed tokens contain only ASCII bytes (ensure_ascii=True escapes every non-ASCII character to \uXXXX), and ASCII decodes identically under latin-1 and UTF-8, so every token already in the wild verifies and decodes to the same value. The compress=True path is covered by the same argument: existing compressed tokens are ASCII.
New tokens with non-ASCII data will contain raw UTF-8 bytes instead of \uXXXX escapes, so they are slightly shorter and round-trip correctly through the new loads.
Suggested test coverage
Tests should be added to the existing signing test module (tests/signing/tests.py) and should cover:
- a regression test that feeds raw UTF-8 bytes (as a UTF-8-emitting external serializer would produce) to loads and asserts the original value is returned, not the latin-1 mojibake — this is the test that fails before the fix and passes after;
- a round-trip of a non-ASCII payload (including combining marks and non-BMP characters) through dumps/loads.
The dumps encode step is not independently testable with the current JSONSerializer (ensure_ascii=True makes its output always ASCII), so no dumps-side regression test is warranted; that change is forward-looking and kept for symmetry with loads.
Related
- PR #21651 — Pull request implementing this fix.
- new-features #187 — Pluggable JSON serialization/deserialization backend (this fix removes the signing carve-out from step 1).
- django-orjson PR #15 — Documented the mojibake from the third-party-package side.
- Commit [58a086acfb] — introduced the latin-1 encode/decode step (Oct 2012).
Change History (1)
comment:1 by , 9 hours ago
| Description: | modified (diff) |
|---|---|
| Has patch: | set |
| Owner: | set to |
| Status: | new → assigned |