perf: Skip request-body compression for small payloads (#988)
Until now the Python client compressed every request body regardless of size. It now skips compression below 1024 B. Closes #934 ## Why 1024 B - **Parity with the JS client.** It uses the same threshold, the same operator (`>=`), and the same byte-based measurement (`MIN_COMPRESS_BYTES` in `apify-client-js/src/utils.ts`), so both clients now put identical bytes on the wire. - **Small bodies grow.** `run.charge` 40 B to 59 B, `kvs.set_record` 27 B to 45 B, single-item `dataset.push_items` 10 B to 30 B. gzip breaks even only at ~88 B of realistic JSON, and high-entropy bodies such as binary key-value store records inflate at any size, by ~23 B for gzip and ~4 B for brotli. - **No packet is saved below ~1 KB**, and removing a packet is the only thing that buys latency. With ~500 B of request headers on a 1400 B MSS: | Raw body | gzipped | Segments raw to gzip | Packet saved? | | --- | --- | --- | --- | | 786 B | 344 B | 1 to 1 | no | | 1071 B | 447 B | 2 to 1 | yes | | 8250 B | 2958 B | 7 to 3 | yes | Compressing everything instead, the reverse direction raised in the issue, doesn't pay off. Over 1200 realistic bodies it inflates 749 of them while total bytes move only from −62.1% to −63.8%, and that gain sits entirely in the 512–1024 B band where no packet is saved anyway. At 1024 B nothing inflates. ## Changes - `MIN_COMPRESSION_SIZE = 1024`. `_prepare_request_call` compresses only at or above it, measured on the encoded bytes so a multibyte `str` is judged correctly. - A caller-supplied `Content-Encoding` is now dropped on a skipped body, where it would otherwise survive and mislabel an uncompressed payload. - The async client skips the `asyncio.to_thread` hop for any body it won't compress, the hop costing 36–68 µs against 5–12 µs of compression. `_is_body_worth_compressing` sits next to the rule it mirrors so the two can't drift apart. A `json=` body still hops, its size being unknown until serialized. - Docs: new "Minimum body size" section. The page claimed the client compresses every request body. ## Verification Against the live API, 8 value shapes round-trip byte-identical with and without compression, small uncompressed bodies are accepted by `dataset.push_items`, `rq.add_request`, `rq.batch_add_requests`, `dataset.update`, `schedules().create`, `schedule.update`, and `webhooks().create`, and latency is unchanged. Tests that compressed tiny bodies were passing vacuously. They now cover the 0/1/1023/1024/1025 boundary for both gzip and brotli, the byte-vs-character threshold, the dropped header, and the thread-hop decisions. `test_run_charge`'s `compression` axis was the suite's only end-to-end compression coverage and had gone vacuous on its 39-byte body, so it's replaced by a test asserting an above-threshold body reaches the server compressed under the configured algorithm. *✍️ Drafted by Claude Code*
V
Vlada Dusek committed
6bd31b2abb840ea9af1efc98bf984c85b8510ce3
Parent: f657f76
Committed by GitHub <noreply@github.com>
on 8/3/2026, 3:21:50 PM