Repository navigation
gh-158897: Optimize io.BufferedReader.readline() by calling memchr() - #158944
Conversation
…chr() Optimize io.BufferedReader.readline() when lines are close to the buffer size (128 kB by default) or longer than the buffer size. Replace the C loop searching for the newline byte in the buffer with a memchr() call which is more efficient.
|
I ran #158897 (comment) benchmark on Linux (Fedora 44) with CPU isolation. Results: This change makes this specific benchmark 4.26x faster! But well, in practice, text files with lines around 128 kB or longer than 128 kB should be rare. I had some issues to get reliable results, so I ran the benchmark with |
|
@leandrodamascena: Would you be able to check if this change solves the performance regression that you spotted in #158897 (comment)? (Can you run your benchmarks on a patched Python?) Tell me if you need help. |
|
The benchmark #158897 (comment) is a microbenchmark which seems to depend heavily on low-level things like CPU cache misses or CPU branch prediction. After I fixed my off-by-one error, I reran the benchmark to be sure that the result remains the same, and the measure on the main branch jumped from Anyway, the updated result: I'm not sure how to get more reliable results. I suppose that building Python with PGO would help to get more efficient machine code most of the time. |
|
Thanks @vstinner! I built the latest version of this PR (with the off-by-one fix) into my CPython OCI images and tested it on AWS Lambda x86_64 with PGO+LTO, comparing it against its base with 25 paired samples per scenario:
So yes, this fixes the regression in my Lambda workload. Since main was about 43% slower than before the regression in the mixed case, this should leave it well ahead of where it was before. Catching regressions early, before they reach a release, is exactly what I hope my workload can do, so getting a fix this quickly is really satisfying. Thanks a lot for looking into this! |
Did you compare the main branch to this PR? Or Python 3.15 to this PR? Anyway, it's faster in all cases, good. Thanks for checking.
According to my experiment, maybe Python 3.15 landed in the lucky case where the hot code uses the right machine code which is "not slow". When I made changes in the main branch, I got various performance results even when I made C code changes unrelated to the readline function. It reminded me the old https://vstinner.github.io/journey-to-stable-benchmark-deadcode.html issue that I solved with PGO build.
Please continue running these benchmark tracking work, it's useful! |
|
Thanks Victor! I compared this PR version (fc422ed) with the main commit it was based on (1818fba), not with Python 3.15. Both images were built with the same PGO+LTO configuration and tested on AWS Lambda x86_64. I’ll keep running these tests. Thanks for the encouragement and for looking into this so quickly! |
|
Merged. Thanks for reviews and benchmarks @maurycy and @leandrodamascena. |
Optimize io.BufferedReader.readline() when lines are close to the buffer size (128 kB by default) or longer than the buffer size. Replace the C loop searching for the newline byte in the buffer with a memchr() call which is more efficient.
BufferedReader.readline()after switching toPyBytesWriter#158897