Skip to content

gh-158897: Optimize io.BufferedReader.readline() by calling memchr() - #158944

Merged
vstinner merged 5 commits into
python:mainfrom
vstinner:readline_memchr
Oct 7, 2026
Merged

vstinner merged 5 commits into
python:mainfrom
vstinner:readline_memchr

Conversation

@vstinner

@vstinner vstinner commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Optimize io.BufferedReader.readline() when lines are close to the buffer size (128 kB by default) or longer than the buffer size. Replace the C loop searching for the newline byte in the buffer with a memchr() call which is more efficient.

…chr()

Optimize io.BufferedReader.readline() when lines are close to the
buffer size (128 kB by default) or longer than the buffer size.
Replace the C loop searching for the newline byte in the buffer with
a memchr() call which is more efficient.
Comment thread Modules/_io/bufferedio.c Outdated
@vstinner

vstinner commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

I ran #158897 (comment) benchmark on Linux (Fedora 44) with CPU isolation.

$ git switch main
$ make
$ ./python bench.py -o main.json -p70
$ git switch readline_memchr 
$ make
$ ./python bench.py -o memchr.json -p70 -v 

Results: Mean +- std dev: [main] 1.78 ms +- 0.14 ms -> [memchr] 417 us +- 34 us: 4.26x faster

This change makes this specific benchmark 4.26x faster!

But well, in practice, text files with lines around 128 kB or longer than 128 kB should be rare.

I had some issues to get reliable results, so I ran the benchmark with -p70 to run more processes than the default (20).

@vstinner

vstinner commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

@leandrodamascena: Would you be able to check if this change solves the performance regression that you spotted in #158897 (comment)? (Can you run your benchmarks on a patched Python?) Tell me if you need help.

@vstinner

vstinner commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

The benchmark #158897 (comment) is a microbenchmark which seems to depend heavily on low-level things like CPU cache misses or CPU branch prediction.

After I fixed my off-by-one error, I reran the benchmark to be sure that the result remains the same, and the measure on the main branch jumped from 1.78 ms +- 0.14 ms to 1.87 ms +- 0.18 ms (same C code, just recompiled).

Anyway, the updated result: Mean +- std dev: [main] 1.87 ms +- 0.18 ms -> [memchr] 330 us +- 30 us: 5.66x faster. It's now 5.7x faster...

I'm not sure how to get more reliable results. I suppose that building Python with PGO would help to get more efficient machine code most of the time.

@leandrodamascena

Copy link
Copy Markdown

Thanks @vstinner! I built the latest version of this PR (with the off-by-one fix) into my CPython OCI images and tested it on AWS Lambda x86_64 with PGO+LTO, comparing it against its base with 25 paired samples per scenario:

  • Mixed lines: 3.981 ms → 0.946 ms (-76.4%, CI -77.9% to -74.5%)
  • Lines close to the buffer size: 247.7 µs → 197.2 µs (-20.5%, CI -21.5% to -19.7%)
  • Long lines: 3.648 ms → 0.784 ms (-78.6%, CI -79.2% to -78.5%)

So yes, this fixes the regression in my Lambda workload. Since main was about 43% slower than before the regression in the mixed case, this should leave it well ahead of where it was before.

Catching regressions early, before they reach a release, is exactly what I hope my workload can do, so getting a fix this quickly is really satisfying. Thanks a lot for looking into this!

@vstinner

vstinner commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

Mixed lines: 3.981 ms → 0.946 ms (-76.4%, CI -77.9% to -74.5%)
Lines close to the buffer size: 247.7 µs → 197.2 µs (-20.5%, CI -21.5% to -19.7%)
Long lines: 3.648 ms → 0.784 ms (-78.6%, CI -79.2% to -78.5%)

Did you compare the main branch to this PR? Or Python 3.15 to this PR?

Anyway, it's faster in all cases, good. Thanks for checking.

So yes, this fixes the regression in my Lambda workload. Since main was about 43% slower than before the regression in the mixed case, this should leave it well ahead of where it was before.

According to my experiment, maybe Python 3.15 landed in the lucky case where the hot code uses the right machine code which is "not slow". When I made changes in the main branch, I got various performance results even when I made C code changes unrelated to the readline function. It reminded me the old https://vstinner.github.io/journey-to-stable-benchmark-deadcode.html issue that I solved with PGO build.

Catching regressions early, before they reach a release, is exactly what I hope my workload can do, so getting a fix this quickly is really satisfying. Thanks a lot for looking into this!

Please continue running these benchmark tracking work, it's useful!

@vstinner
vstinner enabled auto-merge (squash) October 7, 2026 12:46
@leandrodamascena

Copy link
Copy Markdown

Thanks Victor! I compared this PR version (fc422ed) with the main commit it was based on (1818fba), not with Python 3.15.

Both images were built with the same PGO+LTO configuration and tested on AWS Lambda x86_64.

I’ll keep running these tests. Thanks for the encouragement and for looking into this so quickly!

@vstinner
vstinner merged commit 1610807 into python:main Oct 7, 2026
55 checks passed
@vstinner
vstinner deleted the readline_memchr branch October 7, 2026 13:18
@vstinner

vstinner commented Oct 7, 2026

Copy link
Copy Markdown
Member Author

Merged. Thanks for reviews and benchmarks @maurycy and @leandrodamascena.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

extension-modules C modules in the Modules dir performance Performance or resource usage

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants