Repository navigation
using eof? as readiness probe causes timeout exception when used with JDK servers #233
Description
Activity
To start a jetty server to reproduce this you can download jetty 12 for example here:
curl -fLO https://repo1.maven.org/maven2/org/eclipse/jetty/jetty-home/12.1.1/jetty-home-12.1.1.tar.gz tar xvf jetty-home-12.1.1.tar.gz cd jetty-home-12.1.1 openssl req -x509 -newkey rsa:2048 -nodes -subj "/CN=localhost" -keyout "localhost-key.pem" -out "localhost-cert.pem" -days 365 openssl pkcs12 -export -inkey localhost-key.pem -in localhost-cert.pem -out keystore.p12 -name jetty -passout pass:password java -Djavax.net.debug=all -jar ./start.jar --module=ssl,https jetty.ssl.host=localhost jetty.ssl.port=8443 jetty.sslContext.keyStorePath=./keystore.p12 jetty.sslContext.keyStorePassword=password jetty.sslContext.keyStoreType=PKCS12
Then run the client.rb from above. It works in Ubuntu 24.04 about 25% of the time, in rocky 8 pretty much always.
Fun fact: if you strace the server using:
strace -ff -tt -p "$(pgrep -f '*java')" -e trace=network -s 200
it becomes a "heisenbug" :) meaning, it does not appear, at least not in rocky 8. (the timinig is changed slightly, such that the NST arrives in time, or sth)I basically agree
eof?blocking is extremely confusing and made this proposal: https://bugs.ruby-lang.org/issues/20215If you support that, please add your support to the issue, and I'll take it to the developer meeting. With extra support/valid use cases it may be acceptable.
Reacted by Friedrich Raschwitz👋
Independent confirmation against a different server stack, plus a data point that supports the approach in #232.
Environment
- Ruby 3.3.10,
net-http0.9.1,opensslgem 3.2.3, OpenSSL library 3.3.7 - Server: Jetty on JDK 17, TLS 1.3 (
TLS_AES_256_GCM_SHA384), responses withContent-Length, noConnection: close
Symptom
Deterministic on the first reuse of every new connection — not intermittent for us:
req0: 200 64ms port=52082 req1: 200 299951ms port=49320Note the local port changes: req1 blocks on the original socket, then completes on a new one. Response size is irrelevant — a 15-byte
SELECT 1response and a 10,958-byte response stall identically, and neither is chunked.Two things worth adding to the diagnosis:
1. The stall is bounded by the server's idle timeout, not by any client timeout. Ours is exactly Jetty's
maxIdleTime(defaultPT5M), hence ~300s every time. This corroborates step 8 of your trace — the client only unblocks when close_notify arrives. Critically, neitherread_timeout(set to 15s) norwrite_timeout(default 60s) bounds it, because the block is in the liveness probe rather than in the request. So callers cannot defend against this with timeouts, which makes it fail as an unexplained multi-minute stall rather than an error.2. Draining the pending record before reuse fixes it completely, which is direct evidence that the readability signal is the
NewSessionTicketand that a non-blocking drain (as in #232) is sufficient:resp = http.request(req) # req0: 200, 5ms io = http.instance_variable_get(:@socket).io while IO.select([io], nil, nil, 0.5) begin io.read_nonblock(4096) rescue IO::WaitReadable, EOFError break end end resp = http.request(req) # req1: 200, 6ms — same local port, connection reused
Without the drain: ~300s and a new socket. With it: 6ms on the same socket.
Reproduces with bare
Net::HTTP(no Faraday, nonet-http-persistent) and withmax_retries = 0. Forcingmax_version: :TLS1_2also avoids it, consistent with TLS 1.2 delivering tickets during the handshake rather than after the first response.+1 on #232, and support for the
eof?semantics proposal in https://bugs.ruby-lang.org/issues/20215 — a non-blocking readiness check is the right primitive here.Worth noting that this caused a severe production incident for us!
- Ruby 3.3.10,
What is happening
When using keep-alive connections on Jetty/JDK 17+ servers requests will hang and time out on certain platforms, for example Rocky 8.
This could actually happen with any kind of server and on any platform and it can easily be reproduced, see the example code at the bottom.
Here is a small code example to illustrate the occurrence a bit better:
Related issue from ruby-lang
The bug in question was actually already described a few years ago here:
https://bugs.ruby-lang.org/issues/19017
Why is it happening
Requests are issued via transport_request which uses begin_transport as a pre-flight check, reestablishing the connection when certain conditions are met. One of the conditions is checking readability via TCPSocket#wait_readable(0) and subsequently calling SSLSocket#eof? if it is readable.
Net::BufferedIO#eof? ends up calling sysread which eventually issues a SSL_read from OpenSSL. SSL_read will try to get application data from the socket.
The problem is that the TCPSocket may signal readability, despite no application data being present.
In JDK 17+ there is a flag that triggers sending an updated NST (New Session Ticket) upon certain conditions, after sending a HTTP response. If the ticket arrives quickly enough then it will be there right after application data, signalling readability on the TCP socket right before begin_transport is called which leads to the hang.
I have traced such an exchange client-side via OpenSSL tracing, this is how it looks llike:
here eof? is called and the cllient hangs
This is how it plays out for clients on puppetserver, which is based on Jetty. I opened an issue about it a few months ago: OpenVoxProject/openvox-server#25
Possible solutions
eof? is called here in order to check for a EOFError due to the connection being closed unexpectedly or something else. This is not a reliable way to check for this, even if it wouldn't hang (we could turn off SSL_MODE_AUTO_RETRY to make it return when receiving handshake data), because it does not actually read until the end of the stream. A close notify could still be on-line.
In order to properly check for EOFError we could instead call a non-blocking read until EOF or until the TCP socket is no longer readable. See #232
Calling a SSL_read with length 0 would be preferred, in order to be sure not to consume application data without storing it in a useful buffer, but that is not possible atm. It should not be possible for this to occur anyways, as unprompted application data is AFAIK not supported yet in this module (e.g. HTTP/2 or Websockets). So perhaps this is a good enough solution.
Reproducing the issue
How to
Use the following example client in conjunction with the example server, this should work on any system.
When on Rocky 8 you can also use Jetty 12, the following comment details on how to start an example Jetty server for this purpose
Client
Server
build with
gcc -O2 -Wall server.c -o server -lssl -lcryptoServer Certificate
EDIT: Rewrote the issue for better readability.