Skip to content

Bound the server TCP listener backlog at 32 - #229

Open
derekste wants to merge 2 commits into
epics-base:masterfrom
derekste:dev/listener-backlog
Open

derekste wants to merge 2 commits into
epics-base:masterfrom
derekste:dev/listener-backlog

Conversation

@derekste

@derekste derekste commented Oct 1, 2026 •

Copy link
Copy Markdown

Increase the server TCP listener backlog from 4 to a fixed 32. This keeps a small platform-independent limit, including on RTEMS, while avoiding the accept-queue overflows reproduced during simultaneous client startup.

On Linux amd64, 20 interleaved fresh bursts at each of 16, 32 and 64 contexts passed with backlog 32 and zero ListenOverflows/ListenDrops. Backlog 16 passed at 16/32 contexts but failed 6/20 bursts at 64. All 19 IOC CTests and all 25 PVXS core suites (2490 assertions) passed for 32. Measurements and retained evidence.

The committed follow-up was rebuilt on macOS arm64: six focused PVXS suites passed 583 checks, and the existing IOC Redis/PVA source-health regression passed against the rebuilt library, covering partial outage, recovery, aliases, reloads and confirmation epochs.

The final diff changes only the listener constant. This replaces the earlier SOMAXCONN proposal with the bounded value supported by the comparison; it does not introduce a configurable or unlimited backlog. The five-second fixture deadline is not a product startup guarantee.

@mdavidsaver

Copy link
Copy Markdown
Member

@derekste Thank you for keeping this to a one line change. Please try to be similarly concise with the PR description.

When 16 independent PVA client contexts connect simultaneously, the server's four-slot TCP accept queue can overflow and delay some clients beyond the burst fixture's five-second initial-monitor deadline.

This is the desired behavior when a server is being subjected to a DoS attack. Perhaps not by intent, but in effect. I am open to changing the magic 4 to something else, but I do not want to disable this limit completely. Certainly not on RTEMS where file descriptors are precious.

Looking on my Debian 13 host.

$ grep -wR SOMAXCONN /usr/include/
...
/usr/include/bits/socket.h:#define SOMAXCONN    4096
/usr/include/x86_64-linux-gnu/bits/socket.h:#define SOMAXCONN   4096

imo. 4096 is too large.

$ grep -wR SOMAXCONN /usr/x86_64-w64-mingw32/include/
...
/usr/x86_64-w64-mingw32/include/winsock2.h:#define SOMAXCONN 0x7fffffff

And 0x7fffffff is waaaaay too large.

$ find /opt/rtems/5 -name include | xargs grep -wR SOMAXCONN
/opt/rtems/5/i386-rtems5/include/sys/socket.h:#define   SOMAXCONN       128

Even 128 seems too large on RTEMS.

@derekste derekste changed the title Use SOMAXCONN for the server TCP listener backlog Bound the server TCP listener backlog at 32 Oct 9, 2026
@derekste

derekste commented Oct 9, 2026

Copy link
Copy Markdown
Author

Updated this PR from SOMAXCONN to a bounded backlog of 32 and shortened the description around the decision evidence.

Backlog 32 passed all 60 interleaved fresh 16/32/64-context bursts with zero listener overflows; backlog 16 failed 6/20 at 64 contexts. The refreshed library/core build, six focused core suites, and the downstream IOC Redis/PVA regression pass. The downstream 0.9.0 release uses a separate one-line fork commit based on its existing PVXS pin, so this upstream PR can continue on its own review timeline.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants