Commit graph

4 commits

Author SHA1 Message Date
Kamron Batman
24bcfee554
feat(network): lean base pools (#2641)
Some checks are pending
Build / Build (MacOS 15) (push) Waiting to run
Build / Build (MacOS 26) (push) Waiting to run
Build / Build (AlmaLinux 10) (push) Waiting to run
Build / Build (Debian 12) (push) Waiting to run
Build / Build (Debian 13) (push) Waiting to run
Build / Build (Fedora 44) (push) Waiting to run
Build / Build (CentOS 10 Stream) (push) Waiting to run
Build / Build (CentOS 9 Stream) (push) Waiting to run
Build / Build (Ubuntu 26) (push) Waiting to run
Build / Build (Ubuntu 22) (push) Waiting to run
Build / Build (Ubuntu 24) (push) Waiting to run
**Follow-on to #2639 (merged). References a local IORingGroup `1.0.13-preview.11` pack until 1.0.13 (modernuo/IORingGroup#15) is published; do not merge before that switch.**

## Summary

Consumes IORingGroup's lean base pools (modernuo/IORingGroup#15): both network pools now start with one slab, grow a slab at a time with the population, and trim idle slabs back after quiet periods.

- Fixes the transport's send-pool cap: previously only 1024 of the 4096 connections could get a send buffer; connection 1025 was closed at accept.
- Network memory at boot drops from about 96 MB to about 10 MB at the defaults; a full 4096 logged-in connections is about 1.25 GB of base buffers plus the growth budget.
- New settings: `network.initialBufferSlabs` (default 1; slabs of each pool held from boot and the trim floor) and `network.maxBufferSlabs` (default 128; divides the connection maximum into slabs, 32 connections per slab). Both are coerced with a warning; the same value feeds the ring table and the manager so they cannot drift.
- The Debug-only maintenance line includes base-pool capacity and releases.
- `dev-docs/server-requirements.md` rewrites the network memory story and adds the two settings.

## Pre-auth buffers

Every connection starts on the transport's platform-minimum buffers (4 KB receive, 4 KB send) instead of the base pools. It is promoted to full-size buffers (64 KB receive, `network.sendBufferSize` send) when the game server verifies its account — the point where `NetState.Account` is assigned in the `GameServer_AwaitingGameServerLogin` or `GameServer_LoggedIn` state, so a verdict that lands after the parser has moved on still promotes. Nothing ever moves back. The login-server pass stays on the small buffers for its whole lifetime.

Before credentials verify, nothing promotes: the 4 KB send ring is the entire pre-auth send budget, and a connection that overruns it is dropped as exhausted, exactly as the receive side drops a packet header declaring more bytes than the receive buffer can hold (new guard in `HandlePacket`; it also closes the old 65535-byte edge on 64 KB buffers). The stock login sequence sends under 2 KB. After credentials verify, the send path promotes on demand if it ever needs to (unbudgeted, outside the memory ceiling and the shrink bookkeeping — promotion is not growth), and the oversize-packet guard waits on a pending receive promotion or retries a stalled one once for a verified account before disconnecting. A completion that fills the receive buffer arms no receive, so `HandleReceive` now calls `RingSocket.ResumeReceive()` after the parse loop — at 4 KB a burst of small packets fills the buffer in one completion.

Net effect: a flood of unauthenticated connections tops out at about 32 MB across the full 4096-connection cap where the platform minimum is 4 KB (the transport's retained slabs and the base pools used by logged-in players are separate), and never allocates a base-pool slab.

Platform note: on Windows Server 2012 R2 / 2016 the transport's legacy mapping path floors at 64 KB: the pre-auth receive pool is off there (its base is 64 KB), while the pre-auth send buffer starts at 64 KB under the 256 KB base. The server logs the effective sizes at startup.

## Testing

Server.Tests (905) and UOContent.Tests green on the preview pack (one pre-existing `FamiliarAITests` failure from #2644 reproduces on `main`, tracked separately). New tests cover both coercions, that the ring's registration table equals `RequiredRegisteredBuffers` for the configured values, promotion on game-server auth and on a late account, no promotion on the login server, pre-credential overrun ending in exhaustion, post-credential on-demand promotion, the oversize-packet guard through loopback (error, wait, retry with an account, and a promotion made pending mid-parse), and receiving again after a burst fills the initial buffer.
2026-09-20 00:24:15 -07:00
Kamron Batman
31cd19b05b
feat(network): grow the send buffer on demand instead of disconnecting (#2639)
## Problem

A connection's send buffer is a fixed 256 KB. A burst of world traffic (a crowded area, a mass spawn, a war) that outruns the client's acknowledgements fills it, `NetState.Send()` reports "send buffer exhausted", and the player is disconnected. Raising the size for everyone multiplies the per-connection footprint (4096 × 256 KB is already 1 GB at full occupancy, page-locked on Windows).

## What changes

- **Growth.** When a packet does not fit (the write span is too small, the packet is larger than the span, or compression returns 0), `Send()` asks the transport to grow the buffer to the next power-of-two tier and retries, up to `network.sendBufferMaxSize` (2 MB). Compression retries once per tier since its output size is not known in advance, including when the buffer is completely full. Only when growth is refused does the existing exhaustion disconnect run. The success path is unchanged.
- **Memory ceiling.** Growth is refused (with a once-a-minute warning) when the process working set exceeds `network.memoryCeilingPercent` (80) of the memory available to the process (container-aware; `0` turns the check off). The figure is sampled at startup and refreshed each maintenance tick.
- **Shrink.** A grown socket returns to the base buffer once it is drained and 30 s have passed since its last growth, attempted from the `DataSent` handler and from the 5 s alive sweep.
- **Retention.** Every minute a timer calls the transport's `Maintain()`, which trims idle tier slabs down to the peak concurrent usage of the last 15 minutes, so recurring bursts reuse buffers without allocation while rare ones give the memory back. The line logs at Debug, and only when capacity, usage, or the floor changed or a growth was refused (budget, at max, or ceiling), so an idle shard logs nothing.
- **Budget.** `network.sendBufferGrowthBudget` (256 MB) caps the tier pools' capacity; a positive value below one tier slab is raised with a warning, a negative one is clamped to 0 (growth off). Worst case is base × connections plus the budget.
- Settings are coerced with accurate warnings (power of two, minimum, 256 MB transport ceiling). `[dumpnetstates` gains the send buffer size. `dev-docs/server-requirements.md` describes the new memory story.

## Tests

`NetStateSendBufferTests` (real loopback sockets): growth instead of disconnect with a byte-exact stream, compressed growth against the compressor's own output, the grow-then-copy path, growth with a send genuinely in flight, refusal past the maximum, refusal under the ceiling, refusal on a closing socket, shrink after the hold (direct and through the alive sweep), and the setting coercions. Server.Tests 891 passed, UOContent.Tests 1052 passed against the published 1.0.12.

Reviewed per task, whole-branch, and adversarially by a second model (twice, the second time jointly with the transport branch); all findings addressed.
2026-09-12 16:35:06 -07:00
Kamron Batman
240118340e
fix: stop the idle-sleep backoff tripping on healthy hosts (#2572)
## Problem

The late-wake detector added in #2559 suspends idle sleeping on perfectly healthy hosts. The visible symptom is this Warning firing periodically on stable machines:

> This host returned a 2ms idle wait at least 8ms late 2 time(s) in the last second; idle sleeping suspended for 5000ms

Demoting it to Debug would hide the symptom but not the cost: every one of those lines means the shard dropped idle sleeping for 5s and burned a full core for no reason. The detector is what was mis-tuned.

## Cause 1 — lateness was a count, not a rate

An idle loop performs **~400–500 sleeps per second** (2ms each, bounded by the 8ms wheel tick). The trip condition was `late > 1` across two consecutive one-second samples — a **0.4% tail-outlier rate**. A co-tenant burst, a page fault, or another process changing the system timer resolution clears that bar on a healthy host.

A host that genuinely cannot schedule the process — throttled burstable vCPU — returns *most* of its waits late. Signal and noise were two orders of magnitude apart, and the check sat in the noise.

Now gated on the proportion, with the absolute count kept as a floor:

```csharp
if (late <= _lateWakeThreshold)             { _consecutiveBadSamples = 0; return; }  // floor
if (late * 100 < sleeps * _lateWakePercent) { _consecutiveBadSamples = 0; return; }  // rate
```

New `server.lateWakePercent` (default `10`). The floor is what keeps a window with only a handful of sleeps from tripping on a meaningless percentage; `server.lateWakeThreshold` keeps its existing meaning.

## Cause 2 — GC pauses were charged to the host

`dev-docs/debugging-event-loop.md` already documents that the GC collects preferentially **during idle sleeps** — that is the natural pause point it looks for. So the detector was systematically measuring the GC's chosen pause point and billing it to the host's scheduler. Not an occasional coincidence; a designed-in one.

```csharp
var collections = GC.CollectionCount(1);
NetState.WaitForCompletion(requested);
...
if (elapsed - requested >= Timer.TickRate && GC.CollectionCount(1) == collections)
```

Gen1 (which counts gen2 with it) rather than gen0 — gen0 pauses don't approach the 8ms `TickRate` bar anyway, and gating on them would discard useful samples. The second read short-circuits behind the overshoot test, so the common path costs **one** `GC.CollectionCount` per sleep: an internal counter read, single-digit nanoseconds, ~500/sec.

## Cause 3 — every backoff logged at Warning

Tiered to the escalation that already existed, since a single suspension is recoverable and not something an operator can act on:

| Backoff | Level |
|---|---|
| 1–2 | `Debug` |
| 3–5 | `Warning` (now includes the sleep count and "for the Nth time running") |
| ceiling | `Error`, unchanged |
| recovery | `Information` (new) |

Each backoff doubles the suspension, so every line is already a distinct escalation step — no further rate limiting needed.

## Drive-by

The `BackoffResetAfterCleanMs` reset only ran on the path to a *new* backoff, making it unreachable for a host that recovered for good — such a host never cleared its escalation or re-armed `_loggedBackoffCeiling`. It now runs on every health sample, which is also what makes the new recovery line reachable.

## Testing

Full solution builds clean, 0 warnings. No tests added: the state is private static in `Core` coupled to `_tickCount` with no injection point, and nothing covered it before — adding a seam purely to test it seemed worse than the gap. Happy to add one if reviewers disagree.
2026-08-13 19:58:03 -07:00
Kamron Batman
6d846b11e5
perf: Sleep the event loop when idle. Fixes networking micro-stalls. Adds event loop instrumentation. (#2559)
## Problem

`RunEventLoop` span through its body regardless of whether there was anything to do — ~10% of a desktop core for an empty shard, and ~70% of a core on a 3 vCPU VPS. A process that never idles is exactly what burstable vCPU plans throttle, which is how this surfaced: lag spikes that went away when the operator bought more cores. The spin also denied the GC its natural pause points, so memory climbed until a world save forced a collection — alarming in task manager, harmless in practice, and a recurring source of "is my server leaking?" reports.

## Result

Windows desktop, real world of **190,728 items / 33,158 mobiles**, no players, saves and prebake off, three consecutive runs:

| | Legacy spin | Idle sleeping |
|---|---|---|
| **CPU** | 10.42 – 10.50% of one core | **0.78 – 1.00%** |
| **Tick lag** (peak/15s) | 4–10 ms | 5–11 ms |

**~10× less CPU with tick lag unchanged** — the CPU came free rather than being traded for latency. Slower hosts gain proportionally more. Spin mode (`server.eventLoopIdleWaitMs=0`) independently gained **7× the iterations per core** (1.19M → 8.3M cycles/sec) from the ring's AcceptEx rework.

## How

The loop blocks in `NetState.WaitForCompletion` whenever every queue it drains is empty (all the drains are bounded, so leftovers keep it awake). Receive completions, new connections, and cross-thread `LoopContext.Post` (via the ring's sticky `Wake()`) are all in the wait set, so sleeping adds no latency to any of them. Only timer-driven logic sees wheel lag, bounded by the idle wait.

**Health is measured at the only place sleeping can cause harm.** A sleep is bounded by the time to the next wheel turn, so a correctly honoured sleep can never miss a deadline — the only failure mode is the host returning the wait late. That overshoot is measured on every sleep (one extra timestamp read; production's entire accounting cost), and an escalating backoff suspends sleeping when it persists. By construction, server work — saves, heavy staff commands, deep timer callbacks — cannot trip it, so the warning means exactly one thing: *the host is not scheduling the process promptly*, with two known remedies (dedicated CPU, or `=0`). Hosts with no high-resolution wait mechanism at all are detected once at startup and spin instead.

**CPS is removed.** `Core.CyclesPerSecond`/`AverageCPS` measured nothing actionable before and became actively misleading once the loop sleeps (the rate is set by the sleep, not by shard health). The admin gump's Performance page now shows the verdict instead: `Healthy` / `Sleep suspended (host)` / `Spinning (configured)`.

## Configuration

| Setting | Default | Meaning |
|---|---|---|
| `server.eventLoopIdleWaitMs` | `2` | Longest idle block. Measured across 1/2/4/8 ms, 2 is where the trade stops being free. `0` = never sleep: ~98% of a core, zero scheduling overhead — for large shards on dedicated CPU. |
| `server.lateWakeThreshold` | `1` | Idle waits the host may return a full tick late, per second, before sleeping backs off. Raise for jittery hosts; very high disables the backoff. |

## Diagnostics (compiled out by default)

`dotnet build -p:EventLoopProfiling=true` compiles in `EventLoopProfiler` — every hook is `[Conditional("EVENT_LOOP_PROFILING")]`, so normal builds contain zero profiling IL. The profiling build decomposes each second of wall time into **work (per loop phase) / sleep / GC pause / stolen residual**, keeps ~15 minutes of history in a ring buffer, and the `[LoopStats` command prints the last minute and dumps the full history to CSV. `dev-docs/debugging-event-loop.md` is the diagnosis guide (for humans and AI): what production already tells you, when to flip the profiling build, the signature table for host-steal vs deep-processing vs GC vs wake bugs, why dotnet-trace comes last, and the GC/RAM "leak" misconception.

## Verification

- 815 Server.Tests green; both build configurations compile.
- Docker echo harness green on epoll and io_uring (ping-pong mode); kqueue verified manually on an M1 Max.
- A/B measurements and per-change numbers: `measure/event-loop` branch.

## Notes

The full measurement harness and vendored ring sources used to develop this live on the [`measure/event-loop`](https://github.com/modernuo/ModernUO/tree/measure/event-loop) branch, kept for future loop work.
2026-08-09 13:24:59 -07:00