ModernUO/dev-docs/server-requirements.md
Kamron Batman 31cd19b05b
feat(network): grow the send buffer on demand instead of disconnecting (#2639)
## Problem

A connection's send buffer is a fixed 256 KB. A burst of world traffic (a crowded area, a mass spawn, a war) that outruns the client's acknowledgements fills it, `NetState.Send()` reports "send buffer exhausted", and the player is disconnected. Raising the size for everyone multiplies the per-connection footprint (4096 × 256 KB is already 1 GB at full occupancy, page-locked on Windows).

## What changes

- **Growth.** When a packet does not fit (the write span is too small, the packet is larger than the span, or compression returns 0), `Send()` asks the transport to grow the buffer to the next power-of-two tier and retries, up to `network.sendBufferMaxSize` (2 MB). Compression retries once per tier since its output size is not known in advance, including when the buffer is completely full. Only when growth is refused does the existing exhaustion disconnect run. The success path is unchanged.
- **Memory ceiling.** Growth is refused (with a once-a-minute warning) when the process working set exceeds `network.memoryCeilingPercent` (80) of the memory available to the process (container-aware; `0` turns the check off). The figure is sampled at startup and refreshed each maintenance tick.
- **Shrink.** A grown socket returns to the base buffer once it is drained and 30 s have passed since its last growth, attempted from the `DataSent` handler and from the 5 s alive sweep.
- **Retention.** Every minute a timer calls the transport's `Maintain()`, which trims idle tier slabs down to the peak concurrent usage of the last 15 minutes, so recurring bursts reuse buffers without allocation while rare ones give the memory back. The line logs at Debug, and only when capacity, usage, or the floor changed or a growth was refused (budget, at max, or ceiling), so an idle shard logs nothing.
- **Budget.** `network.sendBufferGrowthBudget` (256 MB) caps the tier pools' capacity; a positive value below one tier slab is raised with a warning, a negative one is clamped to 0 (growth off). Worst case is base × connections plus the budget.
- Settings are coerced with accurate warnings (power of two, minimum, 256 MB transport ceiling). `[dumpnetstates` gains the send buffer size. `dev-docs/server-requirements.md` describes the new memory story.

## Tests

`NetStateSendBufferTests` (real loopback sockets): growth instead of disconnect with a byte-exact stream, compressed growth against the compressor's own output, the grow-then-copy path, growth with a send genuinely in flight, refusal past the maximum, refusal under the ceiling, refusal on a closing socket, shrink after the hold (direct and through the alive sweep), and the setting coercions. Server.Tests 891 passed, UOContent.Tests 1052 passed against the published 1.0.12.

Reviewed per task, whole-branch, and adversarially by a second model (twice, the second time jointly with the transport branch); all findings addressed.
2026-09-12 16:35:06 -07:00

129 lines
7.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Server Requirements
Hardware guidance for running a ModernUO shard.
## Tiers
| Use | vCPU | RAM | Storage |
|---|---|---|---|
| Development / test | 2 **dedicated** | 2 GB | SSD |
| Small live shard (< 50 concurrent) | 4 dedicated | 4 GB | NVMe |
| Medium (50200) | 48 | 8 GB | NVMe |
| Large (200+) | 8+, high clock | 16 GB+ | NVMe |
These are starting points. Save size drives RAM more than player count does, and single-thread
clock speed drives tick latency more than core count does. Both are explained below.
## Dedicated vCPU, not burstable
This matters more than any other line on this page.
Budget VPS plans sold as "2 vCPU" are frequently shared or burstable: you get a CPU credit balance
or a cgroup quota, and once it is exhausted the hypervisor throttles you. Throttling shows up in
game as periodic freezes that correlate with nothing in your logs, and it is the single most common
cause of "ModernUO is laggy on my $3/month VPS".
Symptoms worth checking before blaming the server:
- Steal time above ~1% (`top`, the `%st` column on Linux)
- Lag that disappears when you move to a larger plan with the same core count
- Tick lag spikes with no matching CPU spike in the process itself
## Cores
Game logic is **single-threaded**. Every mobile, item, timer, and packet handler runs on one
thread, so a shard's headroom is bounded by how fast one core is. Two fast cores beat four slow
ones.
Cores beyond the first are used by:
- **World saves.** `world.useMultithreadedSaves` (default on) spins up `ProcessorCount - 1`
serialization workers plus one inline on the main thread. On a 2-core box that is one worker; on
a 2-core box with a large world, consider setting it to `false` so saves do not contend with the
loop.
- **The .NET runtime.** Tiered JIT compilation (heaviest in the first minutes after boot) and
background GC.
- **Everything else on the machine**, including your OS and, on Windows, antivirus.
Since ModernUO 2026 the loop sleeps when idle, so an empty shard costs roughly 1% of a core rather
than spinning. That change disproportionately helps small hosts.
## Memory
Three things dominate, and only one of them scales with players.
**World size.** A world of ~190,000 items and ~33,000 mobiles loads in about a second and is not
itself large. Items and mobiles are the cheap part.
**Saves.** Each serialization worker pre-allocates a heap sized to its share of the last save, at
roughly 1.25× total save size, and those buffers are retained afterwards. A 400 MB save therefore
implies about 500 MB of resident serialization heap on top of the live world. **This is the reason
1 GB hosts are not viable for a real shard**, even though an empty one boots fine.
**Map residency.** `TileMatrix` reads map blocks from disk on demand and caches them permanently
there is no eviction. Memory climbs toward full-facet residency as players explore. Felucca's land
tiles alone are around 117 MB, and statics are larger.
Optional systems can add substantially more. The pathfinding prebake
(`pathfinding.prebakeMaps`) peaks above 1 GB of heap while baking. Budget for it or leave it off on
small hosts.
Network buffers are 64 KB receive plus a configurable send buffer per connection. Send memory is
`network.sendBufferSize` at rest and can grow to `network.sendBufferMaxSize` under load. Shared
send-buffer tier memory is capped by `network.sendBufferGrowthBudget`, and growth is refused when
process memory exceeds `network.memoryCeilingPercent` of available memory. The worst case is the
base send-buffer size times the connection count, plus the shared growth budget: at the defaults,
100 players is roughly 32 MB at rest, and the growth budget can add up to another 256 MB under
load.
ModernUO runs **Workstation GC**, which is the right default for small hosts. Do not switch to
Server GC on a 2-core box.
## Storage
Saves are write-heavy bursts. Cheap network-attached storage with throttled IOPS will stall the
save path, and `World.WaitForWriteCompletion` blocks the loop at shutdown. Use local NVMe or SSD.
Budget disk for: the world save, plus archives and backups if `autoArchive` is enabled (retention
defaults keep 24 hourly, 30 daily, and 12 monthly copies), plus the pathfinding cache if enabled.
## Operating systems
See the README for the full supported list. Two things are worth calling out:
- **Windows Server 2012 R2 and 2016 sleep via a raised timer resolution.** Sleeping for a couple
of milliseconds prefers a high-resolution waitable timer, which requires Windows 10 1803 /
Server 2019. On older versions the ring falls back to `timeBeginPeriod(1)`, which raises the
system timer resolution to 1 ms so the plain wait timeout is accurate enough. The trade-off is a
higher interrupt rate (system-wide on those versions) an acceptable price on a dedicated game
server, and the reason the high-resolution timer is preferred where it exists.
Only if *both* mechanisms fail does the server detect it at startup, log it, and spin instead
the same behaviour as setting `server.eventLoopIdleWaitMs` to 0: a full core at idle, and zero
missed deadlines. A host that claims short waits but cannot deliver them is caught at runtime by
the adaptive backoff.
- **Linux kernel 6.1** or newer (Debian 12 and equivalents). io_uring is used where available, with
automatic epoll fallback.
## Tuning for a small host
| Setting | Default | Why change it |
|---|---|---|
| `server.eventLoopIdleWaitMs` | `2` | `0` never sleeps: ~98% of one core, but zero skipped timer slots and zero lag. The choice for a large shard on dedicated CPU that would rather spend a core than risk a late wake. Above `2` the wheel starts losing slots. |
| `server.lateWakeThreshold` | `1` | Floor for the backoff: idle waits the host may return a full tick late, per second, before the rate test below applies at all. Raise on a jittery host; set very high to disable the backoff. |
| `server.lateWakePercent` | `10` | Share of a second's idle waits that must come back late before idle sleeping backs off. An idle loop sleeps hundreds of times a second, so a bare count cannot tell a few tail outliers from a host that never schedules the process a genuinely bad host misses *most* of its waits. `0` leaves `lateWakeThreshold` in sole charge. |
| `world.useMultithreadedSaves` | `true` | Set `false` on 2-core hosts so saves do not contend with the game loop. |
| `pathfinding.prebakeMaps` | varies | Leave off on memory-constrained hosts; it peaks above 1 GB while baking. |
| `network.sendBufferSize` | 256 KB | Lower it if you are memory-bound with many connections. |
| `network.sendBufferMaxSize` | 2 MB (`2097152`) | Ceiling a single connection's send buffer can grow to under load. Lower it on memory-constrained hosts; raise it if slow clients are disconnected with "send buffer exhausted". |
| `network.sendBufferGrowthBudget` | 256 MB (`268435456`) | Cap on the shared memory the larger send-buffer tiers may use. Lower it on memory-constrained hosts. |
| `network.memoryCeilingPercent` | 80% | Refuse send-buffer growth once the process is above this share of available memory; 0 turns the check off. |
| `autoArchive.*` retention | 24h/30d/12m | Reduce if disk is tight. |
## Am I undersized?
Watch the log. The server warns when the host returns idle waits late and suspends idle sleeping,
and says so at startup if the host cannot honour short waits at all. Those warnings mean the host
is not scheduling the process promptly typical of burstable or shared vCPU plans and no
server-side change fixes that. For anything deeper, see
[debugging-event-loop.md](debugging-event-loop.md).