Some checks are pending
Build / Build (MacOS 15) (push) Waiting to run
Build / Build (MacOS 26) (push) Waiting to run
Build / Build (AlmaLinux 10) (push) Waiting to run
Build / Build (Debian 12) (push) Waiting to run
Build / Build (Debian 13) (push) Waiting to run
Build / Build (Fedora 44) (push) Waiting to run
Build / Build (CentOS 10 Stream) (push) Waiting to run
Build / Build (CentOS 9 Stream) (push) Waiting to run
Build / Build (Ubuntu 26) (push) Waiting to run
Build / Build (Ubuntu 22) (push) Waiting to run
Build / Build (Ubuntu 24) (push) Waiting to run
**Follow-on to #2639 (merged). References a local IORingGroup `1.0.13-preview.11` pack until 1.0.13 (modernuo/IORingGroup#15) is published; do not merge before that switch.** ## Summary Consumes IORingGroup's lean base pools (modernuo/IORingGroup#15): both network pools now start with one slab, grow a slab at a time with the population, and trim idle slabs back after quiet periods. - Fixes the transport's send-pool cap: previously only 1024 of the 4096 connections could get a send buffer; connection 1025 was closed at accept. - Network memory at boot drops from about 96 MB to about 10 MB at the defaults; a full 4096 logged-in connections is about 1.25 GB of base buffers plus the growth budget. - New settings: `network.initialBufferSlabs` (default 1; slabs of each pool held from boot and the trim floor) and `network.maxBufferSlabs` (default 128; divides the connection maximum into slabs, 32 connections per slab). Both are coerced with a warning; the same value feeds the ring table and the manager so they cannot drift. - The Debug-only maintenance line includes base-pool capacity and releases. - `dev-docs/server-requirements.md` rewrites the network memory story and adds the two settings. ## Pre-auth buffers Every connection starts on the transport's platform-minimum buffers (4 KB receive, 4 KB send) instead of the base pools. It is promoted to full-size buffers (64 KB receive, `network.sendBufferSize` send) when the game server verifies its account — the point where `NetState.Account` is assigned in the `GameServer_AwaitingGameServerLogin` or `GameServer_LoggedIn` state, so a verdict that lands after the parser has moved on still promotes. Nothing ever moves back. The login-server pass stays on the small buffers for its whole lifetime. Before credentials verify, nothing promotes: the 4 KB send ring is the entire pre-auth send budget, and a connection that overruns it is dropped as exhausted, exactly as the receive side drops a packet header declaring more bytes than the receive buffer can hold (new guard in `HandlePacket`; it also closes the old 65535-byte edge on 64 KB buffers). The stock login sequence sends under 2 KB. After credentials verify, the send path promotes on demand if it ever needs to (unbudgeted, outside the memory ceiling and the shrink bookkeeping — promotion is not growth), and the oversize-packet guard waits on a pending receive promotion or retries a stalled one once for a verified account before disconnecting. A completion that fills the receive buffer arms no receive, so `HandleReceive` now calls `RingSocket.ResumeReceive()` after the parse loop — at 4 KB a burst of small packets fills the buffer in one completion. Net effect: a flood of unauthenticated connections tops out at about 32 MB across the full 4096-connection cap where the platform minimum is 4 KB (the transport's retained slabs and the base pools used by logged-in players are separate), and never allocates a base-pool slab. Platform note: on Windows Server 2012 R2 / 2016 the transport's legacy mapping path floors at 64 KB: the pre-auth receive pool is off there (its base is 64 KB), while the pre-auth send buffer starts at 64 KB under the 256 KB base. The server logs the effective sizes at startup. ## Testing Server.Tests (905) and UOContent.Tests green on the preview pack (one pre-existing `FamiliarAITests` failure from #2644 reproduces on `main`, tracked separately). New tests cover both coercions, that the ring's registration table equals `RequiredRegisteredBuffers` for the configured values, promotion on game-server auth and on a late account, no promotion on the login server, pre-credential overrun ending in exhaustion, post-credential on-demand promotion, the oversize-packet guard through loopback (error, wait, retry with an account, and a promotion made pending mid-parse), and receiving again after a burst fills the initial buffer.
152 lines
9.9 KiB
Markdown
152 lines
9.9 KiB
Markdown
# Server Requirements
|
||
|
||
Hardware guidance for running a ModernUO shard.
|
||
|
||
## Tiers
|
||
|
||
| Use | vCPU | RAM | Storage |
|
||
|---|---|---|---|
|
||
| Development / test | 2 **dedicated** | 2 GB | SSD |
|
||
| Small live shard (< 50 concurrent) | 4 dedicated | 4 GB | NVMe |
|
||
| Medium (50–200) | 4–8 | 8 GB | NVMe |
|
||
| Large (200+) | 8+, high clock | 16 GB+ | NVMe |
|
||
|
||
These are starting points. Save size drives RAM more than player count does, and single-thread
|
||
clock speed drives tick latency more than core count does. Both are explained below.
|
||
|
||
## Dedicated vCPU, not burstable
|
||
|
||
This matters more than any other line on this page.
|
||
|
||
Budget VPS plans sold as "2 vCPU" are frequently shared or burstable: you get a CPU credit balance
|
||
or a cgroup quota, and once it is exhausted the hypervisor throttles you. Throttling shows up in
|
||
game as periodic freezes that correlate with nothing in your logs, and it is the single most common
|
||
cause of "ModernUO is laggy on my $3/month VPS".
|
||
|
||
Symptoms worth checking before blaming the server:
|
||
|
||
- Steal time above ~1% (`top`, the `%st` column on Linux)
|
||
- Lag that disappears when you move to a larger plan with the same core count
|
||
- Tick lag spikes with no matching CPU spike in the process itself
|
||
|
||
## Cores
|
||
|
||
Game logic is **single-threaded**. Every mobile, item, timer, and packet handler runs on one
|
||
thread, so a shard's headroom is bounded by how fast one core is. Two fast cores beat four slow
|
||
ones.
|
||
|
||
Cores beyond the first are used by:
|
||
|
||
- **World saves.** `world.useMultithreadedSaves` (default on) spins up `ProcessorCount - 1`
|
||
serialization workers plus one inline on the main thread. On a 2-core box that is one worker; on
|
||
a 2-core box with a large world, consider setting it to `false` so saves do not contend with the
|
||
loop.
|
||
- **The .NET runtime.** Tiered JIT compilation (heaviest in the first minutes after boot) and
|
||
background GC.
|
||
- **Everything else on the machine**, including your OS and, on Windows, antivirus.
|
||
|
||
Since ModernUO 2026 the loop sleeps when idle, so an empty shard costs roughly 1% of a core rather
|
||
than spinning. That change disproportionately helps small hosts.
|
||
|
||
## Memory
|
||
|
||
Three things dominate, and only one of them scales with players.
|
||
|
||
**World size.** A world of ~190,000 items and ~33,000 mobiles loads in about a second and is not
|
||
itself large. Items and mobiles are the cheap part.
|
||
|
||
**Saves.** Each serialization worker pre-allocates a heap sized to its share of the last save, at
|
||
roughly 1.25× total save size, and those buffers are retained afterwards. A 400 MB save therefore
|
||
implies about 500 MB of resident serialization heap on top of the live world. **This is the reason
|
||
1 GB hosts are not viable for a real shard**, even though an empty one boots fine.
|
||
|
||
**Map residency.** `TileMatrix` reads map blocks from disk on demand and caches them permanently —
|
||
there is no eviction. Memory climbs toward full-facet residency as players explore. Felucca's land
|
||
tiles alone are around 117 MB, and statics are larger.
|
||
|
||
Optional systems can add substantially more. The pathfinding prebake
|
||
(`pathfinding.prebakeMaps`) peaks above 1 GB of heap while baking. Budget for it or leave it off on
|
||
small hosts.
|
||
|
||
Network buffers come from four pools that grow and shrink with the population rather than being
|
||
sized for a full shard. A connection that has not yet presented valid credentials holds a 4 KB
|
||
receive and a 4 KB send buffer (the platform's page size). On Windows Server 2012 R2 / 2016 the
|
||
transport's legacy mapping path floors at 64 KB: the pre-auth receive pool is off there (its base is
|
||
64 KB), while the pre-auth send buffer starts at 64 KB under the 256 KB base. Everything the server
|
||
sends before that must fit in that ring — a connection that overruns it is dropped; the stock login
|
||
sequence uses under 2 KB. Otherwise, when the game server verifies the account the connection is
|
||
promoted to a 64 KB receive buffer and a `network.sendBufferSize` send buffer from the base pools,
|
||
and nothing ever moves back. A flood of unauthenticated connections tops out at about 32 MB across
|
||
the full 4096-connection cap where the platform minimum is 4 KB (the transport's retained slabs and
|
||
the base pools used by logged-in players are separate), and never allocates a base-pool slab.
|
||
At boot the network holds `network.initialBufferSlabs` slab(s) of each pool — at the defaults one
|
||
2 MB receive slab, one 8 MB send slab and two 128 KB pre-auth slabs, about 10 MB — and allocates
|
||
another slab only when the population needs one. Each slab covers 32 connections at the
|
||
4096-connection maximum. After 15 quiet minutes idle slabs are trimmed back towards current usage,
|
||
never past the last 15 minutes' peak, at one slab per pool per minute and never below
|
||
`network.initialBufferSlabs`. Only the newest slab is trimmed, and buffers are handed out from the
|
||
oldest slab first, so ordinary churn empties the newest slabs; a shard that drops from 4096 players
|
||
to a handful takes about two hours to shrink fully, longer if a long-lived connection still holds a
|
||
buffer in a newer slab.
|
||
|
||
Send memory per authenticated connection is `network.sendBufferSize` at rest and can grow to
|
||
`network.sendBufferMaxSize` under load. Shared send-buffer tier memory is capped by
|
||
`network.sendBufferGrowthBudget`, and growth is refused when process memory exceeds
|
||
`network.memoryCeilingPercent` of available memory. The worst case is the receive and base
|
||
send-buffer sizes times the number of logged-in connections, plus the shared growth budget: a full
|
||
4096 logged-in connections is roughly 1.25 GB of base buffers, and the growth budget can add up to
|
||
another 256 MB.
|
||
|
||
ModernUO runs **Workstation GC**, which is the right default for small hosts. Do not switch to
|
||
Server GC on a 2-core box.
|
||
|
||
## Storage
|
||
|
||
Saves are write-heavy bursts. Cheap network-attached storage with throttled IOPS will stall the
|
||
save path, and `World.WaitForWriteCompletion` blocks the loop at shutdown. Use local NVMe or SSD.
|
||
|
||
Budget disk for: the world save, plus archives and backups if `autoArchive` is enabled (retention
|
||
defaults keep 24 hourly, 30 daily, and 12 monthly copies), plus the pathfinding cache if enabled.
|
||
|
||
## Operating systems
|
||
|
||
See the README for the full supported list. Two things are worth calling out:
|
||
|
||
- **Windows Server 2012 R2 and 2016 sleep via a raised timer resolution.** Sleeping for a couple
|
||
of milliseconds prefers a high-resolution waitable timer, which requires Windows 10 1803 /
|
||
Server 2019. On older versions the ring falls back to `timeBeginPeriod(1)`, which raises the
|
||
system timer resolution to 1 ms so the plain wait timeout is accurate enough. The trade-off is a
|
||
higher interrupt rate (system-wide on those versions) — an acceptable price on a dedicated game
|
||
server, and the reason the high-resolution timer is preferred where it exists.
|
||
|
||
Only if *both* mechanisms fail does the server detect it at startup, log it, and spin instead —
|
||
the same behaviour as setting `server.eventLoopIdleWaitMs` to 0: a full core at idle, and zero
|
||
missed deadlines. A host that claims short waits but cannot deliver them is caught at runtime by
|
||
the adaptive backoff.
|
||
- **Linux kernel 6.1** or newer (Debian 12 and equivalents). io_uring is used where available, with
|
||
automatic epoll fallback.
|
||
|
||
## Tuning for a small host
|
||
|
||
| Setting | Default | Why change it |
|
||
|---|---|---|
|
||
| `server.eventLoopIdleWaitMs` | `2` | `0` never sleeps: ~98% of one core, but zero skipped timer slots and zero lag. The choice for a large shard on dedicated CPU that would rather spend a core than risk a late wake. Above `2` the wheel starts losing slots. |
|
||
| `server.lateWakeThreshold` | `1` | Floor for the backoff: idle waits the host may return a full tick late, per second, before the rate test below applies at all. Raise on a jittery host; set very high to disable the backoff. |
|
||
| `server.lateWakePercent` | `10` | Share of a second's idle waits that must come back late before idle sleeping backs off. An idle loop sleeps hundreds of times a second, so a bare count cannot tell a few tail outliers from a host that never schedules the process — a genuinely bad host misses *most* of its waits. `0` leaves `lateWakeThreshold` in sole charge. |
|
||
| `world.useMultithreadedSaves` | `true` | Set `false` on 2-core hosts so saves do not contend with the game loop. |
|
||
| `pathfinding.prebakeMaps` | varies | Leave off on memory-constrained hosts; it peaks above 1 GB while baking. |
|
||
| `network.sendBufferSize` | 256 KB | Lower it if you are memory-bound with many connections. |
|
||
| `network.sendBufferMaxSize` | 2 MB (`2097152`) | Ceiling a single connection's send buffer can grow to under load. Lower it on memory-constrained hosts; raise it if slow clients are disconnected with "send buffer exhausted". |
|
||
| `network.sendBufferGrowthBudget` | 256 MB (`268435456`) | Cap on the shared memory the larger send-buffer tiers may use. Lower it on memory-constrained hosts. |
|
||
| `network.memoryCeilingPercent` | 80% | Refuse send-buffer growth once the process is above this share of available memory; 0 turns the check off. |
|
||
| `network.initialBufferSlabs` | `1` | Slabs of each base pool held from boot, and the floor the trim never goes below. Raise it on a large shard to pre-warm the pools instead of paying for a slab as the population climbs. |
|
||
| `network.maxBufferSlabs` | `128` | Divides the connection maximum into base-pool slabs: a slab holds `MaxConnections / maxBufferSlabs` connections, 32 at the default. Raise it for finer slabs on a small host (the slab floor is 16 buffers); lowering it makes each slab, and the boot allocation, larger. It is not a connection or memory cap — both pools still reach the connection maximum. |
|
||
| `autoArchive.*` retention | 24h/30d/12m | Reduce if disk is tight. |
|
||
|
||
## Am I undersized?
|
||
|
||
Watch the log. The server warns when the host returns idle waits late and suspends idle sleeping,
|
||
and says so at startup if the host cannot honour short waits at all. Those warnings mean the host
|
||
is not scheduling the process promptly — typical of burstable or shared vCPU plans — and no
|
||
server-side change fixes that. For anything deeper, see
|
||
[debugging-event-loop.md](debugging-event-loop.md).
|