feat: Adds Latin1 text support (#2317)

## Summary

- Adds proper Latin1 encoding support, replacing CP1252 usage throughout the codebase
- Adds specialized, optimized string decoding methods with safe string filtering for each encoding type
- Filters invalid Unicode characters (C0/C1 control codes, non-characters) by removal rather than replacement since
the UO client renders nothing for these characters
- Fixes UTF-16 null terminator position handling to correctly advance by 2 bytes

## Changes

TextEncoding.cs

- Added SearchValues-based invalid byte/char detection for efficient filtering
- Added encoding-specific GetString methods: GetStringAscii, GetStringLatin1, GetStringUtf8, GetStringBigUni,
GetStringLittleUni
- Each method supports a safeString parameter for filtering invalid characters
- Little-endian UTF-16 uses direct memory cast for zero-copy decoding on LE systems
- Invalid characters are removed (not replaced with U+FFFD) since the client renders nothing for them

SpanReader.cs

- Added ReadLatin1() and ReadLatin1Safe() methods
- Rewrote encoding-specific read methods to use optimized TextEncoding.GetString* methods
- Fixed UTF-16 null terminator handling: position now correctly advances by byteLength (2) instead of 1

SpanWriter.cs

- Added WriteLatin1 and WriteLatin1Null methods

## Packet Updates

- Updated all packet code to use Latin1 encoding instead of CP1252
- Affected: account packets, equipment packets, menu packets, message packets, mobile packets, player packets, secure
trade packets, vendor packets, gump packets, book packets, mahjong packets

## Filtering Behavior

Invalid characters filtered in safe mode:
```
┌───────────────┬────────────────────────┐
│     Range     │      Description       │
├───────────────┼────────────────────────┤
│ 0x00-0x1F     │ C0 control codes       │
├───────────────┼────────────────────────┤
│ 0x7F          │ DEL                    │
├───────────────┼────────────────────────┤
│ 0x80-0x9F     │ C1 control codes       │
├───────────────┼────────────────────────┤
│ 0xFFFE-0xFFFF │ Unicode non-characters │
└───────────────┴────────────────────────┘
```

Note: Surrogate pairs (0xD800-0xDFFF) are not filtered because proper validation requires context checking for paired
vs unpaired surrogates. The UO client renders nothing for these anyway.

## Test Plan

- All 631 Server.Tests pass
- Verified client rendering behavior using TestUnicodeGump command (pages 1-5)
- Confirmed U+FFFD, unpaired surrogates, and non-characters all render as blank in client
- Verified Latin1 characters (0xA0-0xFF) display correctly
- Verified C1 control codes (0x80-0x9F) are filtered and don't display
This commit is contained in:
Kamron Batman 2026-01-22 15:51:00 -08:00 • committed by GitHub
parent 1e4c32c809
commit 0404251638
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
27 changed files with 735 additions and 96 deletions

View file

@ -309,7 +309,7 @@ public static class OutgoingMobilePackets
writer.Write((byte)0x98); // Packet ID
writer.Write((ushort)37);
writer.Write(m.Serial);
writer.WriteAscii(m.Name ?? "", 29);
writer.WriteLatin1(m.Name ?? "", 29);
writer.Write((byte)0); // Null terminator
ns.Send(writer.Span);
@ -507,7 +507,7 @@ public static class OutgoingMobilePackets
writer.Write((byte)0x11); // Packet ID
writer.Seek(2, SeekOrigin.Current);
writer.Write(beheld.Serial);
writer.WriteAscii(name, 30);
writer.WriteLatin1(name, 30);
writer.WriteAttribute(beheld.HitsMax, beheld.Hits, version == 0, true);
writer.Write(canBeRenamed);
writer.Write((byte)version);