Internal hardware sizing · measured 5 October 2026

10,000 players need 2.55 full generations/sec at peak.

A dedicated H100 pair clears that load. Image cap 16 delivers 4.08/sec at 8.27s p95; cap 32 delivers 4.88/sec at 9.92s. The data supports separate text/image GPUs; it does not establish that both need 80GB.

50,000 creations/day10k daily players × 5 creations; size for 100% cloud
2.55/sec event demand4× average peak, with 10% extra attempts
$66–96k + ~$1k/moTwo-H100 PCIe + host estimate; monthly hosting

What the tested configurations can handle

ConfigurationFull creations/secp95 latencyp99 latencyModeled daily players¹Hardware implication
1 H100 · both models1.3210.25s11.02s5,105Below the 10k all-cloud target
2 H100s · serial image worker2.816.86s7.11s10,996Enough for 10k; little peak margin
2 H100s · image batch cap 164.088.27s8.51s15,709Best measured latency/capacity balance for 10k
2 H100s · image batch cap 324.889.92s10.31s18,850More capacity within the 15s target
2 H100s · image batch cap 645.4815.48s18.04s21,207 *More throughput; misses 15s p95

Each row: 1,000 randomly arriving full jobs, zero failures, at a different offered load. This compares measured operating points, not batch caps at identical load. Text and image run in parallel; p95 means 95% finish within that time. ¹Player counts use rates rounded down to 0.1/sec, 5 creations/day, 4× peak and 10% extra attempts. *Cap 64 fits the 30s target, not 15s.

Measurement scope: H100 SXM 80GB; Qwen3 0.6B BF16 stand-in with exactly 1,792 output tokens; real Anima model, 512×512, two steps. Latency includes service queueing, excludes mobile network/storage. These are short load tests, not daily production observations. PCIe hardware was not measured.

Current sizing recommendation: separate GPUs; use image cap 16 for the 10k baseline, with cap 32 available for more throughput. The measured cap-16 rate is 1.57× peak demand after rounding down. Two H100s are a capacity reference; the lowest-cost combination is still unknown.

Change traffic; see the hardware count

This is capacity arithmetic over measured rates. It does not predict how a different GPU will perform.

2.55/secRequired peak full attempts
1 × 2 GPUsConfigurations / total GPU count
15,709Daily players per configuration
1.6×One configuration / required rate

Traffic assumptions and formula

peak requests/sec = daily players × creations/player/day × cloud share × peak factor × (1 + extra attempts) / 86,400

At the default inputs: average 0.637/sec; ordinary 2× peak 1.273/sec; event 4× peak 2.546/sec. A 50% cloud share halves demand. The 2× ordinary factor is a rounded inference from mobile activity data; 4× event and 10% extra attempts are explicit planning assumptions. No Atelico production arrival trace was used. Multiplying peak benchmark throughput by 24 hours would overstate daily-player capacity when demand is uneven.

Batching tells us which GPU resources matter

Image GPU: batch 16 peaked at 32.4 GiB; batch 32 at 59.2 GiB.

Images per fixed batchImages/secTime for whole batchPeak GPU memory²Memory implication
14.100.24s7.4 GiB24GB-class capacity
24.540.44s8.9 GiB24GB-class capacity
44.650.86s12.4 GiB24GB-class capacity
84.721.69s19.0 GiB24GB-class capacity
164.893.27s32.4 GiB48GB-class capacity
325.555.77s59.2 GiBAbove 48GB; tested on 80GB
645.9810.70s79.0 GiB99.2% of physical H100 memory
128—OOMFailed allocationNo successful batch

²Load + warm peak, sampled every 10ms. GiB is the measured unit; 32.4 GiB is above a nominal 32GB card’s usable capacity. Memory-fit classes describe this implementation’s footprint, not a speed prediction for those cards. Fixed batches exclude queue/fill wait. DiT and full VAE are batched; prompt encoding and PNG encoding are included.

Cheaper image hardware is plausible: cap 16 fits a 48GB-class memory budget and already delivered ample throughput on H100. A 24GB-class budget fits the batch-8 footprint. We have no corresponding throughput measurements on those cheaper image GPUs. Cap 64 consumed 99.2% of physical H100 memory; doubling again to 128 failed.

Text GPU: concurrency increases throughput until latency becomes the limit.

Concurrent text requestsH100 requests/secH100 output tokens/secH100 p953090 requests/sec3090 p95
10.254474.04s0.175.74s
81.662,9794.82s0.948.47s
162.845,0835.65s1.4511.06s
324.397,8627.32s1.7518.36s
646.1310,99210.44s2.4226.55s
1287.7113,82516.61s2.5950.68s
2569.0116,14328.64s——
5129.1116,31859.62s——

Same 1,792-token text workload. These are fixed-concurrency stage sweeps, not full-creation arrival tests. H100 uses continuous batching and reserves ~72.4 GiB; that reservation is not the minimum memory needed. The model also ran on the 24GB 3090. At 512 requests, the H100 run filled 99.98% of its allocated KV cache and recorded 312 preemptions.

A 3090 is too slow as the sole text GPU for the 10k all-cloud event case: at 2.63 arriving text requests/sec, p95 reached 49.4s and the queue grew. Its best tested sub-15s fixed-concurrency point was 1.45/sec. The gap is throughput, not fitting the model in memory. H100 at concurrency 64 achieved 6.13/sec at 10.44s; pushing to 512 barely improves throughput over 256 and takes about a minute.

Where latency and concurrent work accumulate

Two-H100 runCompleted/secIn-flight full jobs: mean / maxText p95Image p95Full-creation p95
Serial image2.8117.3 / 356.85s1.72s6.86s
Image cap 164.0829.4 / 488.27s3.51s8.27s
Image cap 324.8840.8 / 709.62s8.45s9.92s
Image cap 645.4857.7 / 11311.13s15.42s15.48s

At cap 16, text dominates. At cap 64, the image stage dominates full latency; image queue p95 reaches 6.65s. At cap 32 the server held about 41 full jobs concurrently on average, peaking at 70. That is simultaneous generation work, not players with the game open.

Caps are maxima, not forced batch sizes. Mean actual image batches were 2.7, 10.4 and 24.4 for caps 16, 32 and 64; image queue p95 was 1.43s, 3.60s and 6.65s. The 20ms collection window does not bound total queue wait. Text and image overlap, so their latency percentiles should not be added together.

How this changes the hardware choice

OptionWhat the numbers sayInternal conclusion
One H100, both models1.32/sec in the arrival test; highest fixed-concurrency point 2.20/sec at 14.69s p95. Neither establishes 2.55/sec at peak.Insufficient demonstrated capacity for 10k all-cloud at 4× peak. May suit lower cloud share; exact saturation is not measured.
Two H100s, split stages4.08–4.88/sec at 8.27–9.92s p95.$66–96k reference system has measured capacity margin on SXM. Image cap 16 is enough for this traffic.
3090 text + faster image GPUText alone reaches 49.4s p95 near required arrival rate.Rule out this text bottleneck for the 10k all-cloud / 15s case.
Faster text GPU + 48GB image GPUText model fits 24GB; image batch 16 peaks at 32.4 GiB. Performance of the proposed cheaper pair is unknown.Most useful cost-reduction direction. Size each stage for at least 2.55/sec; memory capacity alone cannot identify the winning SKU.

Performance sensitivity for a different GPU variant

Versus 4.8/sec referenceFull jobs/secModeled daily players10k capacity?
0% slower4.8018,850Yes
25% slower3.6014,138Yes
40% slower2.8811,310Yes
50% slower2.409,425No

What-if arithmetic only, not a PCIe/SXM conversion. These rows do not predict latency. A 40% throughput reduction still covers modeled demand; a 50% reduction does not.

The remaining numbers that determine the cheapest pair

  • Text: arrivals/sec and p95 on a cheaper modern GPU at the same 1,792-token length.
  • Images: batch 8/16 throughput and memory on a 24/48GB GPU.
  • Combined: full-creation latency at 2.55/sec and the point where queues grow.

Those are missing benchmark cells. Existing data cannot honestly turn them into a specific cheaper GPU recommendation. PCIe H100, M5, and other untested hardware have no measured rate in this study.

Cost per player: purchase once, host monthly

Cost itemTwo-H100 referenceMeaning
Two H100 80GB PCIe cards$59,510–84,000Public component price range; these are not the measured SXM modules.
Compatible host$6,000–12,000CPU/RAM/storage/chassis/PSU allowance.
Upfront hardware$65,510–96,000Before tax, freight and extended support.
Monthly hosting~$1,000Planning allowance for rack, power and network.
Hosting / player-day at 10k DAU$0.00333 = 0.333¢At 300,000 player-days per 30-day month.
Hosting / full creation$0.000667 = 0.0667¢At 1.5 million completed creations/month.
Equipment over 36 months + hosting0.94–1.22¢ / player-dayComparison metric; hardware cash payment is upfront.

Engineering, separate backend/storage, maintenance and repairs are excluded. The $1k hosting allowance is $12k for 12 months; this is not a GPU rental plan. References checked 5 October: Datenshop, ServerSupply, Columbia Colocation. The lowest-cost adequate system remains unpriced because its GPU performance is unmeasured.

Use the measurements and rerun the tests

Runbook and exact benchmark commands · Compact hardware metrics · Detailed benchmark tables and provenance · Image batching implementation

The full-creation harness measures arrival rate, completed/sec, p50/p95/p99, failures and queue growth. Separate sweeps measure text concurrency and image tensor batches with GPU-memory sampling. The image microbatch gateway is a prototype; production integration is needed to realize those measured batching gains.

Existing M4 Max, 3090 and Thor native-engine references
HardwareFull jobs/sec · C1p95 · C1Full jobs/sec · C4p95 · C4
M4 Max 64GB0.1427.50s0.14129.72s
RTX 3090 24GB0.1447.75s0.14329.20s
AGX Thor 128GB0.06217.68s0.06267.32s

20 jobs/row using native Qwen Q4 + game adapter with variable output lengths. These are a different backend/workload from the CUDA batching study; they do not isolate hardware speed. Their serialized text path explains why more concurrency barely increases throughput. Do not size a Mac/Thor fleet from these as if they were optimized server results.