10,000 players need 2.55 full generations/sec at peak.
A dedicated H100 pair clears that load. Image cap 16 delivers 4.08/sec at 8.27s p95; cap 32 delivers 4.88/sec at 9.92s. The data supports separate text/image GPUs; it does not establish that both need 80GB.
What the tested configurations can handle
| Configuration | Full creations/sec | p95 latency | p99 latency | Modeled daily players¹ | Hardware implication |
|---|---|---|---|---|---|
| 1 H100 · both models | 1.32 | 10.25s | 11.02s | 5,105 | Below the 10k all-cloud target |
| 2 H100s · serial image worker | 2.81 | 6.86s | 7.11s | 10,996 | Enough for 10k; little peak margin |
| 2 H100s · image batch cap 16 | 4.08 | 8.27s | 8.51s | 15,709 | Best measured latency/capacity balance for 10k |
| 2 H100s · image batch cap 32 | 4.88 | 9.92s | 10.31s | 18,850 | More capacity within the 15s target |
| 2 H100s · image batch cap 64 | 5.48 | 15.48s | 18.04s | 21,207 * | More throughput; misses 15s p95 |
Each row: 1,000 randomly arriving full jobs, zero failures, at a different offered load. This compares measured operating points, not batch caps at identical load. Text and image run in parallel; p95 means 95% finish within that time. ¹Player counts use rates rounded down to 0.1/sec, 5 creations/day, 4× peak and 10% extra attempts. *Cap 64 fits the 30s target, not 15s.
Measurement scope: H100 SXM 80GB; Qwen3 0.6B BF16 stand-in with exactly 1,792 output tokens; real Anima model, 512×512, two steps. Latency includes service queueing, excludes mobile network/storage. These are short load tests, not daily production observations. PCIe hardware was not measured.
Change traffic; see the hardware count
This is capacity arithmetic over measured rates. It does not predict how a different GPU will perform.
Traffic assumptions and formula
peak requests/sec = daily players × creations/player/day × cloud share × peak factor × (1 + extra attempts) / 86,400
At the default inputs: average 0.637/sec; ordinary 2× peak 1.273/sec; event 4× peak 2.546/sec. A 50% cloud share halves demand. The 2× ordinary factor is a rounded inference from mobile activity data; 4× event and 10% extra attempts are explicit planning assumptions. No Atelico production arrival trace was used. Multiplying peak benchmark throughput by 24 hours would overstate daily-player capacity when demand is uneven.
Batching tells us which GPU resources matter
Image GPU: batch 16 peaked at 32.4 GiB; batch 32 at 59.2 GiB.
| Images per fixed batch | Images/sec | Time for whole batch | Peak GPU memory² | Memory implication |
|---|---|---|---|---|
| 1 | 4.10 | 0.24s | 7.4 GiB | 24GB-class capacity |
| 2 | 4.54 | 0.44s | 8.9 GiB | 24GB-class capacity |
| 4 | 4.65 | 0.86s | 12.4 GiB | 24GB-class capacity |
| 8 | 4.72 | 1.69s | 19.0 GiB | 24GB-class capacity |
| 16 | 4.89 | 3.27s | 32.4 GiB | 48GB-class capacity |
| 32 | 5.55 | 5.77s | 59.2 GiB | Above 48GB; tested on 80GB |
| 64 | 5.98 | 10.70s | 79.0 GiB | 99.2% of physical H100 memory |
| 128 | — | OOM | Failed allocation | No successful batch |
²Load + warm peak, sampled every 10ms. GiB is the measured unit; 32.4 GiB is above a nominal 32GB card’s usable capacity. Memory-fit classes describe this implementation’s footprint, not a speed prediction for those cards. Fixed batches exclude queue/fill wait. DiT and full VAE are batched; prompt encoding and PNG encoding are included.
Text GPU: concurrency increases throughput until latency becomes the limit.
| Concurrent text requests | H100 requests/sec | H100 output tokens/sec | H100 p95 | 3090 requests/sec | 3090 p95 |
|---|---|---|---|---|---|
| 1 | 0.25 | 447 | 4.04s | 0.17 | 5.74s |
| 8 | 1.66 | 2,979 | 4.82s | 0.94 | 8.47s |
| 16 | 2.84 | 5,083 | 5.65s | 1.45 | 11.06s |
| 32 | 4.39 | 7,862 | 7.32s | 1.75 | 18.36s |
| 64 | 6.13 | 10,992 | 10.44s | 2.42 | 26.55s |
| 128 | 7.71 | 13,825 | 16.61s | 2.59 | 50.68s |
| 256 | 9.01 | 16,143 | 28.64s | — | — |
| 512 | 9.11 | 16,318 | 59.62s | — | — |
Same 1,792-token text workload. These are fixed-concurrency stage sweeps, not full-creation arrival tests. H100 uses continuous batching and reserves ~72.4 GiB; that reservation is not the minimum memory needed. The model also ran on the 24GB 3090. At 512 requests, the H100 run filled 99.98% of its allocated KV cache and recorded 312 preemptions.
Where latency and concurrent work accumulate
| Two-H100 run | Completed/sec | In-flight full jobs: mean / max | Text p95 | Image p95 | Full-creation p95 |
|---|---|---|---|---|---|
| Serial image | 2.81 | 17.3 / 35 | 6.85s | 1.72s | 6.86s |
| Image cap 16 | 4.08 | 29.4 / 48 | 8.27s | 3.51s | 8.27s |
| Image cap 32 | 4.88 | 40.8 / 70 | 9.62s | 8.45s | 9.92s |
| Image cap 64 | 5.48 | 57.7 / 113 | 11.13s | 15.42s | 15.48s |
At cap 16, text dominates. At cap 64, the image stage dominates full latency; image queue p95 reaches 6.65s. At cap 32 the server held about 41 full jobs concurrently on average, peaking at 70. That is simultaneous generation work, not players with the game open.
Caps are maxima, not forced batch sizes. Mean actual image batches were 2.7, 10.4 and 24.4 for caps 16, 32 and 64; image queue p95 was 1.43s, 3.60s and 6.65s. The 20ms collection window does not bound total queue wait. Text and image overlap, so their latency percentiles should not be added together.
How this changes the hardware choice
| Option | What the numbers say | Internal conclusion |
|---|---|---|
| One H100, both models | 1.32/sec in the arrival test; highest fixed-concurrency point 2.20/sec at 14.69s p95. Neither establishes 2.55/sec at peak. | Insufficient demonstrated capacity for 10k all-cloud at 4× peak. May suit lower cloud share; exact saturation is not measured. |
| Two H100s, split stages | 4.08–4.88/sec at 8.27–9.92s p95. | $66–96k reference system has measured capacity margin on SXM. Image cap 16 is enough for this traffic. |
| 3090 text + faster image GPU | Text alone reaches 49.4s p95 near required arrival rate. | Rule out this text bottleneck for the 10k all-cloud / 15s case. |
| Faster text GPU + 48GB image GPU | Text model fits 24GB; image batch 16 peaks at 32.4 GiB. Performance of the proposed cheaper pair is unknown. | Most useful cost-reduction direction. Size each stage for at least 2.55/sec; memory capacity alone cannot identify the winning SKU. |
Performance sensitivity for a different GPU variant
| Versus 4.8/sec reference | Full jobs/sec | Modeled daily players | 10k capacity? |
|---|---|---|---|
| 0% slower | 4.80 | 18,850 | Yes |
| 25% slower | 3.60 | 14,138 | Yes |
| 40% slower | 2.88 | 11,310 | Yes |
| 50% slower | 2.40 | 9,425 | No |
What-if arithmetic only, not a PCIe/SXM conversion. These rows do not predict latency. A 40% throughput reduction still covers modeled demand; a 50% reduction does not.
The remaining numbers that determine the cheapest pair
- Text: arrivals/sec and p95 on a cheaper modern GPU at the same 1,792-token length.
- Images: batch 8/16 throughput and memory on a 24/48GB GPU.
- Combined: full-creation latency at 2.55/sec and the point where queues grow.
Those are missing benchmark cells. Existing data cannot honestly turn them into a specific cheaper GPU recommendation. PCIe H100, M5, and other untested hardware have no measured rate in this study.
Cost per player: purchase once, host monthly
| Cost item | Two-H100 reference | Meaning |
|---|---|---|
| Two H100 80GB PCIe cards | $59,510–84,000 | Public component price range; these are not the measured SXM modules. |
| Compatible host | $6,000–12,000 | CPU/RAM/storage/chassis/PSU allowance. |
| Upfront hardware | $65,510–96,000 | Before tax, freight and extended support. |
| Monthly hosting | ~$1,000 | Planning allowance for rack, power and network. |
| Hosting / player-day at 10k DAU | $0.00333 = 0.333¢ | At 300,000 player-days per 30-day month. |
| Hosting / full creation | $0.000667 = 0.0667¢ | At 1.5 million completed creations/month. |
| Equipment over 36 months + hosting | 0.94–1.22¢ / player-day | Comparison metric; hardware cash payment is upfront. |
Engineering, separate backend/storage, maintenance and repairs are excluded. The $1k hosting allowance is $12k for 12 months; this is not a GPU rental plan. References checked 5 October: Datenshop, ServerSupply, Columbia Colocation. The lowest-cost adequate system remains unpriced because its GPU performance is unmeasured.
Use the measurements and rerun the tests
Runbook and exact benchmark commands · Compact hardware metrics · Detailed benchmark tables and provenance · Image batching implementation
The full-creation harness measures arrival rate, completed/sec, p50/p95/p99, failures and queue growth. Separate sweeps measure text concurrency and image tensor batches with GPU-memory sampling. The image microbatch gateway is a prototype; production integration is needed to realize those measured batching gains.
Existing M4 Max, 3090 and Thor native-engine references
| Hardware | Full jobs/sec · C1 | p95 · C1 | Full jobs/sec · C4 | p95 · C4 |
|---|---|---|---|---|
| M4 Max 64GB | 0.142 | 7.50s | 0.141 | 29.72s |
| RTX 3090 24GB | 0.144 | 7.75s | 0.143 | 29.20s |
| AGX Thor 128GB | 0.062 | 17.68s | 0.062 | 67.32s |
20 jobs/row using native Qwen Q4 + game adapter with variable output lengths. These are a different backend/workload from the CUDA batching study; they do not isolate hardware speed. Their serialized text path explains why more concurrency barely increases throughput. Do not size a Mac/Thor fleet from these as if they were optimized server results.