Performance
A php-via server keeps a context in memory for every open tab and holds one SSE stream per tab. Its capacity depends on what each page view, open tab, action and broadcast costs, and those costs are listed below, measured with 0.14.0 on the code of this website. The full tables are in bench/capacity/RESULTS.md.
Measured on two physical cores of an Intel i5-13500 held at full clock, with one worker, PHP 8.5, OpenSwoole 26.2, opcache on and Brotli level 4. Page views ran with the growth-based cycle collector this website uses (below). A virtual server's core is usually slower: scale the CPU figures as described in measuring your own app. Memory figures carry over.
What each part costs
| Unit | CPU | Memory |
|---|---|---|
| Page view (35 to 173 KB of HTML) | 1.5 to 5.8 ms | 20 to 55 KB until the stream connects or 30 s pass |
| Opening a tab (page view plus stream) | 3.0 to 6.1 ms | see the next row |
| Open tab, Brotli level 4 | none while idle | 590 to 670 KB |
| Open tab, Brotli level 1 | none while idle | 75 to 145 KB |
| Open tab, no Brotli | none while idle | 50 to 95 KB |
Private action with a sync() | 0.11 to 0.14 ms | |
| Action passed to the worker that holds its tab | about twice that of one that reaches it | |
| Broadcast, per receiving tab | about 0.05 ms | |
| Keep-alive comment every 15 s | under 1% of a core for 1,000 tabs |
A broadcast to tabs whose views pass shareRender: true renders once per scope and then writes and
compresses one frame per tab, so the per-tab cost is what grows with the audience. Broadcasts coalesce: a worker renders a busy
scope at most once per tick (25 ms by default), however many actions marked it.
Broadcasting explains the timing.
Memory per open tab and the Brotli level
A tab's context and stream take around 50 to 95 KB. With Brotli on, the stream's encoder adds a window of what it has compressed, which lets a patch that repeats earlier markup shrink to a few bytes. That window is what an open tab mostly costs:
| Home page, 500 tabs, 20 broadcasts per second | Memory per tab | Bytes per tab per second |
|---|---|---|
| Brotli level 4 (default) | 671 KB | 0.69 KB |
| Brotli level 1 | 145 KB | 10.84 KB |
| No Brotli | 93 KB | 25.32 KB |
The encoder grows with traffic: a busy stream, such as a live game board, levels off near 9 MB
at level 4 and 575 KB at level 1. On a server short of memory and not of bandwidth, set
withBrotli(true, 1); the second argument is the level for responses and streams.
Page views that never open a stream
A crawler, a link preview or a prefetch loads the page but never connects its stream. php-via keeps
that context until the connect timeout (withContextTimeouts(connectMs: ...), 30 s by
default), so a crawler fetching 50 pages per second holds 30 to 80 MB.
When the contexts of such a burst expire, php-via destroys them in 10 ms slices, and each one is freed as soon as nothing refers to it. Its internal helpers hold it weakly, so a destroyed context leaves no reference cycles for PHP's cycle collector, whose every run walks all live contexts. While the contexts of a 250,000-view burst expire, the collector runs once. The pauses left in such a burst come from collector runs while the views arrive, which the next section covers.
Cycle collector
PHP's cycle collector frees objects that refer to each other in a loop, and every run walks all live
objects of the worker, every live context included. PHP starts a run once 10,000 possible roots have
piled up, raises that threshold after a run that frees nothing and lowers it again after one that frees
something. A php-via worker keeps those runs and also runs the collector every 30 s:
Config::withGcIntervalMs($ms) sets that interval, and 0 turns the timed runs off.
In a burst of page views, or in an app whose requests leave cycles, PHP's runs come often, and each one
walks every live context. withGcIntervalMs(30_000, onGrowth: true) turns PHP's runs off, and the
worker starts a run itself, checked every 100 ms, when:
- its memory has grown by half since the last run, and by 32 MiB at least,
- half the room that was left below
memory_limitafter the last run is used, - or the interval, 30 s by default, has passed since the last run and possible roots wait.
The collector's work then grows with the memory a worker allocates, not with how many objects its requests touch. One worker served bursts of up to 250,000 page views of a counter page within 20 s, once as it is and once with each view leaving a cycle of 1 KB:
| Burst | Runs | In the collector | Longest pause |
|---|---|---|---|
| No cycles, PHP's runs | 25 | 2.2 to 2.8 s | 0.30 to 0.37 s |
No cycles, onGrowth | 8 to 9 | 0.5 to 0.8 s | 0.24 to 0.40 s |
| A cycle per view, PHP's runs | 201 to 206 | 16.0 to 16.2 s | 0.19 to 0.27 s |
A cycle per view, onGrowth | 10 | 1.1 s | 0.35 to 0.43 s |
Without cycles, onGrowth let the worker serve 29,700 views a second instead of 24,600. With a
cycle per view and PHP's runs, it spent 16 of the 20 s in the collector and served 128,000 views;
with onGrowth it served all 250,000 in 9 s.
Three runs each, held at full clock as above, without Brotli. The longest pause follows the number of live contexts when the last run of the burst comes, which varies from run to run.
A run still walks every live context, so the longest pause stays about one such walk, and runs that
come less often each handle more possible roots. Between onGrowth runs, cycles can take as much memory as the
growth that makes a run due: half of what the worker used after the last run and at least 32 MiB,
or half the room left below memory_limit when that is less. A worker of 5 MB can thus
hold more than six times its live memory in cycles.
With onGrowth, the check runs between coroutines, so a loop that creates cycles without waiting
on I/O, such as a CPU-bound import inside one action, frees them only once it ends. If they outgrow the room
below memory_limit first, the worker dies with a fatal error and every tab on it with it, where
PHP's own runs would have freed them as the loop went. Turn onGrowth on only where no request
runs such a loop, or call gc_collect_cycles() inside it, once per batch of rows say.
Keeping broadcasts narrow
A broadcast to a page's scope re-renders the page and every component on it. Four patterns from this website keep a busy page cheap:
-
Share the render of a view that is the same for every tab:
$c->view(..., shareRender: true)after$c->scope(...). Without it each tab renders its own update. In theshared_readbenchmark ofbench/contention, which renders without writing frames, a broadcast to 2,000 contexts whose view reads five scoped signals from the shared store took 3.6 ms with a shared render and 22 ms when each context rendered its own. Views says when a render must not be shared. -
Give a busy widget a scope of its own and broadcast to that scope. The home page counter lives in
home:counter, so a click re-renders the counters and not the poll next to them. Components shows the code. -
Put presence indicators in their own scope too. A broadcast to
Scope::GLOBALon every connect and disconnect re-renders every open page: with 1,000 readers open, each visitor arriving and leaving cost 506 ms of CPU. With the indicator in its own scope and static pages rendered once, it costs 7 ms. - Render a static page once. A view that returns an empty string on updates sends nothing on the SSE connect or on a broadcast, and its components still patch themselves:
$app->page('/docs/install', function (Context $c): void {
$c->view(fn (bool $isUpdate): string => $isUpdate ? '' : $c->render('docs/install.html.twig'));
});More workers
With withWorkerNum(2) on two cores, page views per second rose from 680 to 1,308 for
/docs/signals and from 177 to 348 for /docs/api, over HTTP/1.1. Scoped signals, session data
and the client list are shared between workers, and the shared context directory has to be sized for your
traffic: Deployment gives the rule.
Two things limit what extra workers buy. Behind a proxy that speaks h2c to php-via, every client arrives on one connection and so on one worker until that connection carries 1,280 streams. Over HTTP/1.1 the clients spread, but a tab's actions usually reach another worker than its stream, which passes them on: with N workers about (N−1)/N of the actions, each at about twice the CPU. Deployment explains both. Handlers that spend their time on CPU gain the most; an app whose actions mostly broadcast gains little.
A small server, worked through
This website runs on a two-core virtual server with 2 GB of memory. With about 1.2 GB left for php-via after the system and the reverse proxy:
| Load | Estimate |
|---|---|
| Open docs tabs, Brotli level 4 | about 2,100 |
| Open docs tabs, Brotli level 1 | about 16,000 |
| Crawler traffic | 50 pages per second hold 30 to 80 MB |
| Visitors arriving and leaving with 1,000 readers open | about 150 per second per core |
| Clicks on a shared widget with 2,000 viewers | about 10 per second per core |
The CPU rows assume a core as fast as the reference machine's; scale them as described below.
Measuring your own app
The harness in
bench/capacity
opens tabs, fires actions and shared clicks at fixed rates, sends visitors through, and reads the
server's CPU and memory from /proc. Its README shows how to pin the server and the
harness to separate cores and hold the server's clock. To compare a server with the reference machine, run
php -d opcache.enable_cli=1 bench/capacity/calibrate.php on it: the reference prints
brotli 7 ms, php 60 ms, and the ratio of your figures to these scales the CPU costs above.