People and screens · Live updates

Bridge and live updates

How one protocol connects every app to the console, and how a change on one screen reaches every other screen in under a second.

Design section
Section 8, agreed 6 Oct 2026 (B1 to B8, G1 to G6)
Main code
admin/routers/bridge.py, packages-ts/omnium-bridge
Main tables
domain_event (the outbox)
Read time
about 20 minutes

In one minute

The bridge is one protocol between a small app (a scorer desk, the games console) and the admin panel that hosts it. The app never calls omnium itself. It sends messages like read, watch and command to the panel, and the panel calls the admin service with the signed-in person's login. This part is built and in use. In the plan, the panel's command messages are run by the command service (Section 4, L15); reads and watches stay with admin-api.

Live updates today work, but each open screen asks the database for news once a second, on its own. That is about 70 to 80 queries a minute for a screen that shows nothing new. An update only says "something changed, read again", so every screen then re-reads. And a stream that gets an error answer can stop for good, with no sign on the screen.

The agreed design turns this around. Postgres tells each server the moment a change is saved (pg_notify, woken 0.20 ms after commit, measured). Each server reads the new changes once, works out each watched view once, and pushes the new value with a version number to every screen that watches it. A screen that sees a skipped number re-reads one record. The client reopens the stream after any close.

Targets: another staff screen updated within 1 s of a save, a second scorer within 0.5 s. Database load stays the same for 5 or 500 open screens. None of the agreed design is in the code yet.

70 to 80
queries a minute per idle screen today (from the code)
0.20 ms
NOTIFY wake-up, median (laptop, 6 Oct)
0.5 s
target: second scorer sees the change
1 s
target: any other staff screen

What this part does

This part answers one question for every open screen: is what I show still true? When one person changes something, every other screen showing it must update quickly. That means a second scorer on the same match, an operator on the schedule, and the console on the medal table.

The user's own words, from the review of Section 7: when two scorers are scoring, the second one's screen should update as soon as the first changes something.

Section 8 answered five questions:

#QuestionWhat goes wrong if we answer it badly
1How fast must another screen show a change?A second scorer works on an old view and taps on top of it
2Does the server ask the database again and again, or is it told?Every open screen adds database load, even idle ones
3Does an update carry "read again", or the new value itself?One ball makes every console re-read a whole schedule day
4How does a screen know it missed nothing?A lost update cannot be noticed today
5What happens to a screen with old code, or a broken stream?A screen can stop updating for good, silently

The bridge in one table Built today

The bridge is the one road between an app and omnium. An app runs in an iframe (a page inside the console page). The console is the host. The app and the host talk with postMessage, and only the host talks HTTP to the admin service.

MessageFromWhat it does
ready then helloapp, then hostThe app says its protocol and version; the host answers with the session: who, which subject, which reads and commands are offered
readappOne read of a named read model, for example games.schedule_day
watch / unwatchappA read model now and whenever it changes; the host sends push messages with new values
commandappOne change, sent through the workflow engine; see Commands
stateappSmall saved app state with a version (for example a desk's layout)
openappAsk the console to open a fixture, a player, a run or an integration
theme, session, session:endhostLight or dark, a fresh session, or "you are signed out / please update"

The message types are in packages-ts/omnium-bridge/src/protocol.ts:124-169. The protocol number is PROTOCOL = 1 (protocol.ts:8). The package is published as @fanos/omnium-bridge 1.0.0; the admin panel in the repo uses its source directly.

How it works

The drawing shows today on the left and the agreed design on the right.

A hand-drawn sketch titled Live updates: today vs agreed, split in two panels by a dashed line. Left panel, Today: every screen asks. Three white boxes, Console 1, Console 2 and Scorer B, each send an arrow labelled every 1 s into a blue cylinder labelled Postgres outbox. A red note under it says 70-80 queries a minute per idle screen. Right panel, Agreed: told once. A blue cylinder Postgres sends an arrow labelled NOTIFY to a yellow box Listener, one per server. An arrow labelled one read goes down to a yellow box Shared watch. From Shared watch three arrows labelled push go down to Console 1, Console 2 and Scorer B. A green note under it says same load for 5 or 500 screens.
Today every screen polls the outbox on its own. In the agreed design, Postgres wakes one listener per server, and one shared watch pushes to every screen.

Today Built today

  1. A command saves its change and one outbox row. The outbox is the domain_event table. The row is written in the same transaction as the change, so it exists exactly when the change does.
  2. Each open screen has its own stream. The admin service runs GET /stream/changes for each screen. Once a second it reads new domain_event rows, up to 500, and filters them in Python by event or by fixture.
  3. The stream sends "these areas changed". A frame names areas like fixture or medals, not data.
  4. The host waits 250 ms, then re-reads. It re-reads, one after another, every watch whose areas were touched. It pushes a new value to the app only if the value changed.
  5. The stream ends itself after 10 s. A proxy on AWS cuts every request at 15 s, so the server hangs up first. The browser dials again with its last place.

Agreed design Agreed, to build

  1. Told, not asking. Every command transaction ends with pg_notify('omnium_change', outbox_id). Postgres delivers it only when the transaction commits.
  2. One listener per server container. Each container keeps one dedicated connection that has run LISTEN omnium_change. A wake-up arrives 0.20 ms after the commit (median, measured).
  3. One read per container. On a wake-up, the container reads the new outbox rows once, for all its screens, with the late-commit-safe reader from Section 4 (L9). A safety read runs every 5 s in case a wake-up was lost.
  4. Shared watches. A screen tells the server what it watches: a read model and its parameters. The server keeps one copy per distinct watch, not one per screen. Fifty consoles on one schedule day are one watch.
  5. Push the value. When a change touches a watch, the server works it out once and pushes the new value to every screen watching it. No screen has to read again.
  6. Narrow updates. A read model says which record each row comes from. A change to one match updates that match's row, not the whole day.
  7. A version on every value. A screen that holds version 41 and receives 43 knows it missed 42, and re-reads that one record.
  8. A stream that never goes quiet. The client reopens the stream after any close, including an error answer. Heartbeats come every 15 s, and they carry the deployed app version and the bridge protocol version.

A worked example

Example · Two scorers on one volleyball match, and a console on the schedule

Example values. Scorer A and scorer B both have the volleyball desk open on the same match. Scorer B's app watches the desk's board read (<desk>.board), at version 205. A console operator has the day's schedule open (games.schedule_day). Scorer A records a point for the home side.

Today (built). Times are counted from the commit. "From the code" means the number is a constant in the code, not a measurement.

StepWhat happensTime after commitSource
1The command commits. Outbox row 88412 is saved with touches: ["fixture"]0pipeline.py:971-979
2Scorer B's stream, on its next poll, reads rows after its place and finds 884120 to 1 sfrom the code, _POLL_SECONDS = 1.0
3The stream sends event: changes with the area fixturesame pollbridge.py:826-856
4Scorer B's host waits for more changes to arrive togetherplus 250 msfrom the code, host.ts:196
5The host re-reads the board over HTTP and pushes it to the appplus one readnot measured
6The console's stream does steps 2 to 4 on its own, then re-reads every watch that depends on fixture: the schedule day, the day list, the event and the choices, one after another0.25 to 1.25 s plus four readsfrom the code; reads not measured

If scorer B's stream was in its 10 s hang-up at that moment, add up to 1 s more (_RETRY_MS = 1000). Between changes, each of the three screens still runs about 70 to 80 queries a minute.

Agreed design (to build).

StepWhat happensTime after commitSource
1The command commits. The same transaction ran pg_notify('omnium_change', '88412')0design, G1
2Each admin container's listener wakes0.20 ms median, 0.53 ms p99measured on a laptop, 6 Oct, 500 tries
3Each container reads the outbox once from its place, and finds 88412 for fixture Xone querynot measured
4The shared watch "board of fixture X" is touched by record id. It is worked out onceone readnot measured
5It is pushed as event: value with version 206 to scorer B (and to scorer A)networknot measured
6The watch "schedule day" lists fixture X among its records. Only that one row is worked out and pushed to the consoleone narrow readdesign, G6
7Scorer B's screen holds 205 and receives 206: no gap, so it applies the value. The point glows for a seconddesign, B5

The target is under 0.5 s for scorer B and under 1 s for the console (B8). That is an estimate from the measured parts, to be proven by the Section 12 live tests. Scorer B's next tap carries the new event number, so it cannot be applied on top of an old view (Section 4, L17).

If scorer B had received 207 instead, it would know it missed 206. The screen would show "Catching up" and re-read that one record.

A hand-drawn flow titled One point, two screens (agreed design). A white box Scorer A, with a person holding a tablet, sends an arrow labelled add point to a yellow box Command service, runs the command. An arrow labelled commit goes to a blue cylinder Postgres, save + NOTIFY. An arrow labelled wake 0.2 ms goes down to a yellow box Listener, in admin-api. An arrow labelled read outbox once goes left to a yellow box Shared watches, in admin-api. Two arrows labelled push go down to two white boxes: Scorer B, new board, v206, and Console, one row changed. A green note at the bottom: target: under 0.5 s.
One point in the agreed design: one commit, one wake-up, one read per server, and two pushes.

Low-level design

The code below is real unless it is marked "Agreed design, not in the code yet". How commands themselves are written is on Writing a workflow in code.

The outbox row Built today

Every accepted command writes one domain_event row in its own transaction. The scoring pipeline does it here:

# packages/core/src/omnium_core/scoring/pipeline.py:971-979
jobs_outbox.emit(
    session,
    jobs_outbox.COMMAND_ACCEPTED,
    sport_id=fixture.sport_id,
    competition_id=fixture.competition_id,
    fixture_id=fixture.id,
    seq=item.seq,
    payload={"command": code, "touches": ["fixture"]},
)

And the workflow engine (Games operations, Results) does it here:

# packages/core/src/omnium_core/workflows/engine.py:171-188
jobs_outbox.emit(
    session,
    jobs_outbox.COMMAND_ACCEPTED,
    sport_id=fixture.sport_id if fixture is not None else None,
    competition_id=fixture.competition_id if fixture is not None else event.id,
    fixture_id=fixture.id if fixture is not None else None,
    seq=item.seq,
    payload={
        "command": code,
        "workflow": workflow.code,
        "actor": actor.name,
        "event": str(event.id),
        # What a host re-reads for it: only watches that depend on these.
        "touches": touches(report.kinds),
    },
)

In plain words:

  • emit (packages/core/src/omnium_core/jobs_outbox.py:44) only calls session.add. The row is saved by the caller's commit, so a rolled-back command leaves no row.
  • touches is a list of areas: fixture, medals, entries, integration. A scoring command always says ["fixture"]. The engine works the areas out from the kinds of records the command wrote (workflows/records.py:94). If a kind is unknown, it says None, which means "re-read everything".
  • The table is domain_event (model in packages/contract/src/omnium_contract/models/scheduler.py:26). Columns used here: id (big serial), kind, competition_id, fixture_id, seq, payload (JSONB). It is insert-only and pruned after 30 days.

Today's stream: one polling loop per screen Built today

The constants first:

# packages/admin/src/omnium_admin/routers/bridge.py:766-771
_POLL_SECONDS = 1.0
_KEEPALIVE_POLLS = 15
_TREE_POLLS = 60
#: How soon a browser dials again after a stream ends. Short, because a stream
#: now ends on purpose every few seconds and the gap is time with no updates.
_RETRY_MS = 1000

The loop that runs for each open screen (trimmed):

# packages/admin/src/omnium_admin/routers/bridge.py:901-958 (trimmed)
async def body() -> AsyncIterator[bytes]:
    cursor = start
    polls = 0
    ends_at = clock.time() + limit if limit > 0 else None
    yield f"retry: {_RETRY_MS}\n: connected\n\n".encode() + _place(cursor)
    while True:
        async with factory() as session:
            if event_row is not None and polls % _TREE_POLLS == 0:
                ids = (await session.execute(subjects.event_tree(event_row.id))).scalars()
                tree = {*ids, event_row.id}
            rows = (
                await session.execute(
                    sa.select(DomainEvent.id, DomainEvent.kind, DomainEvent.fixture_id,
                              DomainEvent.competition_id, DomainEvent.payload)
                    .where(DomainEvent.id > cursor)
                    .order_by(DomainEvent.id)
                    .limit(500)
                )
            ).all()
        polls += 1
        if rows:
            cursor = int(rows[-1].id)
            changes = [c for c in map(_as_change, rows) if wanted(c, ...)]
            if changes:
                yield _frame(cursor, changes)
        ...
        if ends_at is not None and clock.time() >= ends_at:
            yield _place(cursor)      # hang up before the proxy does
            return
        if not rows:
            await asyncio.sleep(_POLL_SECONDS)

What each part does, and what it costs:

  • One query a second per screen. DomainEvent.id > cursor with LIMIT 500 runs every second, even when nothing changed.
  • Every screen reads every row. The query has no filter for the screen's event or fixture. wanted() (bridge.py:795) drops the rows the screen does not need, in Python, after they were read. During a big import every screen reads every row.
  • The event tree. For an event screen, the loop reads the list of competitions under the event when polls % 60 == 0. Because polls starts at 0 in each new stream, and a stream lasts 10 s, this read runs at the start of every stream, so about every 10 s.
  • The 10 s hang-up. ends_at comes from stream_max_seconds. Each new stream also checks the login and looks up the event by code.
  • Late commits are skipped. id > cursor assumes ids commit in order. They do not: a transaction that took id 100 can commit after one that took id 101. Once the cursor passed 101, row 100 is never sent. Section 4 (L9) agreed a late-commit-safe reader.
  • Keepalive. A : keepalive comment is sent every 15 polls. With 10 s streams, this line is not reached (from the code, not tested).

Counting from the code, an event screen that sees no changes runs, per 10 s stream: about 10 outbox reads, 1 tree read, 1 login check and 1 event lookup. That is about 13 queries per 10 s, or about 78 a minute. The design doc rounds this to 70 to 80 a minute. A fixture screen skips the tree and event reads.

The frame it sends carries areas and places, not data:

# packages/admin/src/omnium_admin/routers/bridge.py:826-856 (trimmed)
def _frame(seq: int, changes: list[dict[str, Any]]) -> bytes:
    body = {
        "seq": seq,
        "changes": [
            {"seq": change["seq"], "kind": change["kind"],
             "fixtureId": ..., "competitionId": ...,
             "workflow": change["workflow"], "command": change["command"],
             "touches": change["touches"]}
            for change in changes
        ],
    }
    return f"event: changes\nid: {seq}\ndata: {json.dumps(body)}\n\n".encode()

seq here is the stream's place in the outbox, not a version of any record. A screen cannot tell from it that it missed a value.

Why the stream ends after 10 s Built today

# packages/core/src/omnium_core/settings.py:259-267
# --- The admin panel's live stream (/stream/changes) ---------------------
# How long one stream stays open before the server ends it and the browser
# dials again, carrying on from the last change it saw. 0 keeps it open.
#
# Not 0 by default because of AWS: admin-api sits behind ECS Service Connect,
# whose proxy ends EVERY request after 15 seconds unless the service sets a
# timeout. A stream cut there shows "Reconnecting" and can drop a change.
# Ending first, cleanly, avoids both. Set 0 once that timeout is lifted.
stream_max_seconds: float = Field(default=10.0, ge=0)

The agreed design raises the per-request timeout on the admin service and sets stream_max_seconds = 0 (G5). The infra change is tracked in Section 14 (Deploy).

Today's client: the browser side Built today

The bridge's HTTP client opens the stream with the browser's EventSource:

// packages-ts/omnium-bridge/src/http.ts:244-257
changes(onChange) {
  const live = { event: options.live?.event, fixture: options.live?.fixture };
  const Source =
    options.EventSource === undefined
      ? (globalThis as { EventSource?: EventSourceConstructor }).EventSource
      : options.EventSource;
  if (Source) {
    // A browser's EventSource cannot send a header, so the token rides in
    // the address — for this one route only.
    const stream = new Source(
      `${base}/stream/changes` + query({ ...live, access_token: options.token?.() }),
    );
    stream.addEventListener("changes", (event) => onChange(touchesOf(event?.data)));
    return () => stream.close();
  }
  // ... otherwise poll with ?once=true every 20 s

The host turns "areas changed" into re-reads:

// packages-ts/omnium-bridge/src/host.ts:188-228 (trimmed)
changed(touches: string[] | null = null): void {
  if (touches === null) this.everything = true;
  else if (touches.length === 0) return;
  else for (const area of touches) this.touched.add(area);
  if (this.settleTimer) clearTimeout(this.settleTimer);
  this.settleTimer = setTimeout(() => {
    this.settleTimer = null;
    void this.refresh();
  }, this.options.settleMs ?? 250);
}

refresh(): Promise<void> {
  ...
  for (const [watchId, watch] of [...this.watches]) {
    if (!everything && watch.depends && !watch.depends.some((area) => touched.has(area))) {
      continue;
    }
    const { value, depends } = await this.readWith(watch, watch.name, watch.params);
    const text = JSON.stringify(value);
    if (!this.watches.has(watchId) || text === watch.last) continue;
    watch.last = text;
    this.post({ type: "push", watchId, value });
  }
  ...
}

The console wraps EventSource to show the live dot:

// frontend-admin/src/lib/bridge/frame.ts:194-206
constructor(url: string) {
  this.source = new Source(url);
  this.source.addEventListener("open", () => {
    this.settle();
    onState("open");
  });
  this.source.addEventListener("error", () => {
    if (this.pending !== null) return;
    this.pending = setTimeout(() => {
      this.pending = null;
      onState("retrying");
    }, graceMs);
  });
}

In plain words:

  • The URL is fixed at the start. The login token is put in the address once. The browser's own reconnects reuse that address, with the same token.
  • The re-reads are one after another. refresh() awaits each watch in turn. A console with four schedule watches does four reads in a row for one ball.
  • An error only changes the label. After 5 s of error (RECONNECT_GRACE_MS = 5_000, frame.ts:170), the dot says "Reconnecting". Nothing opens a new stream. By the browser's rules, an answer that is not 200 (an expired login gives 401, a proxy can give 502) closes an EventSource for good. So the screen says "Reconnecting" forever and never updates again.
  • Old code is checked once. The host checks the protocol number only when the app says ready (host.ts:261-264). The app also sends its own version in ready, but nothing compares it, and nothing tells an open app that a new version was deployed.

The wake-up Agreed, to build

Agreed design, not in the code yet. Searching the repo for pg_notify, LISTEN and add_listener in packages/*/src finds nothing today.

-- Agreed design, not in the code yet (from the Section 8 doc).
-- In every command transaction, right after the outbox insert (Section 4):
INSERT INTO domain_event (...) VALUES (...) RETURNING id;
SELECT pg_notify('omnium_change', $outbox_id::text);   -- delivered only if this transaction commits
COMMIT;

The listener, one per admin container:

# Agreed design, not in the code yet: admin/src/omnium_admin/live/listener.py (new)
async def run(dsn: str, on_change: Callable[[], Awaitable[None]]) -> None:
    while True:
        try:
            conn = await asyncpg.connect(dsn)                 # outside the pool, never in a transaction
            wake = asyncio.Event()
            await conn.add_listener("omnium_change", lambda *_: wake.set())
            await on_change()                                  # catch up from the last place after (re)connecting
            while True:
                try:
                    await asyncio.wait_for(wake.wait(), timeout=5.0)   # the safety read every 5 s
                except asyncio.TimeoutError:
                    pass
                wake.clear()
                await on_change()                              # one read for any number of notifications
        except (OSError, asyncpg.PostgresError):
            await asyncio.sleep(backoff())                     # reconnect with growing waits

What matters in it:

  • Its own connection, outside the pool. A listener inside a transaction receives no notifications, and a long transaction that listens can hold up Postgres's notification queue.
  • The wake-up is only a doorbell. The outbox is the truth. A lost notification costs at most 5 s, because the safety read finds the row anyway.
  • Many notifications, one read. wake is a flag. Ten commits in one moment cause one read.
  • Reconnect reads from the last place. After a dropped connection, on_change() runs before waiting again, so nothing is missed.

Shared watches Agreed, to build

Agreed design, not in the code yet. Today each browser's host keeps its own watches (packages-ts/omnium-bridge/src/host.ts:289-302). In the design, the server keeps them, once per distinct view:

# Agreed design, not in the code yet: admin/src/omnium_admin/live/watches.py (new)
# per container, in memory
watches: dict[WatchKey, Watch]            # WatchKey = (read name, canonical params)
class Watch:
    value: bytes                          # canonical JSON of the last value
    version: int                          # the highest record version inside it
    depends: set[str]                     # areas, as read models declare today
    records: set[uuid.UUID]               # which records it shows, for narrow updates
    screens: set[StreamId]

The rules (G2, G3):

RuleValueWhy
Keyread name plus canonical parametersTwo screens asking the same thing share one watch
Touched byarea (depends) and record id (records)One match change touches only views that show that match
Worked outat most 8 at a time per containerA big import cannot flood the database
Pushedonly if the new value differsNo noise for screens
Dropped30 s after its last screen leavesA quick reload keeps the watch warm
Send queue per streamat most 100 valuesA full queue is cleared and the screen told resync, so one slow browser never slows the others
Memoryabout 50 MB per container for 1,000 views of about 50 KBestimate, from the design doc

Narrow updates need one change to read models (G6). Today the schedule reads say only an area:

# packages/core/src/omnium_core/workflows/reads_schedule.py:295, 338, 392
@read_model("games.schedule_days", depends=("fixture",))
@read_model("games.schedule_day", depends=("fixture",))
@read_model("games.schedule_choices", depends=("fixture",))

So any change to any match re-reads all of them. In the design, each read model also declares which records each row comes from. depends stays. The registry is ReadModel in packages/core/src/omnium_core/workflows/reads.py:52-57.

The new stream Agreed, to build

Agreed design, not in the code yet. /stream/changes (bridge.py:859) is replaced by three routes:

GET  /stream?subject=event:asian-games-2026            -> text/event-stream, first frame: "hello" with stream id, app version, protocol version
POST /stream/{id}/watch   {name, params}               -> {watchId, version, value}   (the first value at once)
POST /stream/{id}/unwatch {watchId}

event: value      id: <outbox place>   data: {"watchId": "w7", "version": 205, "value": {...}}
event: resync     data: {"watchId": "w7"}              -> the screen reads that watch again
event: heartbeat  data: {"app": "2026.10.06-3", "protocol": 2}     every 15 s

How the parts fit:

  • Watches are registered over HTTP on an open stream (G4). The stream stays one-way. No WebSockets are needed.
  • id: is the outbox place. The browser sends it back as Last-Event-ID when it reconnects after a network error.
  • version is per record. It is what lets a screen notice a gap (B5).
  • The client wrapper reopens after any close (B6, G5). For an error answer, which the browser never retries, the wrapper opens a new stream itself, with a fresh login from the panel and its last place. It then sends its watches again. Waits between tries grow.
  • 30 s of silence counts as broken. Heartbeats come every 15 s; two missed heartbeats trigger a reopen.
  • Heartbeats carry versions (B7). A newer app shows a quiet "A new version is ready" banner with Reload. A protocol the old code cannot read shows "Please reload", and the panel stops sending that app's commands.

Changes on the client side, from the design doc's file list:

FileChange
packages-ts/omnium-bridge/src/http.tschanges() becomes stream(): reopens after any close, resends watches, applies value frames
packages-ts/omnium-bridge/src/host.tsStops re-reading watches itself; forwards value frames to the app; asks for a resync on a version gap
frontend-admin/src/lib/bridge/frame.tsThe live dot gains "Catching up"; reportingEventSource gives way to the reopening stream

New server files

All new, from the design doc's file list:

FileHolds
core/.../outbox/notify.pynotify(session, outbox_id): the pg_notify call, used by the engine right after the outbox insert
core/.../outbox/reader.pyread_after(): the late-commit-safe read (Section 4), shared with the stats queue
admin/.../live/listener.pyThe listener above: LISTEN, reconnect, the 5 s safety read
admin/.../live/watches.pyThe shared watch map: touch by area and record, work out at most 8 at a time, push, drop after 30 s
admin/.../live/streams.pyStream ids, send queues (100), resync, heartbeats with app and protocol versions

What a screen shows

A hand-drawn state diagram titled What a screen shows (agreed design). A green box Live has an arrow labelled version gap to a yellow box Catching up, which returns to Live with record re-read. Live has an arrow labelled 30 s silence or error down to a yellow box Reconnecting, which returns to Live with stream reopened. Reconnecting has an arrow labelled protocol too old to a red box Please reload. A dashed arrow labelled newer app deployed goes from Live to a white box New version banner.
The live dot's states in the agreed design. Today only Live, Reconnecting and Connecting exist.
StateWhat the person seesWhen
LiveA small green dot by the page title: "Live"The stream is open and heartbeats arrive
Catching upAmber: "Catching up"The screen saw a version gap and is re-reading
ReconnectingAmber: "Reconnecting", with the time since the last updateNo heartbeat for 30 s; shown only after 5 s, so short blips stay quiet
OfflineA bar: "Offline. Showing data from 12:04. Changes will send when you are back"No network; the app's outbox keeps changes (Section 7, V1)
New versionA quiet banner: "A new version is ready" with ReloadA newer app was deployed; waiting changes survive the reload
Must reload"Please reload to keep working"The bridge protocol changed in a way the old code cannot read
Updated by someone elseThe changed value glows for a second; on a match page, "Updated by Priya, 12:04"A push changed a value on screen, never a field being typed in (Section 7, V4)

Today the dot has four labels: "Live", "Reconnecting", "Connecting" and "Checking every 20 s" when there is no EventSource (frame.ts:151-158).

The public API stream is a different thing

GET /v1/stream in the public API (packages/api/src/omnium_api/routers/stream.py:115-151) is for fans, not staff screens. It subscribes to a Redis pub/sub channel and sends {entity, ids, seq}, with a 15 s heartbeat. Redis pub/sub keeps no replay, so a client that missed a message must re-fetch. It answers 503 when no Redis is set. Section 8 does not change it; fan delivery is on Stats, feeds and delivery.

When things go wrong

Every case ends with each screen showing the latest data, or saying clearly that it cannot.

What goes wrongWhat happens (agreed design)Decision
A wake-up is lost (the listening connection dropped for a moment)The 5 s safety read finds the new outbox rows. Screens are at most 5 s late, never wrongB1, B2
The listening connection breaksThe container reconnects, runs LISTEN again, and reads everything after its last placeB1, G1
Two commands commit out of id orderThe late-commit-safe reader (Section 4, L9) still finds the late one. Today id > cursor skips it for goodB2
A stream gets an error answer (expired login, a proxy's 502)The client opens a new stream with a fresh login and its last place. Today this screen goes quiet for goodB6
A screen misses one valueThe next value's version skips one. The screen shows "Catching up" and re-reads that one recordB5
A screen reconnects to a different containerIt sends its watch list again; the new container adds them to its shared watchesB3
A big import changes 1,000 matchesEach container reads the outbox once. Screens watching those matches get their rows; others get nothing. At most 8 watches are worked out at a timeB2, B4, G2
300 people open the console during a finalDatabase work stays per container. Each extra screen costs memory and one open connection to the serverB3
One browser stops readingIts send queue fills at 100, is cleared, and the screen is told resync. Other screens stay on timeG3
A deploy while screens are openStreams reconnect to the new containers within seconds, from their last place. Apps see "A new version is ready"B6, B7
A container diesIts screens reconnect elsewhere and resend their watches. Nothing is lost: the outbox is the truthB3, B6
An app's code is too old for the new bridge protocolThe panel shows "Please reload" and stops sending its commands until it doesB7
Very many commits a second all send NOTIFYA commit that sends a NOTIFY takes a global lock in Postgres. Others hit this at "tens of thousands of simultaneous writers" (Recall.ai, in the research notes). Our write rate is far below that (not measured for omnium)B1

Decisions

All 14 decisions are Agreed (closed 6 Oct 2026). None was dropped.

#DecisionIn plain words
B1Every command transaction sends a NOTIFY with its outbox id; each server container keeps one listening connection, plus a safety read every 5 sThe database wakes the servers within a millisecond of a save, instead of every screen asking every second
B2Each container reads new outbox rows once, with the late-commit-safe reader (Section 4, L9), for all its screensDatabase work no longer grows with the number of open screens
B3Shared watches on the server: one copy per distinct watched view per container, worked out once per change and pushed to every screen watching itFifty consoles on one schedule cost one read
B4A push carries the new value of the view, not only "something changed"; a match row updates only that rowThe second scorer sees the first scorer's ball within about 0.5 s, with no extra read
B5Every pushed value carries its record's version; a skipped number makes the screen re-read that one recordA screen always knows if it missed something
B6The client reopens the stream after any close, including error answers, with a fresh login and its last place; the proxy timeout is raised; heartbeats every 15 sNo screen silently stops updating
B7Heartbeats carry the deployed app version and the bridge protocol version: "A new version is ready" by default, "Please reload" only when the protocol changed incompatiblyOld code is noticed without interrupting someone mid-task
B8Targets: another staff screen updated within 1 s of a save, a second scorer within 0.5 s, proven in the Section 12 tests"Live" has a number we test
G1pg_notify(outbox id) in every command transaction; one dedicated listening connection per container, never in a transaction, reconnecting and reading from its last placeThe wake-up can be lost without losing a change
G2Shared watches keyed by read name and canonical parameters, touched by area and by record id, worked out at most 8 at a time, pushed only if changed, dropped 30 s after the last screenWork grows with what is watched, not with how many people watch it
G3Each stream has a send queue of at most 100; a full queue is cleared and the screen told to resyncOne slow browser never slows the others
G4Watches are registered over HTTP on an open stream; values come back on the stream with their version and outbox placeA simple one-way stream; no WebSockets needed
G5The client wrapper reopens after any close and resends its watches; the proxy timeout is raised and stream_max_seconds set to 0No silent screens, no 10 s gaps
G6Read models declare which records each row comes from, so a change to one match updates only its rowOne ball no longer re-reads a whole day

Built today, or still to build

PieceTodayAgreed designStatus
Bridge protocol between app and hostprotocol.ts:8 (PROTOCOL = 1), host.ts, app.ts; published as @fanos/omnium-bridge 1.0.0Stays; protocol version also goes into heartbeatsBuilt today
Outbox row per accepted commandpipeline.py:971-979, engine.py:171-188, jobs_outbox.py:44Stays; a NOTIFY is added after itBuilt today
Wake-up by NOTIFYNone (no pg_notify or LISTEN in the code)outbox/notify.py plus one listener per container (B1, G1)Agreed, to build
Reading the outboxEach screen polls once a second, bridge.py:910-958One late-commit-safe read per container (B2)Agreed, to build
FilteringIn Python after reading all rows, wanted() at bridge.py:795By area and record id in the shared watch map (G2)Agreed, to build
WatchesKept in each browser's host, re-read one after another, host.ts:188-228Kept on the server, once per distinct view (B3)Agreed, to build
What a push carriesAreas only, _frame at bridge.py:826-856The new value and its version (B4, B5)Agreed, to build
Narrow row updatesSchedule reads depend on fixture as a whole, reads_schedule.py:295, 338, 392Records per row (G6)Agreed, to build
Reopen after an error answerOnly relabels the dot, frame.ts:200-206Client wrapper reopens after any close (B6, G5)Agreed, to build
Stream lengthEnds every 10 s, settings.py:267stream_max_seconds = 0 after the proxy timeout is raised (Section 14)Agreed, to build
Old code noticedProtocol checked once at ready, host.ts:261-264; app version not comparedApp and protocol versions in every heartbeat (B7)Partly built
Live dotLive, Reconnecting, Connecting, Checking every 20 s, frame.ts:151-158Adds Catching up, New version, Must reloadPartly built
Speed targets and testsNone0.5 s and 1 s, proven by the Section 12 live tests (B8)Agreed, to build
Public fan streamRedis pub/sub, api/routers/stream.pyNot part of Section 8Built today

The design doc's build order puts the cheapest big win first:

  1. notify() in the engine after every outbox insert. Harmless while nobody listens.
  2. The listener and the late-commit-safe reader in each admin container, feeding today's /stream/changes. The per-screen polling stops here.
  3. The client wrapper that reopens after any close (fixes silent screens), and the proxy timeout.
  4. Shared watches and the new stream routes; the host forwards values instead of re-reading.
  5. Record keys on the schedule read models (narrow row updates).
  6. Versions on values, gap detection, the Catching up state.
  7. App and protocol versions in heartbeats; the "new version" banner.

Numbers

NumberWhatWhere it comes from
0.20 msMedian time from a commit to a listener wakingMeasured on a laptop, 6 Oct 2026, 500 tries (Section 8 doc)
0.53 ms99th percentile of the sameSame measurement
70 to 80 a minuteQueries per idle open screen todayCounted from the code's constants (Section 8 doc); this page's own count gives about 78 for an event screen
3,500 to 4,000 a minuteQueries for 50 idle screens todayThe line above times 50
12 a minuteSafety reads per container in the agreed designOne every 5 s, from the design
1 s, 250 ms, 10 s, 5 sPoll gap, host wait, stream length, "Reconnecting" graceFrom the code: bridge.py:766, host.ts:196, settings.py:267, frame.ts:170
0.25 to 1.25 s plus re-readsToday's time from a save to another screenFrom the code's constants; re-read times not measured
under 0.5 sAgreed design, save to second scorerEstimate from the measured parts; to be proven in Section 12
about 50 MBMemory per container for 1,000 watched views of 50 KBEstimate, Section 8 doc

The tests that will prove the design (Section 12):

TestProves
Two browsers on one match; one adds a ballThe other shows it within 0.5 s, with no read request in its network log (B4, B8)
200 idle screens open for 5 minutesDatabase queries per minute stay flat as screens are added (B2)
Kill the listening connectionScreens catch up within 5 s; nothing missed (B1)
The stream answers 502 onceThe client reopens and carries on (B6)
Drop one value frame on purposeThe screen sees the gap and re-reads that record (B5)
Change one match on a 900-unit dayOnly that row is sent (G6)
A slow browser that stops readingIts queue resyncs; other screens stay on time (G3)
Deploy a new app version with screens openThe banner appears; nothing waiting is lost (B7)