Start here · One change, end to end

Life of a score

One football goal, followed from the scorer's tap to the client's file, through every part of omnium.

Covers
Every part, for one goal, in 13 stops
Based on
Design sections 2 to 11, and the code on 7 Oct 2026
Main tables
command, fixture, scoreboard, timeline_item, domain_event, feed_document, delivery_state
Read time
about 30 minutes

In one minute

This page follows one goal through the whole system. A scorer taps Goal in the 63rd minute of a Games football match. Thirteen stops later, the goal is in a file on the client's server, in Redis for Google, on a fan's phone, in the stats, and every step of the way is in the logs under one trace id.

The main story is the agreed design: what we are building, from design sections 2 to 11. After each stop, a Today line says what the code on 7 Oct 2026 does instead, with the file and line, and a status tag.

The short version: the goal is saved once, in one transaction, on the main database. One wake-up tells everyone else. Every reader after that works from the saved change, never from the request, so nothing is lost, doubled or skipped. The target from save to client file is 10 seconds.

A second, shorter story at the end shows what happens when a Scout import brings a different score two minutes after a staff fix.

1.48 ms
whole command transaction, median (laptop, 5 Oct)
0.20 ms
commit to wake-up, median (laptop, 6 Oct)
0.5 s
target: the second scorer sees the goal
10 s
target: save to client file (D6)

What this page follows

One change, through every part, in the order it happens. If you read only one page after The big picture, read this one. Each stop links to the page that explains that part in full.

Example · The match, in example values

All ids, names and times below are examples, made up for teaching. Table names, column names, command codes and file paths are real.

WhatExample value
The matchFootball, India v Japan, a group match at a Games. Fixture id 01a1d4e2-…
The moment63rd minute, 19:42:10.000. India scores; the score goes from 0–0 to 1–0
Scorer ATaps Goal on the football desk (scorer app). Their screen last showed event 56
Scorer BScores the same match on a second desk. Also showed event 56
The consoleAn operator has the day's schedule open
The commandfootball.goal, key 7f3a2c1e-…, input: side IND, kind open_play, scorer IND-09, assist IND-17, minute 63
The trace id4bf92f35…
ClientsGoogle (pulls from Redis), NDTV (S3), News18 (FTP), DailyHunt (SFTP)

How to read the tags after each Today line:

TagMeans
Built todayThe code does this today, as designed
Partly builtThe code does part of it, or does it a different way
Agreed, to buildAgreed design, not in the code yet

How it works

A hand-drawn map titled One goal, 13 stops. 1, the scorer app (outbox, key 7f3a), sends the goal down to 2, the Bridge, which passes it to 3, the Engine: lock, check, save. A curved arrow, 4 accepted, goes back to the scorer app. The engine saves to Postgres (main): log, state, outbox. 5, the wake-up (NOTIFY), fans out to 6, Scorer B and the console; 7, the delivery-worker, which builds files; 10, stats; and 12, the alert service. From the delivery-worker, arrows go to 8, Redis (Google pulls), and 9, one queue per client (NDTV, News18, DailyHunt). A dashed copy arrow goes from Postgres to 11, the read-only copy for fans. A band along the bottom reads 13, trace id 4bf92f35 on every step.
The agreed design. One road in (stops 1 to 4), one wake-up (5), many readers out (6 to 12), one trace id through all of it (13).

The 13 stops, in one line each:

  1. The tap. The scorer app saves the goal on the device, in its outbox, with a key made once.
  2. The request. The bridge sends one HTTP command to omnium, with the key and the last event the screen saw.
  3. The engine. Lock the match row, check the key and the version, run the sport's rules, save everything in one transaction.
  4. The answer. "Accepted, event 57", sent only after the commit.
  5. The wake-up. The commit fires a Postgres NOTIFY; one listener per server wakes.
  6. Live screens. Scorer B and the console get the new value pushed. Targets: 0.5 s and 1 s.
  7. The delivery-worker builds. Finds the files that contain this match, builds them from the main database, and saves a new version only if the content changed.
  8. Redis and Google. The new file goes into Redis, write-if-newer. Google's next pull gets it.
  9. Client queues. Each client destination sends the file on its own queue, and records proof of delivery.
  10. Stats. The stats queue works out the heavier numbers: top scorers, the group table.
  11. Fans. The public API answers from Redis, or from the read-only copy, never older than 5 s.
  12. The alert service. Watches every stop, and tells one person once if any stop fails.
  13. The trace id. One id on every log line and row, from the tap to the file.

A worked example: one goal, stop by stop

Each stop has the agreed design first, then the Today line.

1. The tap: the outbox gives it a key

19:42:10.000. Scorer A taps Goal, picks India, the scorer and the assist, and types 63. The desk shows 1–0 at once, with a small "sending" mark. The agreed target for click to screen is under 0.1 s (V6).

Before anything goes over the network, the shared outbox in the bridge package writes one row to IndexedDB, the browser's own small database. It survives a crash, a reload and a new deploy.

// The outbox row for this goal. Shape from the design doc, Section 7 (new; not in the code).
{
  key: "7f3a2c1e-…",            // crypto.randomUUID(), made once, sent on every retry (C2)
  subject: "fixture:01a1d4e2-…", // one send loop per match, oldest first (C12)
  code: "football.goal",
  input: { entry: "IND", kind: "open_play", scorer: "IND-09", assist: "IND-17", minute: 63 },
  saw: 56,                       // the last play event this screen showed (L17)
  createdAt: "2026-10-07T14:12:10.000Z",
  tries: 0,
  state: "waiting"
}

The key is the most important value on this page. Every retry sends the same key, so the server can tell "this goal again" from "a second goal".

Today Partly built: there is no outbox. The bridge makes a new key on every call unless the caller passes one (packages-ts/omnium-bridge/src/app.ts:332), and nothing resends by itself. A person who taps again after a timeout sends a second command. No football desk exists among today's desks (searched boxing-scoring-ui/src for football, 7 Oct); the football.goal command does exist (packages/flows/src/omnium_flows/football/commands.py:57-68).

2. The request: one command through the bridge

The desk runs inside an iframe of the admin panel. It posts a command message to the panel, and the panel sends the HTTP request with the person's sign-in. The app never sees a token or omnium's address.

POST /workflows/football/subjects/fixture/01a1d4e2-…/commands
traceparent: 00-4bf92f35…-…-01

{
  "code": "football.goal",
  "key": "7f3a2c1e-…",
  "input": { "entry": "IND", "kind": "open_play", "scorer": "IND-09", "assist": "IND-17", "minute": 63 },
  "saw": 56,
  "sentAt": "2026-10-07T14:12:10.000Z",
  "source": "desk-1"
}
FieldWhy it is there
codeWhich command: the registry finds football.goal in the flows plugin
keyMakes a retry safe (C2)
inputChecked against GoalInput; kind must be one of open_play, penalty, free_kick, own_goal
sawThe play version: the last event the sender saw (L17). New
sentAtWhen the device made it, saved as timeline_item.sent_at (M7). New
sourceWhich desk, when a match has two
traceparent headerThe W3C trace id the bridge accepts or makes (A9). New

Network time from India: about 200 ms for the round trip to US East (Section 1), or 20 to 60 ms to an AWS region in India (Section 4). Both are estimates from the design docs. Which region prod uses was not checked for this page. The clock below uses the 200 ms case, so the request arrives at about 19:42:10.100.

Today Partly built: the route is built: send_command at packages/admin/src/omnium_admin/routers/bridge.py:213-224, body CommandIn at bridge.py:183 with code, input, an optional key, dry_run and source. A scoring workflow goes to _score, which calls pipeline.command (bridge.py:262-297). There is no saw, no sentAt and no traceparent.

3. The engine: key, lock, version, rules, save

The command service (its own ECS service, L15) runs the one road. For an events match, the middle of the road is the step: the saved state plus this one goal gives the new state. This is the whole transaction, statement by statement. It joins the command road from Section 4, the ball transaction from Section 5 and the wake-up from Section 8.

-- 0. Fast path, before any lock: a finished retry is answered at once (L3).
SELECT outcome, answer, payload_hash FROM command WHERE command_id = '7f3a2c1e-…';   -- no row

BEGIN;                                                    -- READ COMMITTED; limits come from the pool (L16)
-- 1. The match lock and its log number, in one read (L1).
SELECT last_seq FROM fixture WHERE id = $match FOR NO KEY UPDATE;                  -- 56
-- 2. The key again, under the lock. This is the check that counts (L3).
SELECT outcome, answer, payload_hash FROM command WHERE command_id = '7f3a2c1e-…';   -- no row
-- 3. The saved state: one row, never the whole log (S1, M1).
SELECT last_seq, code_sha, snapshot FROM scoreboard WHERE fixture_id = $match;
--    saw 56, last_seq 56: nothing missed. code_sha matches the loaded worker (S6).

-- In the football worker process, no SQL: check(state, event), then step(state, event).

-- 4. The writes. Seq 57 is the new log number.
INSERT INTO timeline_item (fixture_id, seq, unit, event_type, payload, sent_at, received_at, ...)
     VALUES ($match, 57, 'event', 'football.goal', $input, $sent_at, now(), ...);
UPDATE scoreboard SET snapshot = $state, last_seq = 57, code_sha = $sha, state_hash = $h, updated_at = now()
 WHERE fixture_id = $match;
INSERT INTO stat_value (fixture_id, subject_id, subject_role, segment, occurred_on, "values", last_seq, ...)
     VALUES (...)                                          -- only the lines this goal changed (M3)
ON CONFLICT (fixture_id, subject_id, subject_role, segment, occurred_on)
  DO UPDATE SET "values" = EXCLUDED."values", last_seq = EXCLUDED.last_seq;
INSERT INTO published_document ... ;                       -- the match document clients read
INSERT INTO domain_event (..., trace_id) VALUES (..., '4bf92f35…') RETURNING id;   -- the outbox note
SELECT pg_notify('omnium_change', $outbox_id::text);       -- delivered only if we commit (B1)
INSERT INTO command (command_id, sender, subject_id, payload_hash, outcome, answer, seq, ...)
     VALUES ('7f3a2c1e-…', 'desk-1', $match, $hash, 'accepted', $answer, 57, ...);
UPDATE fixture SET last_seq = 57 WHERE id = $match;
COMMIT;                                                    -- synchronous_commit = on

What each part does, in plain words:

  • The lock is the match's own row. FOR NO KEY UPDATE makes a second command on this match wait, but lets plain reads and inserts of child rows through. Another match never waits. Scorer B's goal, if B taps too, waits here for a few milliseconds.
  • The key is checked twice. Once without a lock, to answer most retries at once (0.09 ms at the median, 10 million rows, laptop). Once under the lock, so two copies of one command can never both run.
  • The version check is the play level. The sender saw event 56, and the match is at 56. If scorer B's goal had landed first as 57, this command would roll back and answer conflict, with event 57 inside (L17).
  • The rules run on the saved state. half_is_open today reads one flag from the state. In the design, the sport's check(state, event) does the same in a worker process that can be killed at a limit (L10, M6).
  • One step, not a replay. step(state, event) returns the new state, the lines it changed and any emits. The state stays under 16 KB (M2).
  • Everything commits together. The log row, the state, the lines, the published document, the outbox note, the wake-up and the command row. Either all of them are saved, or none.

The worker call is one request and one reply:

{"id": "7f3a", "op": "check_and_step", "sport": "football", "state": {}, "events": [{}]}
{"id": "7f3a", "ok": true, "refusal": null, "step": {"state": {}, "lines": [], "emits": []}, "cpu_ms": 0.08}

The message shape is from the design doc; the bodies are left empty here, and cpu_ms is the design doc's own example value.

Two trips, not a dozen (M9). As written above, the transaction is about a dozen statements; Sections 4 and 5 counted about 9 trips for their shorter versions. The agreed form sends it in 2 trips: one statement that locks and reads the key, the state and the versions together, and one call to a database function that does every write and the commit. On a laptop, the Section 4 command transaction took 1.48 ms at the median and 3.98 ms at the 99th percentile. On prod, about 2 to 3 ms is expected (estimate). The target for one event is under 5 ms of server time at any point in a match (S9).

The goal commits at about 19:42:10.103 on the example clock.

Today Partly built: the live path is pipeline.command (packages/core/src/omnium_core/scoring/pipeline.py:845). It sets two SET LOCAL limits (pipeline.py:841-842), takes an advisory lock on the first 8 bytes of the fixture id (pipeline.py:869), looks for the key in timeline_item (pipeline.py:879, ledger.py:299), reads the whole match log (pipeline.py:897), appends the row, then refolds the whole log again (pipeline.py:686). Football's state action reads the whole log (flows/football/actions.py:34). It writes the scoreboard, the published document (pipeline.py:965) and the outbox note (pipeline.py:971-979). There is no command table, no version check and no NOTIFY. Live scoring desks were not on prod during the Games.

4. The answer: only after the commit

19:42:10.200. The answer reaches the desk. The sending mark turns into a tick for 2 s, then the field goes back to idle. The outbox deletes the row. The agreed target for save to tick is under 1 s (V6).

{ "outcome": "accepted", "seq": 57 }

The design fixes the outcome field and its five values, and an accepted answer carries the new seq. The rest of the body is not fixed yet. The other four answers, and what the desk does with each:

OutcomeHTTPWhat the desk does
accepted200Tick; delete the outbox row
duplicate200The same as accepted: this key was saved before
refused422Puts the old value back and shows the rule's sentence, for example "No half is open"
conflict409Shows the events it missed, asks the scorer to check, pauses that match's queue
try_again503Keeps the row; sends again with the same key after 0.2 s, 0.4 s, 0.8 s, and so on

If the answer is lost in a stadium tunnel, the desk resends with key 7f3a2c1e-…. The quick check finds the command row and answers duplicate, with the first answer. One goal, not two.

Today Partly built: the answer can go out before the commit. The commit runs in the teardown of the FastAPI dependency get_write_session (packages/admin/src/omnium_admin/deps.py:28), which FastAPI 0.139 runs after the response is sent (checked on Commands, 7 Oct). Answers have two shapes: HTTP 200 with accepted: false, or an HTTP error from _fail (bridge.py:77-91).

5. The wake-up: NOTIFY, then one read per server

19:42:10.103 plus 0.2 ms. The commit delivers the pg_notify('omnium_change', …). Every server container that keeps a listening connection wakes. Measured on a laptop on 6 Oct: 0.20 ms at the median, 0.53 ms at the 99th percentile, 500 tries.

# Agreed design, not in the code yet: admin/src/omnium_admin/live/listener.py (new). Trimmed.
await conn.add_listener("omnium_change", lambda *_: wake.set())
await on_change()                                  # catch up from the last place first
while True:
    try:
        await asyncio.wait_for(wake.wait(), timeout=5.0)   # safety read every 5 s
    except asyncio.TimeoutError:
        pass
    wake.clear()
    await on_change()                              # one read for any number of wake-ups

The wake-up is only a doorbell. The outbox row is the truth. On a wake-up, each container reads new domain_event rows once, for all its screens, with the late-commit-safe reader: it reads only rows from transactions older than the oldest one still running, so a slow transaction that commits late is never skipped (L9).

The same doorbell wakes the delivery-worker (stop 7), the stats queue (stop 10) and, through its own channel, the alert service (stop 12).

Today Agreed, to build: there is no pg_notify or LISTEN anywhere in packages/*/src. Each open screen polls domain_event once a second with id > cursor (bridge.py:766, bridge.py:924), which can skip a late commit.

6. Live screens: scorer B and the console

The container's shared watches see that the outbox row touches fixture 01a1d4e2-…. The watch "board of this match" is worked out once and its new value is pushed to every screen that watches it. The schedule watch lists this match among its records, so only that one row is worked out and pushed to the console (G6).

event: value   id: <outbox place>   data: {"watchId": "w7", "version": 57, "value": {...}}

Scorer B's screen held the version before and receives the next one (57 here is an example version): no gap, so it applies the value. The goal glows for a second. If a version were skipped, the screen would show "Catching up" and re-read that one record (B5). Scorer B's next tap carries "saw 57", so it can never be saved on top of a view that missed this goal.

Targets (B8): under 0.5 s from the save to scorer B, under 1 s to any other staff screen. These are estimates from the measured parts, to be proven by the Section 12 tests.

Today Partly built: it works, more slowly. The screen's stream finds the row on its next 1 s poll (_POLL_SECONDS, bridge.py:766) and sends only "the area fixture changed". The host waits 250 ms (packages-ts/omnium-bridge/src/host.ts:196), then re-reads every watch that depends on fixture, one after another. From the code's constants: 0.25 to 1.25 s plus the re-reads, which were not measured.

7. The delivery-worker builds: only the touched files, only if changed

The delivery-worker is woken by the same doorbell. It reads the outbox with the same safe reader and asks which files contain this match:

SELECT doc_key FROM feed_dependency
 WHERE record_kind = 'fixture' AND record_id = $match;

In this example two files come back: the result details file and the calendar of this Games. Which files hold a live score depends on each feed's definition; the two here are examples. The medal table, team medals and the rest are not touched.

Then, per file:

  1. Wait a moment. A 1 s timer per file gathers a burst. It never runs more than 3 s from the first change (D2), so a steady stream of goals cannot hold a file back.
  2. Build from the main database. The delivery-worker calls the feed resolvers directly, in its own process, never through the public API and never from the read-only copy (D1). A copy could be a few milliseconds behind and miss this very goal.
  3. Hash the body, without its time stamp. Same as the stored content_hash: stop, nothing is saved or sent. Different: save (D3, E1).
  4. Save in one transaction: the body, the new content_hash, changed_at = the goal's commit time (never "now"), the next seq, and the file's feed_dependency rows (E2). Say the new versions are 9120 and 9121.
  5. After the commit, put both new versions on every destination queue that wants them, Redis included.

Cost, for scale: hashing the 4.3 MB calendar took 1.74 ms (measured, design doc). Building that calendar took 1.3 s (a comment in the code, api app.py:157-165). The result file is smaller; its build time was not measured. The agreed budget is a build under 3 s and a send under 5 s, inside the 10 s target (D6).

Today Partly built: after the commit, dirty.flush (deps.py:32) fires the invalidation listeners, and the feed listener in the admin service arms a 1 s debounce (feed_debounce_seconds, packages/core/src/omnium_core/settings.py:143) with no maximum wait. The pass then rebuilds all 7 files of each changed Games (publishing/feed.py:351), each one over HTTP through the public API (feed.py:158-194), and each takes a new version number whether it changed or not (feed.py:295).

8. Redis and Google: write-if-newer, then pull

Redis is one more destination, with its own queue, retries and proof (E8). Postgres is written first and Redis right after, because the two cannot share a transaction (D11). The Redis job runs one Lua script:

-- Agreed design, not in the code yet. KEYS[1] = feed:{doc_key}; ARGV = version, etag, gzip body
local cur = redis.call('HGET', KEYS[1], 'v')
if cur and tonumber(cur) >= tonumber(ARGV[1]) then return 0 end
redis.call('HSET', KEYS[1], 'v', ARGV[1], 'etag', ARGV[2], 'gz', ARGV[3])
return 1

It writes version 9120 only if Redis holds something older, so a retry or a second copy can never put an old file back. Measured: 0.30 ms at the median for a 3.3 KB file (laptop, Redis 7 in Docker, 7 Oct).

Google's next pull. Google calls /getfeeds with its secret key (D14) and the ETag of the file it holds. The API server checks the key in memory, with no database read (D13). It asks Redis only for the version: 0.27 to 0.30 ms. The version is newer than its copy in memory, so it fetches the gzip body once, keeps it, and answers 200 with the new ETag. On Google's poll after that, the ETag matches: 304, no body (E11).

Section 1 sets the target for API and push clients at 2 s from the save, for 95 in 100 updates. With the 1 s file timer, the Redis copy is ready a little over 1 s after the save on the example clock (estimate). How often Google polls was not checked.

Today Partly built: every /getfeeds hit reads Postgres four times before any cache (packages/api/src/omnium_api/routers/getfeeds.py:159-175). The Redis cache is a 3 s timer, never refreshed by a rebuild, and off when no REDIS_URL is set. /getfeeds needs no key (packages/api/src/omnium_api/auth.py:42).

9. Client queues: each destination on its own

Each client destination is one row in delivery_assignment, and gets its own queue and task (D4, E3). Versions 9120 and 9121 go on each queue that wants them:

DestinationWhat happens (example)
NDTV over S3Sends both files. S3 answers with an ETag, saved as the receipt. Age about 2 s
News18 over FTPSends both, one at a time. The remote file size after upload is the receipt
DailyHunt over SFTPThe host hangs. The send stops at the 60 s limit; only this queue waits and backs off

Every transport has a 10 s connect and a 60 s send limit (E4). A queue keeps only the newest waiting version of each file: if a second goal makes version 9135 while 9120 still waits for DailyHunt, 9120 is dropped and 9135 is sent.

After each send, the proof goes into delivery_state in one statement (E5). The columns below are the agreed ones; the statement is a sketch written for this page:

-- Sketch from E5 (columns from Section 9; the design doc gives no SQL for it).
UPDATE delivery_state
   SET last_seq = 9120, last_delivered_at = now(), status = 'sent', failures = 0,
       sent_hash = $hash, sent_bytes = $bytes, receipt = $s3_etag,
       age_ms = (extract(epoch FROM now() - $changed_at) * 1000)::int
 WHERE assignment_id = $ndtv_s3 AND doc_key = $doc_key;

age_ms is the time from the goal's save to the confirmed send: the number the 10 s target is measured on. An operator who is asked "did NDTV get the goal?" reads this row on the client page.

Today Partly built: the delivery worker asks the database who is owed a file every 5 s (delivery_poll_seconds, settings.py:238). A file is owed when its seq is newer (publishing/delivery/dispatcher.py:142), and every rebuild makes a new seq, so all 7 files go to every destination. Sends run in one shared round of 8 that waits for the slowest (publishing/delivery/runner.py:214-224), and SFTP has no time limit (transports.py:436-437). One hanging host stalls every client.

10. Stats: the heavier numbers, right after

The goal's own transaction already wrote the lines it changed: the scorer's and the assist's lines in stat_value (S10, M3). The heavier stats wait for nobody's answer. The stats queue reads the same outbox row with the late-commit-safe reader (D7) and queues the match's stat runs, such as football's football.top_scorers.node and football.table.node. The stats engine (packages/stats) runs them on its Procrastinate queue in Postgres, in lanes, so a long backfill never sits in front of a live match.

If a stat job fails after its tries, the alert service raises stat-job-failed, naming the stat and the match.

How long this takes was not measured; the design sets no separate target for it.

Today Partly built: the scheduler's tail reads domain_event rows with id above its bookmark (packages/stats/src/omnium_stats/triggers.py:65-71), listening to command.accepted, fixture.completed and fixture.reopened (triggers.py:30). It waits 2 s after a pass that found nothing (IDLE_SECONDS, packages/scheduler/src/omnium_scheduler/__main__.py:39). A late commit can be skipped; a 15-minute sweep repairs it. A failed job is tried 3 times (packages/stats/src/omnium_stats/worker.py:87), then writes a stat.failed event that only the Runs screen shows.

11. Fans: Redis first, then the read-only copy

A fan's app asks the public API for the score. The public API reads Redis first (D8). On a miss it reads the read-only copy through Sessions.read(), which uses the copy only while it is at most 5 s behind (DB1, DB3). The goal reached the copy about a millisecond after the commit: 0.90 ms at the median under 5,700 writes a second (laptop, 7 Oct). On AWS, copy lag is not measured yet.

Copy lagWhat the fan's read does
Milliseconds, or under 5 sReads the copy, with its data age
Over 5 sReads with a freshness limit move to the main database, at most 4 connections per server
Over 10 s for 60 sSame, plus an alert

Staff screens, scorers, the delivery-worker and alert checks never read the copy (DB2). A live match's file that misses Redis is read from the main database, never the copy.

Today Agreed, to build: there is one database and no copy. If REPLICA_DATABASE_URL were set, the public API would send all its reads to it (packages/api/src/omnium_api/app.py:50-60), and the feed builder reads through the public API, so client files could go out without the goal. That is why DB5 switches a copy on only after the delivery-worker builds from the main database directly.

12. The alert service: watching every stop

Nobody watches a dashboard to know this goal arrived. The alert service watches for them. Code adds one alert_signal row in the same transaction as a failure (alerts.raise_), and a checker looks at our own numbers every 15 s. Signals fold into one open alert per problem, sent once to Slack, email or the console (A1 to A3).

If this stop failedWhat noticesAlert
1. The device cannot sendThe desk shows "Offline, N changes waiting"; nothing is lostNone to staff while it is offline; scorers never see system alerts
3. Commands slow downp95 of a person's command over 30 ms, last hour and last 5 minutescommand-slow
3. A lock wait over 1 s, or a transaction open over 10 sThe checker reads lock waitslock-wait
3. Our code fails the command 5 timesRaised with raise_alone, because the command rolled backcommand-refused
3. The match's sport code is missing after a deployThe engine refuses to score and raisessport-code-missing
3. Saved state drifts from the logThe background check replays every live match every 5 minutesreplay-mismatch
7. The delivery-worker's build loop stops or hangsEach background process beats every 10 s; 3 missed beats, or no progress for 60 s with work waitingworker-silent, worker-stuck
8. Redis writes failThe Redis destination fails 3 timesredis-write-failed
9. A client is behindA destination over 60 s behind, or 3 failures in a rowdelivery-behind
10. A stat job failsAfter its triesstat-job-failed
11. The copy lagsLag over 10 s for 60 sthe copy-lag rule
12. The alert service itself stopsIts beat to CloudWatch stops (after Section 14); until then a red console banner after 60 sAWS alarm

In the DailyHunt case at stop 9: the checker finds the destination 60 s behind, opens one delivery-behind alert, waits 30 s to group, and sends one message: "DailyHunt SFTP: 2 files behind". When the host recovers, the alert resolves itself.

Today Agreed, to build: alerts reach nobody. Prometheus has 5 rules and no receiver (deploy/observability/prometheus/prometheus.yml.tmpl:10-11), and the dead-letter Slack alerter only writes a log line (packages/core/src/omnium_core/publishing/delivery/deadletter.py:161).

13. The trace id: one search, the whole path

The bridge accepted traceparent with trace id 4bf92f35… at stop 2. That id is saved on the command row and the domain_event row, carried by the delivery-worker when it builds (each feed_document keeps the last 20 trace ids it was built from) and when it sends (delivery_attempt), and stamped on every JSON log line (A9, N7).

The design names the fields every line carries: trace id, command id, subject, workflow, client, outcome and time taken. It does not fix their key names, so the keys below are illustrative:

{"svc": "commands",  "trace_id": "4bf92f35…", "command_id": "7f3a2c1e-…", "subject": "fixture:01a1d4e2-…", "workflow": "football", "outcome": "accepted", "ms": 2.4}
{"svc": "delivery",  "step": "build", "trace_id": "4bf92f35…", "subject": "fixture:01a1d4e2-…", "outcome": "built", "ms": 210}
{"svc": "delivery",  "trace_id": "4bf92f35…", "client": "NDTV S3", "outcome": "sent", "ms": 640}
{"svc": "delivery",  "trace_id": "4bf92f35…", "client": "DailyHunt SFTP", "outcome": "timeout", "ms": 60000}

One search for 4bf92f35 shows the goal's whole path. An alert carries its trace id too, so "show logs" on the DailyHunt alert opens exactly these lines. Writing one JSON log line cost 6.9 µs on a laptop (7 Oct).

Today Partly built: domain_event has a trace_id column, but the id is made fresh at the outbox (packages/core/src/omnium_core/jobs_outbox.py:68) because the scoring pipeline passes none (pipeline.py:971-979). The stats engine carries that id into its runs (triggers.py:87); the bridge, the file builders and delivery do not. Only the public API writes JSON logs throughout, with its own request id; the workers write mostly plain text.

The timeline

From the tap to the client's file, on the example clock. Every time says where it comes from.

StepPartWhat happensTime (source)
1Desk (scorer app)Outbox row saved; 1–0 on screenunder 0.1 s from the tap (target, V6)
2Bridge, networkRequest travels to omniumabout 100 ms one way to US East (estimate, Section 1); 10 to 30 ms to an Indian region (half of the Section 4 estimate)
3Command service, PostgresLock, checks, step, one transaction1.48 ms median, 3.98 ms p99 (measured, laptop, 5 Oct); 2 to 3 ms on prod (estimate, M9); under 5 ms per event (target, S9)
4Bridge, desk"Accepted, 57"; tickabout 19:42:10.200 here; save to tick under 1 s (target, V6)
5ListenerNOTIFY wakes each server0.20 ms median, 0.53 ms p99 (measured, laptop, 6 Oct)
6Shared watchesScorer B and the console get the new valueunder 0.5 s and under 1 s from the save (targets, B8)
7Delivery-worker, buildFinds 2 files, builds, hashes, saves 9120 and 91211 s timer, at most 3 s (agreed, D2); build under 3 s (target, D6); hash of 4.3 MB 1.74 ms (measured)
8Redis, GoogleWrite-if-newer; next pull gets the file0.30 ms write (measured, laptop, 7 Oct); API and push under 2 s from the save, 95 in 100 (target, Section 1)
9Client queuesNDTV, News18 send; DailyHunt times outunder 10 s from the save (target, D6); send under 5 s (target); 60 s send limit (agreed, E4)
10Stats queueTop scorers, group tablenot measured; no separate target
11Read-only copy, public APIFans see 1–0copy lag 0.90 ms median at 5,700 writes a second (measured, laptop, 7 Oct); fallback at 5 s (agreed, DB3)
12Alert serviceDailyHunt 60 s behind: one alertchecker every 15 s, 30 s group wait (agreed, A2, A6)
13LogsOne trace id on every line6.9 µs per JSON line (measured, laptop, 7 Oct)
A hand-drawn time ruler titled From the save to the client file (targets). A long arrow runs left to right, labelled time after the save. Eight evenly spaced ticks, left to right, with green dots: 0, saved; 0.2 ms, wake-up; 0.5 s, scorer B; 1 s, console; 1 to 3 s, file built; 2 s, Google; 10 s, client file. The last tick, in red: 60 s, alert. Below the first tick, a white box Tap has an arrow up to it, labelled network + 3 ms. A note at the bottom says targets, not measurements.
The agreed targets, counted from the save, drawn evenly spaced rather than to scale. Only the 0.2 ms wake-up is a measurement (laptop); the rest are targets to prove in the Section 12 tests.

A second story: an import after a staff fix

The same match. Full time is 20:35, and the desk's events say India 2, Japan 1. At 20:41 an operator checks the official match sheet: Japan scored a late goal the desk missed. The real result is 2–2. Two minutes later a Scout run arrives from the results site, which still says 2–1.

A hand-drawn sketch titled A fix, then an import 2 minutes later. Top row: a white box Console: away 2 sends an arrow labelled edit_side to a yellow Engine box, which sends an arrow labelled write + pin to a blue cylinder Match row, 2-2, result pinned. Bottom row: a white box Scout run: 2-1 points to a yellow box Fast path, row changed since?, which points with an arrow labelled yes to a yellow box Batch: compare. A dashed arrow labelled reads now comes down from the Match row cylinder into Batch: compare. Batch: compare points to a diamond pinned?. From the diamond, an arrow labelled yes goes to a green box Keep 2-2, and an arrow labelled saved too goes down to a blue box Claim: Scout 2-1. From the claim an arrow labelled differs too long goes to a red box Alert. From Keep 2-2 a dashed arrow labelled if official goes to a red box Review item.
The hand edit wins, the import's value is kept as a claim, and a person is told if the two keep disagreeing.

Example · The console fix, then the Scout run (agreed design)

20:41:00. The fix. The operator sets Japan's score to 2 on the match page. The console sends results.edit_side through the same outbox and the same road as the goal: a new key, the match lock, and the rows version check, because this is an edit (saw carries the side row's changed_seq, C3, L4). The writer saves the score. Under Section 2, the value is a claim from the source "hand edit", which sits at the top of the result list by default (T4). It shows on the "open overrides" screen. The side row's changed_seq moves to the new log number. An events match takes a scorecard edit like this as a hand edit (S14).

The outbox note goes out exactly as at stops 5 to 9: the console and desks update, the delivery-worker rebuilds this match's files with 2–2, and every client gets the new version.

20:43:00. The Scout run. Scout rings the doorbell: run passed. The import fetches the rows with no database transaction open (I1, K4), records the run as received (I2), and maps each row to a key and a hash.

  1. The fast path does not skip this match. The site's row may even be the same as last time. But the side row's changed_seq moved since this integration last wrote it, so the key goes on to the compare (K2).
  2. One batch, on the import pool. It locks the match row with FOR NO KEY UPDATE, like any command (C6). If a person were saving this match at that moment, the batch would wait a few milliseconds.
  3. Compare with the database now, minus pins (I3). Scout says Japan 1; the row holds 2, set by hand. The hand edit is higher on the list, so nothing is written for the score. Other fields on the row, such as the status, are still compared and written if they really differ and are not held.
  4. Scout's value is not thrown away. It is saved as a claim from Scout (T1). The run screen shows it under "Kept as set by hand".
  5. The compare job (T6). Scout says 2–1 and we publish 2–2. If that lasts longer than the set time for this sport, the operator gets one alert: "Scout says 2–1, we publish 2–2". It never changes what clients get.
  6. If the result was already official. If the operator had marked the result official before the run, and this integration does not have may_change_official on, a different value from Scout opens a review item (I9, T9). Clients see no change until a person accepts it.
  7. Later, the site shows 2–2. The compare screen shows both agree. The operator releases the override. Scout's newest claim is already saved, so the list decides again at once, with no wait for the next run.

Today Partly built:

  • The console's results.edit_side writes the score and pins result on that side (packages/core/src/omnium_core/workflows/writers/fixtures.py:312-318).
  • The import first compares each row with its own last read, not with the database (packages/core/src/omnium_core/integrations/service.py:835-851). If the site's row has not changed since the last run, it is not sent at all.
  • If it is sent (say the site's status changed), _update_side finds result pinned, keeps 2–2 and adds 1 to kept_by_hand (packages/core/src/omnium_core/ingest/fixtures.py:1490-1513).
  • Scout's 2–1 is kept only in integration_row.raw. There is no compare, no alert and no review item, and no writer reads results_official, so an import can change an official score unless a person pinned it.
  • The import holds an advisory lock on the whole event until its job commits (packages/core/src/omnium_core/workflows/engine.py:116), so every console edit on that Games waits for it.

When things go wrong

Every case ends with the goal saved once, and every reader getting it, or a person told. This is the agreed design.

What goes wrongWhereWhat happens
The answer is lost in a tunnel4The desk resends with the same key and gets duplicate. One goal
The server crashes before the commit3Postgres rolls back everything, including the command row. No answer was sent. The resend runs normally
Scorer B taps Goal at the same moment3B waits for the match lock, then gets conflict with event 57 inside. B's screen shows A's goal and asks B to check
A deploy happens mid-match3The match keeps its rules and code (code_sha). Commands cut off before commit are rolled back and resent
The database switches to its standby3, 4Try again until it is back; the desk keeps the goal in its outbox
A NOTIFY is lost5The 5 s safety read finds the outbox row. Screens are at most 5 s late, never wrong
Scorer B's push is dropped6The next value skips a version; B's screen shows "Catching up" and re-reads one record
The delivery-worker crashes mid-build7Nothing half-built is saved. On restart it reads the outbox from its last place and builds again
A rebuild gives the same bytes7Same hash: no new version, nothing sent
Redis is down8API servers serve their copy in memory, with its age; misses read Postgres, at most 4 at once per server; an alert
DailyHunt's host hangs9Only DailyHunt's queue waits. delivery-behind at 60 s. When it returns, it gets the newest file, not every version
A stat job fails10Retried, then stat-job-failed names the stat and the match
The copy falls 7 s behind11Fan reads with a freshness limit move to the main database; the delivery-worker never used the copy
The alert service dies12Signals keep landing as rows. On restart it folds from where it stopped. Its silence is alarmed from outside

Decisions

The decisions this page leans on, by stop. All are Agreed. Each page under "Where to read more" has the full list.

#DecisionIn plain words
V1One shared outbox on the device: key made once, saved before sending, one at a time per match, retried with the same keyA dropped network never loses a goal and never sends it twice
C2Every command has a key the sender makes once; the first answer is kept 30 daysA resend is answered "duplicate"
L1The match lock is the match's own row, FOR NO KEY UPDATEEvery match has its own lock
L3The key is checked before the lock and again under it; the command row commits with the change"Was it saved?" always has a true answer
L17Three levels of version check: none, rows, playTwo scorers never save on top of an unseen goal
C4, L11"Accepted" goes out only after COMMIT returnsAccepted always means saved
S1, M1One saved state row per match, in the same transaction as every eventThe 63rd-minute goal costs what the first did
M9A command makes 2 database tripsAbout 2 to 3 ms on prod (estimate)
B1, G1Every command transaction sends NOTIFY with its outbox id; one listener per container, plus a 5 s safety readServers are told within a millisecond
L9Outbox readers read by transaction id, below the oldest running oneA late commit is never skipped
B4, B8Push the new value; targets 0.5 s for a second scorer, 1 s for other staffLive has a number we test
D1, D2, D3One build path, in the delivery-worker, from the main database; only touched files, 1 s wait, 3 s at most; a new version only when the hash changesUnchanged files are never sent again
D4, D5, D6One queue per destination; proof of delivery; 10 s from save to file, alert at 60 sOne slow client delays only itself
D11, E11Postgres first, then Redis write-if-newer; gzip and 304Pull clients read memory, never an old file
D7Stats read the outbox safely; a failed stat job alertsNo stat silently missing
DB2, DB3Fans read the copy; over 5 s behind, reads fall back to the main databaseFans are never more than 5 s behind
A1, A3One alert service; one open alert per problemOne message per problem, not fifty
A9One W3C trace id from the bridge to deliveryOne search shows the whole path
T1, T4, T6Every source's value is a claim; a hand edit is a source at the top; losers are comparedThe fix wins, the import's value is kept, a person is told
I3, K2, I9Imports compare with the database now; skip only when nothing touched the record; an official result becomes a review itemA staff fix is never undone silently

Built today, or still to build

Checked in the code on 7 Oct 2026, branch feat/result-feeds, commit 76ef1e2.

StopTodayAgreed designStatus
1. Tap and keyNew key per call (app.ts:332); no device outbox; no football deskOutbox in IndexedDB (V1)Partly built
2. Requestsend_command, CommandIn (bridge.py:183, :213)Adds saw, sentAt, traceparentPartly built
3. Enginepipeline.command (pipeline.py:845): advisory lock, whole-log reads, refoldRow lock, command table, version check, saved state, 2 tripsPartly built
4. AnswerMay go out before commit (deps.py:28); two shapes (bridge.py:77-91)After commit; five outcomesPartly built
5. Wake-up1 s polling per screen (bridge.py:766, :924)NOTIFY and one listener per containerAgreed, to build
6. Live screens"Area changed", 250 ms wait, re-reads (host.ts:196)Pushed values with versionsPartly built
7. BuildAll 7 files, new seq every time, over HTTP (feed.py:295, :351, :158-194)feed_dependency, content_hash, main databasePartly built
8. Redis and pull4 Postgres reads per hit, 3 s cache, no key (getfeeds.py:159-175, auth.py:42)Memory and Redis, write-if-newer, secret keyPartly built
9. Client queuesOne shared round of 8 (runner.py:214-224); SFTP no limitOne queue per destination, proof, limitsPartly built
10. StatsTail by id (triggers.py:65-71); no alert on failureLate-commit-safe reader; alertPartly built
11. Fans and the copyOne database; replica URL would route all API reads (api app.py:50-60)Sessions.read(), heartbeat, 5 s fallbackAgreed, to build
12. Alerts5 Prometheus rules, no receiver; Slack seam logs only (deadletter.py:161)One alert serviceAgreed, to build
13. Trace idMade at the outbox (jobs_outbox.py:68); stats carry it (triggers.py:87)From the bridge to deliveryPartly built
Import after a fixPin kept; compare with own last read (service.py:835-851); no alert or reviewClaims, compare with the database, T6 alert, I9 reviewPartly built

Numbers

NumberWhatWhere it came from
under 0.1 sClick to screenTarget (V6)
about 200 msRound trip, India to US EastEstimate (Section 1, design doc point 7)
20 to 60 msNetwork, India to an AWS region in IndiaEstimate (Section 4)
0.09 msKey lookup, key not seen, 10 million rows, medianMeasured, laptop, 5 Oct
1.48 ms / 3.98 msWhole command transaction, median / p99Measured, laptop, 5 Oct
2 to 3 msOne command with 2 trips on prodEstimate (M9)
under 5 msServer time per event, any point in a matchTarget (S9)
under 30 msA person's command, p95Target (Section 4)
under 1 sSave to tickTarget (V6)
0.20 ms / 0.53 msCommit to wake-up, median / p99Measured, laptop, 6 Oct, 500 tries
under 0.5 s / under 1 sSave to second scorer / to other staffTargets (B8)
1 s, 3 sFile timer, and its maximumAgreed (D2)
1.74 mssha256 of a 4.3 MB fileMeasured, design doc
0.30 msRedis write-if-newer, 3.3 KB file, medianMeasured, laptop, 7 Oct
0.27 to 0.30 msRedis version check, medianMeasured, laptop, 7 Oct
2 sAPI and push clients, 95 in 100 updatesTarget (Section 1)
10 sSave to client fileTarget (D6)
60 sBehind before delivery-behind firesAgreed (D6, E6)
0.90 msCopy lag, median, 5,700 writes a secondMeasured, laptop, 7 Oct
5 sCopy lag at which reads move to the main databaseAgreed (DB3)
6.9 µsOne JSON log lineMeasured, laptop, 7 Oct
1 s, 250 ms, 5 s, 2 sToday: screen poll, host wait, delivery poll, stats tail idleFrom the code (bridge.py:766, host.ts:196, settings.py:238, scheduler/__main__.py:39)

Where to read more