Running it · What is next
Build order and what is parked
What we build first, in what order, what each step needs, how we know it is done, and the three design sections still waiting their turn.
In one minute
Design sections 1 to 11 were agreed between 3 and 7 Oct 2026. Sections 12 (testing), 13 (security) and 14 (deploy and operations) were parked on 7 Oct: "keep this aside, first we focus on building the system". Section 12 is written but none of it is agreed. Sections 13 and 14 were never written.
This page gathers the "order of work" lists from every section page into one build order of 11 phases. The order follows the dependencies the pages state. For example, the delivery-worker must build from the main database (Section 9, D1) before a read-only copy is switched on (Section 10, DB5). The order is proposed. It has not been agreed yet.
Three things matter most:
- One risk is live today. Prod's Jenkins job offers
DB_TASK=seed, which wipes the domain tables with no approval step (Jenkinsfile_prod:74-78). The fix (Section 12, T16) is parked with Section 12. - The parked sections block the end of the road. The read-only copy needs AWS work (Section 14). The shadow run needs proof (Section 12). Role-aware screens need sign-in (Section 13).
- Several gaps were found while writing these docs. They are listed at the end, each with the phase it must be settled before.
Where we are
The design sections
| Section | Topic | State | Decisions | Page |
|---|---|---|---|---|
| 1 | Goals, users and numbers | Agreed, to build 3 Oct | 15 (D1 to D15) | Goals |
| 2 | Truth and ownership | Agreed, to build 3 Oct | 11 (T1 to T11), plus 9 edge rules | Truth and ownership |
| 3 | Imports from Scout | Agreed, to build 6 Oct | 24 (I, U, K) | Imports |
| 4 | Commands and the workflow engine | Agreed, to build 6 Oct | 30 (C1 to C12, L1 to L18) | Commands |
| 5 | Live scoring engine | Agreed, to build 6 Oct | 23 (S, M) | Scoring engine |
| 6 | Sport plugins and releases | Agreed, to build 6 Oct | 22 (P, U, R) | Sport plugins |
| 7 | Console and scorer apps | Agreed, to build 6 Oct | 17 (V1 to V10, F1 to F7) | Console |
| 8 | Bridge and live updates | Agreed, to build 6 Oct | 14 (B1 to B8, G1 to G6) | Live updates |
| 9 | Stats, feeds and delivery | Agreed, to build 7 Oct | 27 (D1 to D15, E1 to E12) | Feeds |
| 10 | Database and the read-only copy | Agreed, to build 7 Oct | 25 (DB, DE), plus 1 chat rule; 3 extras still proposed | Database |
| 11 | Monitoring, logs and alerts | Agreed, to build 7 Oct | 22 (A1 to A13, N1 to N9) | Monitoring |
| 12 | Testing, rehearsal and developer experience | Parked 7 Oct | 26 proposed (T1 to T16, H1 to H10), none agreed | This page |
| 13 | Security and access | Parked 7 Oct, not written | none | This page |
| 14 | Deploy, environments, operations | Parked 7 Oct, not written | none | This page |
What exists in code today
The code that ran the Asian Games 2026 is the starting point. Checked on branch feat/result-feeds, commit 76ef1e2.
| Part | Today | Where |
|---|---|---|
| Workflow engine, apply mode | Built and on prod: key, advisory lock, rules, log row, outbox | packages/core/src/omnium_core/workflows/engine.py:89 |
| Event-by-event scoring | A second engine that refolds the whole match each command; not on prod | packages/core/src/omnium_core/scoring/pipeline.py:845 |
| Outbox | domain_event, written in the same transaction as the change | packages/core/src/omnium_core/jobs_outbox.py:44 |
| Bridge | Protocol v1, published as @fanos/omnium-bridge 1.0.0 | packages-ts/omnium-bridge/src/protocol.ts:8 |
| Console | Built (separate repo omnium-console); one 594 KB JS file | Console |
| Imports | Scout jobs, no-code mapping, pins, guards, held rows | packages/core/src/omnium_core/integrations/service.py |
| Client files | Built through the public API over HTTP, sent by the delivery worker | packages/core/src/omnium_core/publishing/feed.py:158-194 |
| Sport plugin | fanos-omnium-flows 0.6.0: 10 workflows, 245 components, 20 goldens | Sport plugins |
| Monitoring | Prometheus with 5 rules and no receiver; Grafana with 3 dashboards | Monitoring |
| Migrations | 59 on this branch; migration 060 is on main (PR #155) and nothing reads it yet | packages/core/alembic/versions/, Data model |
| The gate | make check: 1,605 tests, 1,604 passed, 1 skipped, in 126 s | Laptop, 7 Oct (Section 12 design doc) |
None of the new tables the sections agreed exists yet. A search of packages/ on 7 Oct found no match for command_failure, integration_key_state, changed_seq, plugin_release, release_default, feed_dependency, alert_signal, process_beat, replica_heartbeat or pg_notify.
The build order
Proposed This order is proposed. It has not been agreed with the user. It is built from the dependencies the section pages state. A phase starts when the phases it needs are done. Effort and dates were not estimated, and no page gives them.

The order at a glance
| # | Phase | Sections | Needs phase | Waits on a parked section |
|---|---|---|---|---|
| 0 | Safety fixes and quick wins | all | none | Seed guard is Section 12, T16 |
| 1 | The one road | 4 | 0 (helpful, not required) | no |
| 2 | Alerts, beats and logs | 11 | 1 (the late-commit-safe reader) | AWS pieces: Section 14 |
| 3 | Scorecards and rule releases | 5, 6 | 1 | Who may approve a prod release: Section 13 |
| 4 | Screens and live updates | 7, 8 | 1 | Sign-in: Section 13. Proxy timeout: Section 14 |
| 5 | Who wins: claims and result status | 2 | 1 | no |
| 6 | Imports | 3 | 1, 2, 5 | no |
| 7 | Files out: build in the delivery-worker, queues, Redis pull | 9 | 1, 2, 4 | Pull keys: Section 13. Load test: Sections 12, 14 |
| 8 | Events mode | 5, 6 | 1, 3 | Prod approval: Section 13 |
| 9 | Read-only copy and retention | 10 | 2, 7 | The copy on AWS: Section 14 |
| 10 | Shadow run, then switch clients | 1 | 1 to 4, and 8 (see below) | Proof: Section 12 |
Phases 2, 3 and 4 can run side by side once phase 1 is done. Phase 0 and the first six steps of phase 9 need nothing and can start any time.
What month 1 means. Target 9 (Section 1) names the first release as "the base first (scoring worker, workflows, console), ready in one month and proven with football and curling". In this order that is phases 1 to 4, plus the minimum alert set D8 asks for in month 1 (phase 2). Whether phases 5 to 9 must be done before the shadow run is not stated on any page. That is open question 1 below.
Phase 0: Safety fixes and quick wins Proposed
- Builds: the fixes in Quick wins below. Each is small, already agreed or a plain bug, and stands alone.
- Decisions: Imports migration step 1; Scoring migration step 1; Section 9 order steps 1 to 4 (E1, E3, E4); Section 10 order step 1 (DB13, DE8, DE10); Section 8 step 1 (B1, harmless while nobody listens); Section 11 steps 2 and 3 (A7, A9, N5); Section 6, P11.
- Depends on: nothing.
- Done when: each fix has its own test. For example: an SFTP test host that accepts and then hangs times out at 60 s, and other destinations deliver on time (Section 9, D4). A rebuild with no real change sends nothing (D3).
make new-sport, thenmake check, passes (P11). - Pages: Feeds, Database, Imports, Scoring engine, Sport plugins.
Phase 1: The one road (Section 4) Proposed
Everything else is built on this phase. Imports send their batches through road.run. The console's outbox needs required keys, five answers and conflicts. Live updates, stats and alerts all use the late-commit-safe outbox reader.
- Builds: the new
commandandcommand_failuretables;changed_seqonfixtureandfixture_competitor; thexidcolumn ondomain_eventand the reader incore/.../outbox/reader.py;road.run; the row lockFOR NO KEY UPDATEin a fixed order; a required key; the version check at three levels; five answers; refusals saved with a code and a sentence; two connection pools with pool-level time limits; lock metrics and alerts. - Decisions: C1 to C12, L1 to L9, L11 to L14, L16 to L18, and M9 (Section 5). L10 (sport code in killable worker processes) lands in phase 8, where Section 5's migration order puts it (step 4). L15 (commands in their own ECS service) needs deploy work: see Section 14.
- Depends on: nothing. Phase 0 makes it calmer, not possible.
- Done when:
- The two-writer test (L12): two connections send to one match at once, many times, and every round ends in one order or the other. Section 12's local harness measured 0 wrong rounds out of 1,000 with the lock, and 1,000 out of 1,000 with the lock removed (laptop, 7 Oct).
- Kill the server between the writes and the commit, and after the commit, then retry with the same key: never half saved, never saved twice.
- A used key with a different body is refused as
key_reused(C2). - "Accepted" goes out only after
COMMITreturns (C4, L11). Today the commit runs after the answer (packages/admin/src/omnium_admin/deps.py:28). - A note whose transaction commits late is still read by the outbox reader (L9).
- Pages: Commands, Workflow code.
Phase 2: Alerts, beats and logs (Section 11) Proposed
Alerts come early because imports (I10, I13), files out (Section 9, D6, E6), stats (D7), the scoring background check (S5) and copy lag (DB3) all raise alerts through this one service.
- Builds: the alert tables (
alert_signal,alert,alert_event,alert_send,process_beatand the rule and route tables);alerts.raise_; the first rules with runbooks; the alert service; shared JSON logs; beats from every loop; the Health page; the Slack bot-token client; the trace id from the bridge to delivery; the nightly cleanup. This follows the monitoring page's order of work, steps 1 to 8. - Decisions: A1 to A13, N1 to N9.
- Depends on: phase 1's late-commit-safe reader (N1). The
ops.*commands can use today's engine. - Done when: 8 senders raise one key at once and get one alert with count 8, and no sender waits on a lock. A signal that commits late is still folded. A worker that beats but makes no progress with work waiting raises
worker-stuck. Two alert services run, and only the lock holder sends. Then one week on staging with sending off, to measure the noise, before Slack goes on (monitoring step 4). - Waits on Section 14: the CloudWatch alarm on the alert service's beat (A8), email through SES (N6), the Slack token in Secrets Manager, and lasting storage for Prometheus (N8). Until then, every console screen shows a red banner when the alert service has not beaten for 60 s.
- Pages: Monitoring.
Phase 3: Scorecards and rule releases (Sections 5 and 6) Proposed
Most matches are scorecards. This phase makes the scorecard path complete and gives every match a fixed release of the rules.
- Builds: Section 5 migration steps 2 and 3: the
engine/package withroad.pyandscorecard.py; Games and Results move onto it;workflows/engine.pyis removed;results.set_lineandwriters/lines.py; the mode set per competition. Section 6 migration steps 1 to 4, 6 and 8:code_shaover the whole package; the newplugin_releaseandrelease_defaulttables;fixture.release_id; the pin at creation;results.change_typeandresults.set_release; the releases screen. Step 9, the new-sport scaffold, is in phase 0. - Decisions: S11 to S14 (Section 5); P1 to P3, P5, P7, P8, P10, P12, U1 to U3, R1 to R3, R6 (Section 6).
- Depends on: phase 1.
- Done when: two operators fixing different batters never block each other, and two fixing the same batter get "conflict" (S13). Changing a shared helper changes
code_shafor every sport (P3). Taking a release back moves the defaults back, and matches on it keep it (P7). A change of type on a scorecard unit with places lists the places first, then removes only what was confirmed (R6). The goldens give the same answers through the new path. - Settle first: the
release_defaultNULL problem (open question 5). - Pages: Scoring engine, Sport plugins, Workflow code.
Phase 4: Screens and live updates (Sections 7 and 8) Proposed
- Builds: the console's build order, steps 1 to 9: the outbox in IndexedDB and
useChange;useFieldon every text field;results.edit_fixture; the five field states;React.lazyroutes and virtual lists; parallel re-reads; read-only screens for viewers; the scorecard grid and Ctrl+K; the shared UI package and the desks moved into the repo. The live-updates build order, steps 1 to 7:pg_notifyafter each outbox insert; one listener per container; the client wrapper that reopens after any close; shared watches; record keys on the schedule reads; versions and gap detection; versions in heartbeats. - Decisions: V1 to V10, F1 to F7 (Section 7); B1 to B8, G1 to G6 (Section 8).
- Depends on: phase 1 (required keys, five answers, the version check, the late-commit-safe reader).
- Done when: offline in the browser, 6 changes, then back online: all 6 saved in order, none twice (V1). A reload with 3 changes waiting sends all 3 (F1). Two browsers on one match: the other shows a new ball within 0.5 s, with no read request in its network log (B4, B8). 200 idle screens for 5 minutes: database queries a minute stay flat (B2). Kill the listening connection: screens catch up within 5 s, nothing missed (B1).
- Waits on parked sections: viewers get read-only screens only when people sign in (V8; Section 13). The stream's 10 s cut-off goes only after the proxy timeout is raised (G5, F5; the live-updates page tracks that in Section 14).
- Pages: Console, Live updates.
Phase 5: Who wins, claims and result status (Section 2) Proposed
No page gives a build order for Section 2. Its place here is my proposal.
- Builds: every source value saved as a claim; field groups; a priority list per group, per sport, competition and match; the hand edit as a source on each list; the "open overrides" screen; result status (live, unofficial, official, protested) with a version; the official lock and review items; the link from each published value to its winning claim.
- Decisions: T1 to T11 and the edge rules E2, E3, E4, E7, E10, E11, E13, E14, E16 (all Section 2).
- Depends on: phase 1. The command transaction saves the claims and the winner (C10).
- Why before imports: an import writes claims once this exists, and I9 needs the "may change official" setting (Section 2, T9 and T10). Building imports first would mean changing how they write twice.
- Done when: the page lists no tests, so these checks come from the decisions. An older claim from a source never replaces a newer one from the same source (T7). An official result refuses a different claim from a source and opens a review item (T9). Every published value names its winning claim, its source and why (T11).
- Pages: Truth and ownership.
Phase 6: Imports (Section 3) Proposed
- Builds: the imports migration order, steps 2 to 10: the fetch outside any transaction, with retries; the run recorded first; the new
integration_key_statetable; the fast path; batches of 25 matches throughroad.run; missing rows become review items; the settle window; row-log trimming and the 30-day delete; the console's two-line health, run screen and "Waiting for you" list. The import queue moves into the command service, in a process of its own with its own heartbeat, and admin-api only receives the doorbell (8 Oct). Proposed with it: a 60 s lease instead of today's one-hour block on a run cut off by a restart. - Decisions: I1 to I13, U1 to U3, K1 to K8 (all Section 3).
- Depends on: phase 1 (
road.run,changed_seq, the import pool, L7), phase 2 (I10 and I13 raise through the alert service), phase 5 (I9 reads the official setting). - Done when: a crash at batch 7 of 12 loses nothing and repeats nothing (I5). A run of 4 rows writes (I12). A row missing from two full runs opens a review item and cancels nothing (I6). An HTTP call to Scout that hangs holds no lock (K4). The fast path answers 1,000 keys in about 2 ms (measured on a laptop, 6 Oct).
- Settle first: I9 against I12 (open question 3).
- Pages: Imports.
Phase 7: Files out: build in the delivery-worker, queues and Redis pull (Section 9) Proposed
- Builds: the Section 9 order, steps 5 to 9: proof of delivery and age in
delivery_stateand the client screens; alerts through the central service; the build in the delivery-worker on the Section 8 wake-up, calling resolvers directly on the main database; the newfeed_dependencytable and building only touched files; the stats reader fix; version and provisional or official marks per client. Then the Redis pull path (client files and, since 8 Oct, the public match documents), per-client send settings and secret keys, which the design doc has not yet placed in its order. - Decisions: D1, D2, D5 to D15, E2, E5 to E12 (all Section 9). D3, D4, E1, E3 and E4 start in phase 0.
- Depends on: phase 1 (the stats reader uses L9), phase 2 (D6 and E6 alerts), phase 4 (the build is woken by the Section 8 wake-up, B1 and G1).
- Done when: 50 changes in 2 seconds build each file at most once per 3 s (D2, E3). A host down for 10 minutes then gets only the newest version of each file, and the alert fired at 60 s (E3, E6). A change to one match on a 7,000-unit Games rebuilds only the files that contain it (E2). Replaying a day of Games data puts 95% of change-to-file ages under 10 s (D6). A pull load test holds 5,000 hits a second (D15).
- Waits on parked sections: how pull keys (D14) are issued, stored and revoked is Section 13. The D15 load test needs a load environment (Section 12, T10, and Section 14).
- Pages: Feeds.
Phase 8: Events mode (Sections 5 and 6) Proposed
- Builds: Section 5 migration steps 4 to 9: the step-function contract, sport worker processes and
code_shachecked on every event; football's step function;events.py,state.pyandcorrections.py; the background checkverify.py; cricket's step function; the other sports move over;scoring/pipeline.pyis removed. Section 6 steps 5 and 7: releases side by side in workers; the CI release workflow and the replay corpus. - Decisions: S1 to S10, M1 to M8 (Section 5); L10 (Section 4); P4, P6, P9, R4, R5, R7 (Section 6).
- Depends on: phases 1 and 3.
- Done when: a full-length synthetic match per sport: the last 100 events cost at most 2 times the first 100, and the state stays under 16 KB (M8). A correction at a random event equals a fresh replay of the corrected log. A sport with a deliberate endless loop costs one command "try again", and the server keeps answering others. A match started on release A stays on A after B is installed (P6). A scoreboard row damaged by hand is found within one interval and rebuilt through a command, with an alert (S5). A recorded cricket match runs through the new engine in a test (Section 1, D3).
- Waits on Section 13: releases reach prod only on a person's approval (P5, R7). Who may press it is Section 13's question.
- Pages: Scoring engine, Sport plugins.
Phase 9: Read-only copy and retention (Section 10) Proposed
- Builds: the database order of work. Steps 1 to 6 can start any time: quick wins;
Sessions.write()andSessions.read()with each route choosing; storing only changed import rows; monthly partitions forintegration_rowanddomain_event; the migration runner limits and its CI check; the health page anddb_size_daily. Step 7 is phase 7's build from the main database. Step 8: the newreplica_heartbeat, the lag watcher and the fallback, tested against a real copy in Docker. Step 9: the copy itself on AWS. - Decisions: DB1 to DB15, DE1 to DE10, and the 7 Oct chat rule (Section 10).
- Depends on: phase 2 (the alert service writes the heartbeat, DE2) and phase 7 (the delivery-worker builds from the main database, D1). DB5 says a real copy is switched on only after both of these hold: the delivery-worker builds from the main database directly, and every read route has passed against a real copy.
- Done when: the DB5 conditions hold. A copy over 5 s behind moves reads with a freshness limit to the main database. Dropping a month of a partitioned table replaces the delete (59 ms against 9.3 s, laptop, 7 Oct).
- Waits on Section 14: the copy on AWS (DB6), cost and setup. No AWS step without asking first.
- Settle first: the
/v1documents route (open question 2). - Pages: Database.
Phase 10: Shadow run, then switch clients (Section 1) Proposed
- What happens: football and curling run in shadow next to the legacy .NET system on the same real matches, and the outputs are compared after every match. Clients switch one by one after two clean weeks (D4). Client feeds moved from the legacy system keep exactly the same responses (D12). One named person is on duty per live day (D8). Three device setups are certified (D10).
- Decisions: D2 to D4, D8, D10 to D12, D15 (Section 1).
- Depends on: "the base is fully built and tested" (D15). D3 adds that a recorded cricket match must run through the foundation in a test before it counts as done (phase 8).
- Waits on Section 12: "tested" has no agreed meaning while Section 12 is parked. The shadow run cannot honestly start until it has one.
- Pages: Goals.
Quick wins that need no new design
Each item is already agreed, or is a plain bug. None needs a new table design, and none depends on another phase.
| Fix | From | Code today |
|---|---|---|
_retire_gone gets a savepoint; skipped unnamed units stop being reported as accepted | Imports, migration step 1 | integrations/service.py:684, :701-712 |
ledger.append checks the actors before it writes | Scoring, migration step 1; Section 4 bug list | scoring/ledger.py:235-256 |
| Time limits on SFTP and S3; one reused S3 client per destination | Feeds, step 1 (E4) | transports.py:169-171, :425-437 |
| One send task per destination instead of one shared round | Feeds, step 2 (E3) | runner.py:214-224 |
content_hash on feed_document; send only on a change; "last updated" from real changes | Feeds, step 3 (D3, E1) | feed.py:295 |
| Remove the listener send path; add the event filter to the dispatcher | Feeds, step 4 | listener.py:92, dispatcher.py:121 |
| One failed feed no longer stops every feed after it in a pass | Lessons (F28); no decision names it | publishing/feed.py:320-328 |
application_name and a 60 s idle-transaction limit on every pool | Database, step 1 (DB13, DE8) | core/db.py:18-37 |
A daily partition job; a daily call to the outbox prune that nothing calls today | Database, step 1 | jobs_outbox.py:75 |
Delete finished stats-queue jobs; drop the duplicate index on raw_ingested_result | Database, step 1 (DB8, DE10) | models/ingestion.py:129-148 |
pg_notify after every outbox insert, harmless while nobody listens | Live updates, step 1 (B1) | none today |
| Shared JSON logging in every service, then beats from every loop | Monitoring, steps 2 and 3 (A7, A9) | api/.../observability/logging.py:28 only |
| Parallel re-reads in the host, 4 at a time | Console, step 6 (F5) | host.ts:215-229 |
| The new-sport scaffold makes the six files, a test and a golden | Sport plugins (P11) | packages/flows/templates/sport/, last changed 28 Aug |
| Restore the test summary line; add Redis to CI; write the test-database step in the docs; seed refusal locally and on staging | Section 12 order, step 1 Parked | Makefile:62, pyproject.toml:137, .github/workflows/ci.yml |
File paths are short forms of the paths on each page. The last row is parked with Section 12.
Parked: testing (Section 12) Parked
Section 12 is written but parked. On 7 Oct the user said: "keep this aside, first we focus on building the system". All 16 product decisions (T1 to T16) and 10 engineering decisions (H1 to H10) were Proposed. None was agreed. The doc has no comment threads. Everything below is parked.
What the doc found today
The bottom of the test ladder is solid. make check runs lint, strict type checks, 7 import rules, the flows structure check and the tests. CI starts Postgres 17, builds the test database and stops if it is missing. Above that, most things are missing.
| Gap | Where (checked again for this page, 7 Oct) |
|---|---|
One click can wipe prod. Prod's Jenkins job offers DB_TASK=seed, which wipes the domain tables, with no approval step | Jenkinsfile_prod:74-78; the wipe is a TRUNCATE of every mapped table, packages/core/src/omnium_core/seed/writer.py:147-163 |
| Load was tested once, on 7 Sep, with no pass marks; not nightly, not in CI | loadtest/harness.py, loadtest/run_L*.json |
| Browser tests cannot block a merge | .github/workflows/ci.yml:159 (continue-on-error: true) |
| No Redis in CI, so the Redis rate-limit test skips | ci.yml has no Redis service |
The test summary line is hidden by a doubled -q | Makefile:62, pyproject.toml:137 |
| 26 test files commit and then delete their rows; a crash leaves rows behind | Section 12 doc, e.g. admin/tests/test_ssot_loop.py:142 (not re-checked) |
| Deploys have no smoke test | Section 12 doc, Jenkinsfile_stg:672-697 (not re-checked) |
| No rehearsal mode, no failures on purpose, no property tests, no data generators | Section 12 doc |
| Coverage is configured but not enforced | pyproject.toml:144 |
The 7 Sep load test (one laptop, one scoring process, from loadtest/PHASE0-REPORT.md, as quoted in the doc): 5 live matches gave 9.9 accepted a second with a 130 ms 99th percentile; 100 matches gave 33.1 a second and 4.4 s; 250 matches gave 38.5 a second and 7.5 s. The report's own finding: the refold on every ball was the ceiling. Sections 4 and 5 rebuild exactly that path, so these numbers are the baseline to beat.
What the doc proposed
| Area | Proposal | Ids |
|---|---|---|
| Test sizes | Every test is small, medium or large; a small test that opens a socket fails | T1, H1 |
| Honest CI | Redis in CI; any skip not on an allow-list fails the run; counts shown; browser tests block merges | T2, T3, H2 |
| Scoring proofs | Goldens, property tests, the Section 6 compatibility replay, and the two-writer test (200 rounds per pull request, 1,000 nightly) | T4, H5 |
| Speed proofs | A load test with a pass mark for every speed promise; busy day, spike, 8-hour soak, breakpoint | T5, H6 |
| Failures on purpose | Nightly: kill a sport worker, cut the database, hang a client host, fill Redis, stop the alert service; a fault proxy; sandbox destinations with five modes | T6, H7 |
| Test data | Generators per sport, a "Games in a box", recorded provider payloads; never a copy of prod | T7, H4 |
| Rehearsal | competition.is_rehearsal; a trigger on delivery_state refuses a real destination; a replayer at 1 to 10 times speed | T8, H8 |
| Ready checklist | A proof per item and a named sign-off, refused while an item is red | T9, H9 |
| Where tests run | Laptop, CI, staging for a smoke test only, a load environment for nightly and event tests | T10 |
| Flaky tests, slow creep | Quarantine with an owner, at most 5; alert on a 20% nightly slowdown | T11, T12 |
| Developer start | make dev starts everything and plays a demo match | T13 |
| Config and migrations | Each environment has every setting it needs (names only); migrations tested on prod-sized generated data | T14, T15 |
| Wipe guard | Seed and every wipe refuse on prod | T16, H10 |
What it measured (laptop, 7 Oct): the two-writer test with the row lock, 1,000 rounds in 9.31 s, 0 wrong; with the lock removed, 1,000 of 1,000 wrong. A toy property test found a real bug on its first run: an undo left "home: 0" where there was nothing before. After the fix, 1,000 cases of up to 600 events took 1.16 s.
What it would have proven. Its pass marks are the numbers the agreed sections promised: a person's command under 30 ms at the 95th percentile, the lock under 20 ms at the 99th, a second scorer under 0.5 s, files under 10 s, 5,000 pull hits a second. While Section 12 is parked, these stay promises. Several "done when" checks above depend on them.
Parked: security (Section 13) Parked
Section 13 was parked on 7 Oct before it was written. It has no decisions. These items were handed to it by earlier sections, or found while writing these pages.
| Item | Today | Handed over by |
|---|---|---|
| Sign-in on the admin API | Off by default: admin_auth_required: bool = False (packages/core/src/omnium_core/settings.py:56), and ADMIN_AUTH_REQUIRED=false in .env.deploy.example:83. With it off, every call is made as "admin" (admin/auth.py:31). On 24 Sep the prod admin API answered with no login, still true on 5 Oct (open-issues list S6); not re-checked on prod | As-built review; Lessons |
| A secret key per pull client | /getfeeds is exempt from the API key; the client name in the address is the only check (packages/api/src/omnium_api/auth.py:42). Section 9 agreed the key (D14). How keys are issued, stored and revoked is Section 13's | Feeds, D14 |
Scout's /v1 has no sign-in | Scout adds no auth middleware or dependency (scout/app/main.py:105-118, app/api/router.py). Omnium's client says Scout "has no authentication of any kind, by design" and is meant to be reached only from an internal network (packages/core/src/omnium_core/ingest/scout.py:65-67). Whether each Scout server answers from the internet: not checked here | As-built review |
| Roles | Who gets operator, admin or viewer; the console shows every control today | Console, V8 |
| Who may approve a release for prod | P5 and R7 need a person's approval | Sport plugins |
| How CI signs in to upload test results | Section 12, H9 | Section 12 |
| Personal data rules for the replay corpus | Section 6's nightly export of finished prod matches waits for these (Section 12, T7) | Section 12 |
Parked: deploy and operations (Section 14) Parked
Section 14 was parked on 7 Oct before it was written. These items were handed to it.
| Item | Handed over by | What the code and Terraform say |
|---|---|---|
| The AWS setup for the read-only copy; a Multi-AZ DB cluster recommended (one writer, two readable standbys) | Database, DB6 | In the Terraform, prod's database is an aws_db_instance with multi_az = true (omnium/production/us-east-1/main.tf:315): one standby that cannot be read. DB6's Multi-AZ DB cluster is a different AWS product. REPLICA_DATABASE_URL points at the writer itself (main.tf:116-117, :260) |
| Redis with a standby | Feeds, D8 | The Terraform defines a dedicated Valkey for prod: create_valky = true, cache.t4g.small, num_cache_nodes = 1 (terraform.tfvars:83-86). One node means no standby. The Section 9 doc recorded "prod has no Redis" |
| Lasting storage for Prometheus | Monitoring, N8 | Staging's Terraform keeps Prometheus on EFS (omnium/staging/us-east-1/prometheus-efs.tf, main.tf:236). Prod's does not |
| The alert service as its own service | Monitoring, N3 | Not in the code |
The container change decided on 7 Oct: add the command service (2 copies) and the alert service; switch the stats-worker on, with the timed jobs; run the import queue in the command service (8 Oct); turn off admin-api's import thread (integration_queue_in_admin); remove the scheduler | The big picture | Prod's Terraform runs the scheduler at 1 copy and sets the stats-worker to 0 copies (terraform.tfvars:211-233). There is no command service or alert service yet |
| The CloudWatch alarm on the alert service's beat | Monitoring, A8 | Not in the code |
| Email through SES | Monitoring, N6 | Not in the code |
| The Slack token in Secrets Manager | Monitoring | Today Slack is a webhook setting, off and never called (core/notify/slack.py:64) |
| A load-test environment | Section 12, T10 | Staging's Terraform has a load-test stack behind create_loadtest: a shadow public API with its own database and Valkey, and an Artillery task (omnium/staging/us-east-1/loadtest.tf). It covers the public API only, not the command road |
| The proxy timeout, so streams need no 10 s cut-off | Live updates, G5 | settings.py:267; the 15 s proxy cut is from a code comment, not checked in AWS |
Removing the prod seed choice from Jenkins | Section 12, T16 | Jenkinsfile_prod:74-78 |
| Staging that can hold a rehearsal | Lessons | Staging's 1 GB database ran out of memory twice on 21 Sep |
| The doorbell trap: wiring prod removed staging's Scout webhooks | Lessons | No decision yet |
Open questions found while writing these docs
Each was flagged on a page, or found in the code for this page. None is settled. The last column says which phase must wait for the answer.
| # | Question | Found on | Settle before |
|---|---|---|---|
| 1 | Which phases must be done before the shadow run? Target 9 names "scoring worker, workflows, console" as the base. Whether imports, files out, the copy and events mode are part of it is not stated. Which mode football and curling are scored in for the shadow run is not stated either | Goals | Phase 10 |
| 2 | /v1 documents also read the copy if one is set. The public API gives every request a session on the copy when REPLICA_DATABASE_URL is set (api/app.py:50-60, deps.py:10-21), and /v1/fixtures/... reads published_document through it (routers/documents.py:49-52). The 7 Oct chat rule says a live match never reads the copy. The database order of work does not name moving this route. Since 8 Oct these documents are served from Redis first, written by the delivery-worker; a miss should read the main database, like client files | Database | Phase 9 |
| 3 | I9 against I12. I9 says an official result an import may not change opens a review item. I12 says a run that would change such a result is held for a person, with nothing written. Which wins when one row in a large run does this? | Imports | Phase 6 |
| 4 | The event lock. Section 4's text says event-wide data "locks the event's stream row". Its code locks the competition row with plain FOR UPDATE, which also blocks inserts of rows that point at it | Commands | Phase 1 |
| 5 | release_default has a UNIQUE key on (sport, competition, from_at). A sport-wide row has competition_id NULL, and a plain Postgres UNIQUE treats NULLs as different, so two sport-wide rows with the same time would both save. It needs UNIQUE NULLS NOT DISTINCT or a partial index | Sport plugins | Phase 3 |
| 6 | "Accepted" before the commit. The admin API commits in get_write_session after yield (packages/admin/src/omnium_admin/deps.py:28), which FastAPI 0.139 runs after the answer is sent | Commands, Goals | Phase 1 fixes it (C4, L11) |
| 7 | Fold mode checks actors after it saves. ledger.append inserts the log row, then checks actors (scoring/ledger.py:235-256). A refused actor answers "not accepted", but the session still commits, so the row would stay. Read in the code, not run | Workflow code | Phase 0 |
| 8 | make new-sport templates are out of date. packages/flows/templates/sport/ still makes compute.py, triggers.py and procedures.py, and a new sport fails the structure gate | Sport plugins, Workflow code | Phase 0 (P11) |
| 9 | Which time budget does a long cricket match hit first? The design doc says a 50 ms limit at about ball 550; the code gives an action 200 ms and a validation 50 ms (scoring/program.py:86-92). Not checked | Scoring engine | Phase 8 |
| 10 | No break-glass path when CI is down and a fix is urgent: R7 removes publishing from a laptop | Sport plugins | Phase 8 |
| 11 | The Redis pull path, per-client settings and secret keys (Section 9, D11 to D15, E7 to E12) are not yet placed in the design doc's order of work | Feeds | Phase 7 |
| 12 | Who sets up the two new ECS services, the command service (L15) and the alert service (N3), and removes the scheduler? The list was fixed on 7 Oct 2026: six services, with imports in the command service and the file build in the delivery-worker. It is on The big picture. Setting them up on AWS is Section 14's job | Commands, Feeds, Monitoring | Phases 1, 2, 7 |
| 13 | A draw is stored two ways: the football plugin puts both sides on rank 1; imported matches leave rank empty. One rule for every source is not decided | Data model | Phase 5 |
| 14 | Migration 060 on main adds a second id map (xref.entity_map) beside external_id_map, and a second stat store (stat) beside stat_value. Which one wins is not decided | Data model | Phase 6 |
| 15 | Picking who won on the match page sends two results.edit_side commands. Section 7 did not decide whether this becomes one command | Console | Phase 4 |
| 16 | Three Section 10 extras are still proposed: return to the copy only under 2 s, one copy per API server, pause imports while the copy is over 5 s behind | Database | Phase 9 |
| 17 | Prod facts not checked: the Postgres version (the Terraform says 18.3, terraform.tfvars:97, while CI and laptops run 17; DB15 wants one), point-in-time recovery (the Terraform says 14 days, main.tf:324; DB14 wants 14), whether pg_stat_statements is on, and whether prod copies every commit to its standby before confirming | Database, Commands | Phase 9 |
| 18 | Curling has no plugin and no catalogue entry today. How it is added for D2 is not stated (as a scorecard sport by settings, P1 would need no code) | Goals | Phase 3 |
Read next
- The big picture: every part on one page, and what the rebuild changes.
- Commands and the workflow engine: phase 1, the road everything else is built on.
- Lessons from the Games: why each phase exists.
- Goals, users and numbers: the targets the shadow run must meet.