Running it · Lessons
Lessons from the Games
What running the Asian Games 2026 taught us about our scoring workflow, and where each lesson landed in the agreed design.
In one minute
omnium carried the Asian Games. Three paying clients (NDTV, News18 and DailyHunt) got the same 7 files to the end. On the last day the medal table matched the official site exactly, and India's 85 medals were the same in every file.
But almost every write came from a Scout import or a hand edit in the console, not from live scoring. So the Games tested imports, the command road, feeds and delivery. Those are where it hurt.
The five worst lessons, ranked by harm in the as-built review of 29 Sep:
- Imports fought each other and fought hand edits. One import wiped the results of 114 finished matches.
- Failures were silent. A Scout job was rejected for 6 days before anyone saw it.
- The write APIs had no login.
- A unit's shape and its pairs could not change after they were first written.
- Prod was run by hand, and staging did not match prod.
Each lesson below links to the design page that fixes it, with the decision id. Most fixes are agreed design, not built yet. A few were fixed in the code during the Games; those are marked built.
The Games in numbers
Every number here comes from a written source. The source is in the last column.
| What | Number | Source |
|---|---|---|
| Time to build omnium before the Games | about 10 weeks | As-built review, 29 Sep |
| Paying clients served | 3: NDTV (S3), News18 (FTP and webhook), DailyHunt (SFTP) | Issue list, 5 Oct |
| Files per client | 7 (calendar, results, medal table and others) | Issue list, 5 Oct |
| Client send rules that got the final files | 40, all ok, 14:08:13 to 14:08:43 UTC on 5 Oct | Final fixes note, 5 Oct |
| India's medals at the close | 21 gold, 27 silver, 37 bronze (85) | Issue list, 5 Oct |
| Fixes in the week of 18 to 24 Sep | 83, of which 33 to feeds | As-built review |
| Commits that were fixes, 15 Aug to 24 Sep | 184 of 534 | As-built review |
| Import rows saved per day, never cleaned up | about 500,000 | As-built review |
| Data sent per rebuild, changed or not | about 24 MB to the 4 destinations | Section 9 design doc |
| Final clean-up on 5 Oct | 17 fix rows: 359 side changes by script, 63 by console | Final fixes note |
| India place or mark differences against the site | 60 before the final fixes, 1 after | Final fixes note, 19:44 IST |
| Imports switched off | 5 Oct, 12:59 IST | Issue list |
| Sending to clients switched off | 5 Oct, 19:51 IST | Final fixes note |

What went wrong, and what we changed
One row per lesson. "Where it is fixed" links to the page that explains the fix, with the decision id from that design section.
| Lesson | What happened | Where it is fixed |
|---|---|---|
| Imports undid each other's values | 24 Sep: the draw job wrote 114 finished matches with no score and no winner. Two were India's (F7) | Code fix 24 Sep Built today. Design: Truth and ownership T2, T3 Agreed, to build |
| Wrong data never healed | An import compared with its own last read, not our data. 27 Sep: 4 boxing bouts 6 hours late, India's Lovlina among them (F14) | Imports I3, K1, K2 Agreed, to build |
| Late corrections never reached us | Protests and corrections after an event: recurve scores about 55 short, heptathlon after a disqualification, sailing protests (F17). Fixed by script on 5 Oct | Imports I8, K7 Agreed, to build |
| A replay rewrote every row | 23 Sep: a replay meant to fix 2 units rewrote 29. Three got worse | Truth and ownership T7; Commands C11; Imports I3, I12 Agreed, to build |
| Failures were silent | The entries job was rejected 22 to 28 Sep. omnium showed "on time". A wrong partner went out on India's gold (F15) | Monitoring A1, A7; Imports I10, U1 Agreed, to build |
| Scout jobs died with no sign | 19 Sep: empty runs parked 3 jobs, one for 37 hours. 25 to 27 Sep: a memory leak slowed Scout, then it was killed | Scout fixes in Scout's repo, partly deployed. Imports I1, I10; Monitoring A7 Agreed, to build |
| Every rebuild resent every file | All 7 files went to every client on every rebuild. Identical files were sent again and again | Feeds and delivery D3, E1 Agreed, to build |
| One hanging host stalled every client | 24 Sep and 28 Sep: DailyHunt's SFTP host hung, and NDTV and News18 got nothing for minutes | Feeds and delivery D4, D12, E3, E4 Agreed, to build |
| Some edits did not rebuild the files | 27 Sep: a team fix at 22:25 IST was still missing from India's medal file at 22:35 (F27) | Feeds and delivery D2, E2 Agreed, to build |
| One failed feed stopped the others | 28 Sep 08:42 UTC: the first feed failed, and the error handler crashed, so no feed of the Games was rebuilt in that pass (F28) | No decision names it. Still in the code (feed.py:320-328). D6's 60 s freshness alert would catch it Partly built |
| The delivery worker could not build files on AWS | Until 1 Oct 06:49 UTC a missing setting made every build fail. Clients got files in bursts, with 4 to 13 minute gaps | Feeds and delivery D1 Agreed, to build |
| The delivery worker ran out of memory | A new S3 connection per send. On 23 Sep memory jumped from 40% to 70% at NDTV's first send, and AWS replaced the worker at least twice on 24 Sep | Feeds and delivery E4 Agreed, to build |
| A switched-off client was forgotten | 24 Sep: NDTV was switched off on purpose and then forgotten. It got no files from 16:34 to 18:08 IST | Feeds and delivery D9 Agreed, to build |
| A unit's shape was fixed at creation | Rounds first listed as two-slot matches later became ranked lists (F16, F18). Some were fixed by direct SQL on 28 Sep and 1 Oct. Mixed archery pairs printed no names (F31) | Sport plugins P10, R6, U3 Agreed, to build |
| Prod was changed outside the command log | 4 raw SQL changes between 24 Sep and 1 Oct. Fixes went in by curl and runbooks | Commands C1; Imports K6 Agreed, to build |
| Staging was not a safe place to rehearse | 21 Sep: the 1 GB staging database ran out of memory twice. Staging held no data after 25 Sep (T7) | Database DB15 Agreed, to build. Staging itself is parked with Sections 12 and 14 Parked |
| Wiring prod deleted staging's webhooks | 21 Sep: connecting prod's doorbells removed staging's webhook from 10 Scout jobs | No decision names it yet Parked |
| The write APIs had no login | 24 Sep: the prod admin API answered with no login (S6). Still true on 5 Oct | Parked with Section 13, security. See What is next Parked |
| "Accepted" could go out before the save | Found in the code on 30 Sep, after the Games' busiest days. Not seen to lose data on prod | Commands C4, L11 Agreed, to build |

The deep dives below take the eight lessons that cost the most. Each one says what happened, why, the real code behind it, and what the agreed design does instead.
Imports undid each other, and hand edits held only by luck
What happened. On 24 Sep we found 114 finished matches on prod with no score and no winner: wushu 50, table tennis 20, boxing 17, fencing 17, 3x3 basketball 4, cricket 3, squash 3. Each one was last written by the draw job. Two were India's: the women's cricket semifinal and APARNA's wushu bout. None was a medal match, so the medal table was safe.
Why. Several Scout jobs wrote the same matches: the schedule, the draw, the result pages and the medal winners jobs. No field had an owner. The draw job sends each match with names and status only, and the import read "no score" as "the score is now empty".
The only rule between writers was the pin. A person always wins, and pins the field. Between two imports, there was no rule at all:
# packages/core/src/omnium_core/workflows/records.py:173-181
def may_write(ctx: WriteContext, row: Any, name: str) -> bool:
"""The hand-edit rule for one field. A person always may, and pins it."""
if ctx.actor.by_person:
pin(row, name)
return True
if held(row, name):
ctx.report.counts["kept_by_hand"] += 1
return False
return True
- A person's edit pins the field. A pinned field is skipped by every import.
- The last line is the problem. Any import may write any field that is not pinned, so the last job to run wins.
Fixed in the code on 24 Sep Built today. A row that says nothing about how a match went no longer clears its result:
# packages/core/src/omnium_core/ingest/fixtures.py:919-937 (docstring trimmed)
def _says_how_it_went(unit: _Unit) -> bool:
"""Whether this row says anything about how a match went: a score, a winner or a place.
...
"""
if not unit.head_to_head:
return True
return bool(unit.scoreline) or any(
side.score or side.winner or side.rank is not None for side in unit.sides
)
This fixes one job. It does not give fields an owner.
The agreed design Agreed, to build. Every source's value is saved as a claim (T1). Values are split into field groups such as schedule, line-up and result (T2). Each group has its own priority list (T3). A draw job can then own the line-up and never touch the result. A hand edit is a source near the top of each list (T4). See Truth and ownership.
Wrong data never healed
What happened. On 27 Sep, 4 boxing quarter-finals showed start times 6 hours late, India's Lovlina Borgohain among them (F14). The site's schedule had the right time all along. Sailing races moved on 30 Sep stayed on the old day (F35). Results the site corrected after protests never reached us (F17).
Why. An import decides "changed or not" by comparing a row with its own last read of that row, not with what our database holds:
# packages/core/src/omnium_core/integrations/service.py:835-851 (docstring trimmed)
async def _changed(
session: AsyncSession,
integration: Integration,
target: ImportTarget,
ready: list[Row],
replay_of: uuid.UUID | None = None,
) -> list[Row]:
...
if not target.max_rows or replay_of is not None:
return ready
last = await _last_hashes(session, integration)
return [row for row in ready if not (row.key and last.get(row.key) == row.row_hash)]
_last_hashesreads the hash of each row as this integration last saved it.- A row with the same hash is dropped as "unchanged".
- So when another job (here, the draw job) wrote a wrong time, the schedule job saw its own row unchanged and never put the right time back.
The agreed design Agreed, to build. Change detection compares each row with what our database holds now, minus pinned fields (I3). A fast path may skip a row only when nothing has touched that record since this integration last wrote it (K1, K2). A finished unit is read again until the source marks it official, and then for a settle window of 72 hours by default (I8, K7). See Imports from Scout.
A replay rewrote every row
What happened. On 23 Sep we replayed a 29-row run of the result pages job on prod, to fix 2 pentathlon units. It rewrote all 29. Three got worse: a source typo in a name, a name with a repeated word, and a finished fencing bout set back to LIVE, which took it out of India's results.
Why. A replay sends rows omnium already holds from a stored run. It skips the "only changed rows" filter, the "newest run wins" check and the size check, and it makes a new key every time:
# packages/core/src/omnium_core/integrations/service.py:566
guarded = not dry_run and replay_of is None
# packages/core/src/omnium_core/integrations/service.py:854-857
def _key(fetched: Fetched, version: IntegrationVersion, replay_of: uuid.UUID | None) -> str:
"""What makes sending the same run twice write nothing. A replay always sends."""
key = f"{fetched.run_id or 'held'}:v{version.number}"
return f"{key}:replay:{uuid.uuid4().hex[:12]}" if replay_of is not None else key
guardedis false for a replay. It switches off two checks (service.py:587and598): "is this run older than the last one we used?" and "is this run far smaller than usual?". So an old run can overwrite newer data._changed(shown above) returns every row for a replay.- The key gets a random part, so the same replay twice writes twice.
- "Try" could not preview it. Try goes through the normal path, which skips unchanged rows, so it showed almost nothing.
What we did instead, from then on. Every later prod fix was the narrowest one: one command on one fixture, with every touched row listed first. The final clean-up on 5 Oct used that rule (see "What went right").
The agreed design Agreed, to build. An older claim from a source never replaces a newer one from the same source (T7), so a stored old run cannot set a finished bout back to LIVE. Imports compare with our data (I3). A run that would reopen more than 5 finished matches, or clear values on more than 10 records, is held for a person (I12). Every command can be tried first with nothing saved (C11).
Failures were silent
What happened. The Scout entries job was rejected on every run from 22 Sep 03:16 UTC to 28 Sep. 135 empty slots on the site failed Scout's name check, at 98.95% against a 99% bar. A rejected run never reaches omnium, so our entries stopped. omnium showed the integration as "on time" the whole time. We found it on 27 Sep, because India's 10m air pistol mixed team gold named the wrong partner (F15).
That was not the only silent failure:
| Date | What failed | How we found out |
|---|---|---|
| 19 Sep | Empty runs parked 3 Scout jobs for good. The medal winners job sat for 37 hours | A person looked |
| 21 Sep | A database restart on the Scout box stopped Scout's workers. Its health check still said ok | A person looked |
| 25 to 27 Sep | Each failed AI repair left a process of about 110 MB behind. Scout slowed from 25 Sep 18:04 UTC and was killed for memory on 27 Sep 06:32 UTC | A person looked, after the kill |
| 1 Oct | The delivery worker failed every build. Clients got late files | Reading logs |
Why. Nothing in omnium sends an alert to a person. The Slack sender for failed deliveries only writes a log line:
# packages/core/src/omnium_core/publishing/delivery/deadletter.py:160-161
async def alert(self, entry: DeadLetterEntry) -> None:
log.warning("slack alert (seam, not sent): %s", self.format_message(entry))
The agreed design Agreed, to build. One central alert service in our repo (I13, A1). Every part adds an alert_signal row in the same transaction as the failure. Every background process beats every 10 seconds, so a silent or stuck worker is found in 30 to 60 seconds (A7). Imports alert on a failed run and on 2 Scout runs in a row that failed Scout's checks (I10). The integration card shows Scout's health and omnium's health on two lines, green only when both are fine (U1). Section 1 also agreed a minimum alert set and one named person on duty per live day in month 1 (D8). See Monitoring, logs and alerts.
Every rebuild resent every file, and missed some changes
What happened. Every rebuild sent all 7 files to every client, changed or not. In the first 3 hours of News18's delivery (22 to 23 Sep), the same team_medals file went to them 23 times, identical each time. In the other direction, some real changes did not start a rebuild at all. On 27 Sep a hand fix to India's pair was written at about 22:25 IST, and India's medal file still had the old name at 22:35 (F27).
Why, part 1. Each feed takes a new number on every rebuild, whatever its content:
# packages/core/src/omnium_core/publishing/feed.py:293-313 (trimmed)
for definition in definitions:
try:
seq = await feed_store.next_seq(session)
body = await self._executor.execute(...)
...
await feed_store.upsert(
session,
...
body=payload,
etag=_etag(payload),
seq=seq,
)
- The sender sends a file when the number is higher than the last one it sent.
- The file's hash (
etag) is saved, but nothing compares it before sending.
Why, part 2. The "has anything changed" check looks at only three tables:
-- packages/core/src/omnium_core/publishing/feed.py:403-419 (trimmed)
SELECT (SELECT max(f.updated_at) FROM fixture f ...) AS units,
(SELECT max(fc.updated_at) FROM fixture_competitor fc ...) AS sides,
(SELECT max(m.updated_at) FROM medal_standing m ...) AS medals,
(SELECT min(d.updated_at) FROM feed_document d ...) AS feeds
A team member change or a person's new name is in none of these tables, so the files did not rebuild.
The agreed design Agreed, to build. A new version only when the file's content hash changes (D3, E1). Each feed declares the records it is built from, in a new feed_dependency table, so a change rebuilds exactly the files that contain it (D2, E2). See Stats, feeds and delivery.
One hanging host stalled every client
What happened. On 24 Sep at 16:21 IST we switched on one DailyHunt SFTP rule. Their server did not accept our address yet, so the connection hung. NDTV and News18 got nothing for that whole round, at least 5 minutes. On 28 Sep a test stalled every client again, from 08:31 to 08:42 UTC.
Why. Two things together. A send round waits for every send to finish:
# packages/core/src/omnium_core/publishing/delivery/runner.py:202-224 (trimmed)
async def flush(self) -> list[DeliveryResult]:
...
gate = asyncio.Semaphore(max(1, self._settings.delivery_concurrency))
async def one(job: _PendingJob) -> DeliveryResult:
async with gate:
return await self.deliver_now(job.spec, job.payload, trigger=job.trigger)
return list(await asyncio.gather(*(one(job) for job in jobs)))
And the SFTP sender is built with no time limit at all:
# packages/core/src/omnium_core/publishing/delivery/transports.py:432-437
def build_ftp_transport(settings: Settings) -> FtpTransport:
return FtpTransport(timeout=settings.delivery_transport_timeout)
def build_sftp_transport(settings: Settings) -> SftpTransport:
return SftpTransport()
asyncio.gatherwaits for the slowest send. Each send also retries up to 5 times inside the round.- FTP gets a timeout. SFTP does not, so a host that hangs holds the round for minutes.
- Switching a rule off does not stop a send already in flight.
DailyHunt's allow-list was fixed on 30 Sep. All 7 of their feeds went live on 1 Oct. The risk (S9) stayed in the code to the end.
The agreed design Agreed, to build. Each client destination has its own queue and task, so a slow destination never blocks another (D4, E3). Every transport has a 10 s connect limit and a 60 s send limit (E4). Each destination has its own settings for limits, retries and alert levels (D12, E12). A returning host gets only the newest version of each file, not every old one (E3).
The delivery worker: a missing setting and a memory leak
What happened. Two problems hid each other.
- It could never build files on AWS. Until 1 Oct, its feed builder called the public API at a default address that does not exist inside its container. Every build failed, and only the admin service built files, right after its own writes. Import changes waited for the next admin write. On 1 Oct, between 03:00 and 06:51 UTC, clients saw 21 gaps of 4 minutes or more. After the fix at 06:49 UTC (one setting added, memory raised from 0.5 to 1 GB), there was 1 such gap in the first 22 minutes.
- It used too much memory. Prod memory jumped from 40% to 70% at 16:58 IST on 23 Sep, the minute NDTV's first S3 send went out. On 24 Sep AWS replaced the worker at least twice.
Why, part 1. The address falls back to localhost:
# packages/core/src/omnium_core/settings.py:298-302
@property
def api_url(self) -> str:
"""The public API, as the admin service should call it. No trailing slash."""
found = self.omnium_api_url or self.public_api_base_url or "http://localhost:8000"
return found.rstrip("/")
Why, part 2. Every S3 send builds a new session and client:
# packages/core/src/omnium_core/publishing/delivery/transports.py:169-176
session = self._session_factory()
try:
async with session.client("s3", **client_kwargs) as s3:
await s3.put_object(
Bucket=bucket,
Key=object_key,
Body=payload.body,
ContentType=payload.content_type,
)
Measured on a laptop on 25 Sep: about 130 MB of memory that Python keeps per new client, against about 20 MB with one shared client.
The agreed design Agreed, to build. The delivery-worker alone builds every file, from the main database, by calling the resolvers directly, with no HTTP call to the public API (D1). One reused S3 client per destination (E4). A worker that stops beating raises an alert (A7).
Staging was not a safe place to rehearse
What happened. We wanted to try every risky prod fix on staging first. Staging could not carry that.
| Date | What happened |
|---|---|
| 21 Sep | The staging database, a 1 GB db.t4g.micro shared with another product, ran out of memory and restarted twice (05:13 and 12:40 UTC). Every "staging is down" report traced back to this box |
| 21 Sep | Connecting prod's doorbells deleted staging's webhooks from all 10 shared Scout jobs |
| after 25 Sep | Staging held no new data (T7). Its workflows were never installed (T1). Prod hand fixes were never copied to it |
Why the doorbell deleted staging's webhook. Prod's database was first copied from staging, so each prod integration row held staging's webhook id. Connecting a doorbell deletes the id the row holds:
# packages/admin/src/omnium_admin/routers/integrations.py:2186-2192
created = await scout.create_webhook(
integration.scout_base_url, integration.scout_job_id, url
)
if integration.webhook_id:
# The old one may be gone already; either way Scout stops using it.
with contextlib.suppress(scout.ScoutError):
await scout.delete_webhook(integration.scout_base_url, integration.webhook_id)
Nothing checks that the old id belongs to this environment.
Where it landed. Partly agreed, mostly parked:
- One Postgres major version in CI, staging and prod (DB15) Agreed, to build. Staging ran 18.3 when seen on 21 Sep; CI and laptops run 17.
- A rehearsal of a full Games day on staging, and a bigger staging database, wait for Section 12 (testing) and Section 14 (deploy). Both are parked Parked. See Build order and what is parked.
- The doorbell trap has no decision yet Parked. The working rule: after connecting a doorbell on either side, check the Scout job's webhooks and confirm both are there.
What went right
These held under real load and real mistakes. The agreed design keeps all of them.
| What worked | The proof at the Games |
|---|---|
| Delivery keeps its state in the database and polls | Clients caught up by themselves after outages and restarts. A restart lost nothing |
| Two safety switches on delivery | Our own test box was not on prod's allowed list, so prod refused every send to it from 22 Sep to the end, with "destination is not allowed here" |
| Pins: a hand edit wins until it is handed back | Operators trusted it (as-built review). Before the 5 Oct script, a read-only check confirmed none of its 359 planned sides was locked by hand |
| A fixed key on every command | The 5 Oct script sent each side with the key fix1005- plus the side id. On staging, the same key sent twice answered "duplicate" and applied once. On prod: 371 accepted, 0 duplicates, and a second run found nothing to send |
| One road: every change is a command | Apart from 4 raw SQL changes, every hand change on prod went through console commands, each one logged and pinned |
| Scout's run grading, and omnium's guards | A run that failed Scout's checks was never used. A bad import was rolled back whole, with a plain-words reason |
| Snapshots before and after every prod change | 63 dated snapshot folders. Each fix was checked by comparing files: only the planned rows changed, and India stayed at 85 |
| Saved GraphQL queries for client files | 7 file shapes served to 3 clients, with a transform per client where a client needed its own form |
Working habits that held. These came out of the incidents above and were the rules by the end of the Games (issue list, 5 Oct):
- Never replay a whole import run on prod. It rewrites every unit in it.
- Before a prod change, list every row it touches. Take a snapshot before and after.
- Use staging first when it holds the same data.
- Raw SQL must set
updated_at = now(), or the files do not rebuild. - Check medals in the client file, not only in the console.
- A client switched off on purpose is named in every status update until it is back on.
The design turns several of these habits into rules in the code: T7 and I12 for the first, C11 for the second, and D9 for the sixth.
Read next
- Truth and ownership: field groups and priority lists, the main answer to "imports undid each other".
- Imports from Scout: comparing with our data, batches, held rows and import alerts.
- Stats, feeds and delivery: one build path in the delivery-worker, send only on change, one queue per destination.
- Monitoring, logs and alerts: the central alert service and heartbeats.
- Build order and what is parked: what we build first, and where staging and security wait.