Running it · Lessons

Lessons from the Games

What running the Asian Games 2026 taught us about our scoring workflow, and where each lesson landed in the agreed design.

Period
19 Sep to 5 Oct 2026 (Games closed 5 Oct)
Based on
As-built review (29 Sep), prod issue list (5 Oct), incident notes
Lands in
Design sections 2, 3, 4, 6, 9, 10 and 11
Read time
about 20 minutes

In one minute

omnium carried the Asian Games. Three paying clients (NDTV, News18 and DailyHunt) got the same 7 files to the end. On the last day the medal table matched the official site exactly, and India's 85 medals were the same in every file.

But almost every write came from a Scout import or a hand edit in the console, not from live scoring. So the Games tested imports, the command road, feeds and delivery. Those are where it hurt.

The five worst lessons, ranked by harm in the as-built review of 29 Sep:

  1. Imports fought each other and fought hand edits. One import wiped the results of 114 finished matches.
  2. Failures were silent. A Scout job was rejected for 6 days before anyone saw it.
  3. The write APIs had no login.
  4. A unit's shape and its pairs could not change after they were first written.
  5. Prod was run by hand, and staging did not match prod.

Each lesson below links to the design page that fixes it, with the decision id. Most fixes are agreed design, not built yet. A few were fixed in the code during the Games; those are marked built.

114
finished matches wiped by one import (found 24 Sep)
6 days
entries job rejected before anyone saw it (22 to 28 Sep)
85
India's medals, the same in all 7 files and on the site (5 Oct)
371
final fix commands sent by script, 0 applied twice (5 Oct)

The Games in numbers

Every number here comes from a written source. The source is in the last column.

WhatNumberSource
Time to build omnium before the Gamesabout 10 weeksAs-built review, 29 Sep
Paying clients served3: NDTV (S3), News18 (FTP and webhook), DailyHunt (SFTP)Issue list, 5 Oct
Files per client7 (calendar, results, medal table and others)Issue list, 5 Oct
Client send rules that got the final files40, all ok, 14:08:13 to 14:08:43 UTC on 5 OctFinal fixes note, 5 Oct
India's medals at the close21 gold, 27 silver, 37 bronze (85)Issue list, 5 Oct
Fixes in the week of 18 to 24 Sep83, of which 33 to feedsAs-built review
Commits that were fixes, 15 Aug to 24 Sep184 of 534As-built review
Import rows saved per day, never cleaned upabout 500,000As-built review
Data sent per rebuild, changed or notabout 24 MB to the 4 destinationsSection 9 design doc
Final clean-up on 5 Oct17 fix rows: 359 side changes by script, 63 by consoleFinal fixes note
India place or mark differences against the site60 before the final fixes, 1 afterFinal fixes note, 19:44 IST
Imports switched off5 Oct, 12:59 ISTIssue list
Sending to clients switched off5 Oct, 19:51 ISTFinal fixes note
A hand-drawn timeline from 19 Sep to 5 Oct. Red boxes mark problems: 19 Sep, Scout jobs parked by empty runs. 21 Sep, staging database out of memory. 22 Sep, entries job rejected for 6 days. 23 Sep, replay rewrites 29 units. 24 Sep, 114 finished results wiped. 25 to 27 Sep, Scout slow, then killed. Green boxes mark good moments: 1 Oct, delivery worker builds files. 5 Oct, 17 fixes, Games closed.
The main incidents of the Games, by the date each started or was found.

What went wrong, and what we changed

One row per lesson. "Where it is fixed" links to the page that explains the fix, with the decision id from that design section.

LessonWhat happenedWhere it is fixed
Imports undid each other's values24 Sep: the draw job wrote 114 finished matches with no score and no winner. Two were India's (F7)Code fix 24 Sep Built today. Design: Truth and ownership T2, T3 Agreed, to build
Wrong data never healedAn import compared with its own last read, not our data. 27 Sep: 4 boxing bouts 6 hours late, India's Lovlina among them (F14)Imports I3, K1, K2 Agreed, to build
Late corrections never reached usProtests and corrections after an event: recurve scores about 55 short, heptathlon after a disqualification, sailing protests (F17). Fixed by script on 5 OctImports I8, K7 Agreed, to build
A replay rewrote every row23 Sep: a replay meant to fix 2 units rewrote 29. Three got worseTruth and ownership T7; Commands C11; Imports I3, I12 Agreed, to build
Failures were silentThe entries job was rejected 22 to 28 Sep. omnium showed "on time". A wrong partner went out on India's gold (F15)Monitoring A1, A7; Imports I10, U1 Agreed, to build
Scout jobs died with no sign19 Sep: empty runs parked 3 jobs, one for 37 hours. 25 to 27 Sep: a memory leak slowed Scout, then it was killedScout fixes in Scout's repo, partly deployed. Imports I1, I10; Monitoring A7 Agreed, to build
Every rebuild resent every fileAll 7 files went to every client on every rebuild. Identical files were sent again and againFeeds and delivery D3, E1 Agreed, to build
One hanging host stalled every client24 Sep and 28 Sep: DailyHunt's SFTP host hung, and NDTV and News18 got nothing for minutesFeeds and delivery D4, D12, E3, E4 Agreed, to build
Some edits did not rebuild the files27 Sep: a team fix at 22:25 IST was still missing from India's medal file at 22:35 (F27)Feeds and delivery D2, E2 Agreed, to build
One failed feed stopped the others28 Sep 08:42 UTC: the first feed failed, and the error handler crashed, so no feed of the Games was rebuilt in that pass (F28)No decision names it. Still in the code (feed.py:320-328). D6's 60 s freshness alert would catch it Partly built
The delivery worker could not build files on AWSUntil 1 Oct 06:49 UTC a missing setting made every build fail. Clients got files in bursts, with 4 to 13 minute gapsFeeds and delivery D1 Agreed, to build
The delivery worker ran out of memoryA new S3 connection per send. On 23 Sep memory jumped from 40% to 70% at NDTV's first send, and AWS replaced the worker at least twice on 24 SepFeeds and delivery E4 Agreed, to build
A switched-off client was forgotten24 Sep: NDTV was switched off on purpose and then forgotten. It got no files from 16:34 to 18:08 ISTFeeds and delivery D9 Agreed, to build
A unit's shape was fixed at creationRounds first listed as two-slot matches later became ranked lists (F16, F18). Some were fixed by direct SQL on 28 Sep and 1 Oct. Mixed archery pairs printed no names (F31)Sport plugins P10, R6, U3 Agreed, to build
Prod was changed outside the command log4 raw SQL changes between 24 Sep and 1 Oct. Fixes went in by curl and runbooksCommands C1; Imports K6 Agreed, to build
Staging was not a safe place to rehearse21 Sep: the 1 GB staging database ran out of memory twice. Staging held no data after 25 Sep (T7)Database DB15 Agreed, to build. Staging itself is parked with Sections 12 and 14 Parked
Wiring prod deleted staging's webhooks21 Sep: connecting prod's doorbells removed staging's webhook from 10 Scout jobsNo decision names it yet Parked
The write APIs had no login24 Sep: the prod admin API answered with no login (S6). Still true on 5 OctParked with Section 13, security. See What is next Parked
"Accepted" could go out before the saveFound in the code on 30 Sep, after the Games' busiest days. Not seen to lose data on prodCommands C4, L11 Agreed, to build
A hand-drawn map. Six red boxes on the left list problems: imports undo each other, replay rewrites everything, failures are silent, every file resent, one host stalls all, unit shape is fixed. Five yellow boxes on the right list design sections: truth and ownership, imports, monitoring and alerts, feeds and delivery, sport plugins. Arrows: imports undo each other goes to truth and ownership and to imports. Replay rewrites everything goes to truth and ownership and to imports. Failures are silent goes to monitoring and alerts. Every file resent and one host stalls all both go to feeds and delivery. Unit shape is fixed goes to sport plugins.
The six biggest lessons and the design pages that answer them.

The deep dives below take the eight lessons that cost the most. Each one says what happened, why, the real code behind it, and what the agreed design does instead.

Imports undid each other, and hand edits held only by luck

What happened. On 24 Sep we found 114 finished matches on prod with no score and no winner: wushu 50, table tennis 20, boxing 17, fencing 17, 3x3 basketball 4, cricket 3, squash 3. Each one was last written by the draw job. Two were India's: the women's cricket semifinal and APARNA's wushu bout. None was a medal match, so the medal table was safe.

Why. Several Scout jobs wrote the same matches: the schedule, the draw, the result pages and the medal winners jobs. No field had an owner. The draw job sends each match with names and status only, and the import read "no score" as "the score is now empty".

The only rule between writers was the pin. A person always wins, and pins the field. Between two imports, there was no rule at all:

# packages/core/src/omnium_core/workflows/records.py:173-181
def may_write(ctx: WriteContext, row: Any, name: str) -> bool:
    """The hand-edit rule for one field. A person always may, and pins it."""
    if ctx.actor.by_person:
        pin(row, name)
        return True
    if held(row, name):
        ctx.report.counts["kept_by_hand"] += 1
        return False
    return True
  • A person's edit pins the field. A pinned field is skipped by every import.
  • The last line is the problem. Any import may write any field that is not pinned, so the last job to run wins.

Fixed in the code on 24 Sep Built today. A row that says nothing about how a match went no longer clears its result:

# packages/core/src/omnium_core/ingest/fixtures.py:919-937 (docstring trimmed)
def _says_how_it_went(unit: _Unit) -> bool:
    """Whether this row says anything about how a match went: a score, a winner or a place.
    ...
    """
    if not unit.head_to_head:
        return True
    return bool(unit.scoreline) or any(
        side.score or side.winner or side.rank is not None for side in unit.sides
    )

This fixes one job. It does not give fields an owner.

The agreed design Agreed, to build. Every source's value is saved as a claim (T1). Values are split into field groups such as schedule, line-up and result (T2). Each group has its own priority list (T3). A draw job can then own the line-up and never touch the result. A hand edit is a source near the top of each list (T4). See Truth and ownership.

Wrong data never healed

What happened. On 27 Sep, 4 boxing quarter-finals showed start times 6 hours late, India's Lovlina Borgohain among them (F14). The site's schedule had the right time all along. Sailing races moved on 30 Sep stayed on the old day (F35). Results the site corrected after protests never reached us (F17).

Why. An import decides "changed or not" by comparing a row with its own last read of that row, not with what our database holds:

# packages/core/src/omnium_core/integrations/service.py:835-851 (docstring trimmed)
async def _changed(
    session: AsyncSession,
    integration: Integration,
    target: ImportTarget,
    ready: list[Row],
    replay_of: uuid.UUID | None = None,
) -> list[Row]:
    ...
    if not target.max_rows or replay_of is not None:
        return ready
    last = await _last_hashes(session, integration)
    return [row for row in ready if not (row.key and last.get(row.key) == row.row_hash)]
  • _last_hashes reads the hash of each row as this integration last saved it.
  • A row with the same hash is dropped as "unchanged".
  • So when another job (here, the draw job) wrote a wrong time, the schedule job saw its own row unchanged and never put the right time back.

The agreed design Agreed, to build. Change detection compares each row with what our database holds now, minus pinned fields (I3). A fast path may skip a row only when nothing has touched that record since this integration last wrote it (K1, K2). A finished unit is read again until the source marks it official, and then for a settle window of 72 hours by default (I8, K7). See Imports from Scout.

A replay rewrote every row

What happened. On 23 Sep we replayed a 29-row run of the result pages job on prod, to fix 2 pentathlon units. It rewrote all 29. Three got worse: a source typo in a name, a name with a repeated word, and a finished fencing bout set back to LIVE, which took it out of India's results.

Why. A replay sends rows omnium already holds from a stored run. It skips the "only changed rows" filter, the "newest run wins" check and the size check, and it makes a new key every time:

# packages/core/src/omnium_core/integrations/service.py:566
    guarded = not dry_run and replay_of is None

# packages/core/src/omnium_core/integrations/service.py:854-857
def _key(fetched: Fetched, version: IntegrationVersion, replay_of: uuid.UUID | None) -> str:
    """What makes sending the same run twice write nothing. A replay always sends."""
    key = f"{fetched.run_id or 'held'}:v{version.number}"
    return f"{key}:replay:{uuid.uuid4().hex[:12]}" if replay_of is not None else key
  • guarded is false for a replay. It switches off two checks (service.py:587 and 598): "is this run older than the last one we used?" and "is this run far smaller than usual?". So an old run can overwrite newer data.
  • _changed (shown above) returns every row for a replay.
  • The key gets a random part, so the same replay twice writes twice.
  • "Try" could not preview it. Try goes through the normal path, which skips unchanged rows, so it showed almost nothing.

What we did instead, from then on. Every later prod fix was the narrowest one: one command on one fixture, with every touched row listed first. The final clean-up on 5 Oct used that rule (see "What went right").

The agreed design Agreed, to build. An older claim from a source never replaces a newer one from the same source (T7), so a stored old run cannot set a finished bout back to LIVE. Imports compare with our data (I3). A run that would reopen more than 5 finished matches, or clear values on more than 10 records, is held for a person (I12). Every command can be tried first with nothing saved (C11).

Failures were silent

What happened. The Scout entries job was rejected on every run from 22 Sep 03:16 UTC to 28 Sep. 135 empty slots on the site failed Scout's name check, at 98.95% against a 99% bar. A rejected run never reaches omnium, so our entries stopped. omnium showed the integration as "on time" the whole time. We found it on 27 Sep, because India's 10m air pistol mixed team gold named the wrong partner (F15).

That was not the only silent failure:

DateWhat failedHow we found out
19 SepEmpty runs parked 3 Scout jobs for good. The medal winners job sat for 37 hoursA person looked
21 SepA database restart on the Scout box stopped Scout's workers. Its health check still said okA person looked
25 to 27 SepEach failed AI repair left a process of about 110 MB behind. Scout slowed from 25 Sep 18:04 UTC and was killed for memory on 27 Sep 06:32 UTCA person looked, after the kill
1 OctThe delivery worker failed every build. Clients got late filesReading logs

Why. Nothing in omnium sends an alert to a person. The Slack sender for failed deliveries only writes a log line:

# packages/core/src/omnium_core/publishing/delivery/deadletter.py:160-161
    async def alert(self, entry: DeadLetterEntry) -> None:
        log.warning("slack alert (seam, not sent): %s", self.format_message(entry))

The agreed design Agreed, to build. One central alert service in our repo (I13, A1). Every part adds an alert_signal row in the same transaction as the failure. Every background process beats every 10 seconds, so a silent or stuck worker is found in 30 to 60 seconds (A7). Imports alert on a failed run and on 2 Scout runs in a row that failed Scout's checks (I10). The integration card shows Scout's health and omnium's health on two lines, green only when both are fine (U1). Section 1 also agreed a minimum alert set and one named person on duty per live day in month 1 (D8). See Monitoring, logs and alerts.

Every rebuild resent every file, and missed some changes

What happened. Every rebuild sent all 7 files to every client, changed or not. In the first 3 hours of News18's delivery (22 to 23 Sep), the same team_medals file went to them 23 times, identical each time. In the other direction, some real changes did not start a rebuild at all. On 27 Sep a hand fix to India's pair was written at about 22:25 IST, and India's medal file still had the old name at 22:35 (F27).

Why, part 1. Each feed takes a new number on every rebuild, whatever its content:

# packages/core/src/omnium_core/publishing/feed.py:293-313 (trimmed)
            for definition in definitions:
                try:
                    seq = await feed_store.next_seq(session)
                    body = await self._executor.execute(...)
                    ...
                    await feed_store.upsert(
                        session,
                        ...
                        body=payload,
                        etag=_etag(payload),
                        seq=seq,
                    )
  • The sender sends a file when the number is higher than the last one it sent.
  • The file's hash (etag) is saved, but nothing compares it before sending.

Why, part 2. The "has anything changed" check looks at only three tables:

-- packages/core/src/omnium_core/publishing/feed.py:403-419 (trimmed)
SELECT (SELECT max(f.updated_at) FROM fixture f ...)              AS units,
       (SELECT max(fc.updated_at) FROM fixture_competitor fc ...) AS sides,
       (SELECT max(m.updated_at) FROM medal_standing m ...)       AS medals,
       (SELECT min(d.updated_at) FROM feed_document d ...)        AS feeds

A team member change or a person's new name is in none of these tables, so the files did not rebuild.

The agreed design Agreed, to build. A new version only when the file's content hash changes (D3, E1). Each feed declares the records it is built from, in a new feed_dependency table, so a change rebuilds exactly the files that contain it (D2, E2). See Stats, feeds and delivery.

One hanging host stalled every client

What happened. On 24 Sep at 16:21 IST we switched on one DailyHunt SFTP rule. Their server did not accept our address yet, so the connection hung. NDTV and News18 got nothing for that whole round, at least 5 minutes. On 28 Sep a test stalled every client again, from 08:31 to 08:42 UTC.

Why. Two things together. A send round waits for every send to finish:

# packages/core/src/omnium_core/publishing/delivery/runner.py:202-224 (trimmed)
    async def flush(self) -> list[DeliveryResult]:
        ...
        gate = asyncio.Semaphore(max(1, self._settings.delivery_concurrency))

        async def one(job: _PendingJob) -> DeliveryResult:
            async with gate:
                return await self.deliver_now(job.spec, job.payload, trigger=job.trigger)

        return list(await asyncio.gather(*(one(job) for job in jobs)))

And the SFTP sender is built with no time limit at all:

# packages/core/src/omnium_core/publishing/delivery/transports.py:432-437
def build_ftp_transport(settings: Settings) -> FtpTransport:
    return FtpTransport(timeout=settings.delivery_transport_timeout)


def build_sftp_transport(settings: Settings) -> SftpTransport:
    return SftpTransport()
  • asyncio.gather waits for the slowest send. Each send also retries up to 5 times inside the round.
  • FTP gets a timeout. SFTP does not, so a host that hangs holds the round for minutes.
  • Switching a rule off does not stop a send already in flight.

DailyHunt's allow-list was fixed on 30 Sep. All 7 of their feeds went live on 1 Oct. The risk (S9) stayed in the code to the end.

The agreed design Agreed, to build. Each client destination has its own queue and task, so a slow destination never blocks another (D4, E3). Every transport has a 10 s connect limit and a 60 s send limit (E4). Each destination has its own settings for limits, retries and alert levels (D12, E12). A returning host gets only the newest version of each file, not every old one (E3).

The delivery worker: a missing setting and a memory leak

What happened. Two problems hid each other.

  1. It could never build files on AWS. Until 1 Oct, its feed builder called the public API at a default address that does not exist inside its container. Every build failed, and only the admin service built files, right after its own writes. Import changes waited for the next admin write. On 1 Oct, between 03:00 and 06:51 UTC, clients saw 21 gaps of 4 minutes or more. After the fix at 06:49 UTC (one setting added, memory raised from 0.5 to 1 GB), there was 1 such gap in the first 22 minutes.
  2. It used too much memory. Prod memory jumped from 40% to 70% at 16:58 IST on 23 Sep, the minute NDTV's first S3 send went out. On 24 Sep AWS replaced the worker at least twice.

Why, part 1. The address falls back to localhost:

# packages/core/src/omnium_core/settings.py:298-302
    @property
    def api_url(self) -> str:
        """The public API, as the admin service should call it. No trailing slash."""
        found = self.omnium_api_url or self.public_api_base_url or "http://localhost:8000"
        return found.rstrip("/")

Why, part 2. Every S3 send builds a new session and client:

# packages/core/src/omnium_core/publishing/delivery/transports.py:169-176
        session = self._session_factory()
        try:
            async with session.client("s3", **client_kwargs) as s3:
                await s3.put_object(
                    Bucket=bucket,
                    Key=object_key,
                    Body=payload.body,
                    ContentType=payload.content_type,
                )

Measured on a laptop on 25 Sep: about 130 MB of memory that Python keeps per new client, against about 20 MB with one shared client.

The agreed design Agreed, to build. The delivery-worker alone builds every file, from the main database, by calling the resolvers directly, with no HTTP call to the public API (D1). One reused S3 client per destination (E4). A worker that stops beating raises an alert (A7).

Staging was not a safe place to rehearse

What happened. We wanted to try every risky prod fix on staging first. Staging could not carry that.

DateWhat happened
21 SepThe staging database, a 1 GB db.t4g.micro shared with another product, ran out of memory and restarted twice (05:13 and 12:40 UTC). Every "staging is down" report traced back to this box
21 SepConnecting prod's doorbells deleted staging's webhooks from all 10 shared Scout jobs
after 25 SepStaging held no new data (T7). Its workflows were never installed (T1). Prod hand fixes were never copied to it

Why the doorbell deleted staging's webhook. Prod's database was first copied from staging, so each prod integration row held staging's webhook id. Connecting a doorbell deletes the id the row holds:

# packages/admin/src/omnium_admin/routers/integrations.py:2186-2192
        created = await scout.create_webhook(
            integration.scout_base_url, integration.scout_job_id, url
        )
        if integration.webhook_id:
            # The old one may be gone already; either way Scout stops using it.
            with contextlib.suppress(scout.ScoutError):
                await scout.delete_webhook(integration.scout_base_url, integration.webhook_id)

Nothing checks that the old id belongs to this environment.

Where it landed. Partly agreed, mostly parked:

  • One Postgres major version in CI, staging and prod (DB15) Agreed, to build. Staging ran 18.3 when seen on 21 Sep; CI and laptops run 17.
  • A rehearsal of a full Games day on staging, and a bigger staging database, wait for Section 12 (testing) and Section 14 (deploy). Both are parked Parked. See Build order and what is parked.
  • The doorbell trap has no decision yet Parked. The working rule: after connecting a doorbell on either side, check the Scout job's webhooks and confirm both are there.

What went right

These held under real load and real mistakes. The agreed design keeps all of them.

What workedThe proof at the Games
Delivery keeps its state in the database and pollsClients caught up by themselves after outages and restarts. A restart lost nothing
Two safety switches on deliveryOur own test box was not on prod's allowed list, so prod refused every send to it from 22 Sep to the end, with "destination is not allowed here"
Pins: a hand edit wins until it is handed backOperators trusted it (as-built review). Before the 5 Oct script, a read-only check confirmed none of its 359 planned sides was locked by hand
A fixed key on every commandThe 5 Oct script sent each side with the key fix1005- plus the side id. On staging, the same key sent twice answered "duplicate" and applied once. On prod: 371 accepted, 0 duplicates, and a second run found nothing to send
One road: every change is a commandApart from 4 raw SQL changes, every hand change on prod went through console commands, each one logged and pinned
Scout's run grading, and omnium's guardsA run that failed Scout's checks was never used. A bad import was rolled back whole, with a plain-words reason
Snapshots before and after every prod change63 dated snapshot folders. Each fix was checked by comparing files: only the planned rows changed, and India stayed at 85
Saved GraphQL queries for client files7 file shapes served to 3 clients, with a transform per client where a client needed its own form

Working habits that held. These came out of the incidents above and were the rules by the end of the Games (issue list, 5 Oct):

  1. Never replay a whole import run on prod. It rewrites every unit in it.
  2. Before a prod change, list every row it touches. Take a snapshot before and after.
  3. Use staging first when it holds the same data.
  4. Raw SQL must set updated_at = now(), or the files do not rebuild.
  5. Check medals in the client file, not only in the console.
  6. A client switched off on purpose is named in every status update until it is back on.

The design turns several of these habits into rules in the code: T7 and I12 for the first, C11 for the second, and D9 for the sixth.