Skip to content

feat(climate-news): Sprint 3a — content-age probe (7d budget)#3598

Merged
koala73 merged 2 commits into
mainfrom
feat/health-readiness-content-age-sprint-3a-climate-news
May 8, 2026
Merged

feat(climate-news): Sprint 3a — content-age probe (7d budget)#3598
koala73 merged 2 commits into
mainfrom
feat/health-readiness-content-age-sprint-3a-climate-news

Conversation

@koala73

@koala73 koala73 commented May 4, 2026

Copy link
Copy Markdown
Owner

Sprint 3a of the 2026-05-04 health-readiness probe plan. Stacked on PR #3597 (Sprint 2 disease-outbreaks pilot).

Summary

  • Adds a content-age contract on seed-climate-news.mjs so /api/health surfaces STALE_CONTENT when the freshest cached climate-news item is older than 7 days.
  • Pure helper lives in scripts/_climate-news-helpers.mjs (climateNewsContentMeta with injectable nowMs + CLIMATE_NEWS_MAX_CONTENT_AGE_MIN). Seeder imports it, test imports it — no replicas, no drift (matches the Sprint 2 post-refactor pattern from PR feat(disease-outbreaks): Sprint 2 — content-age pilot (synthetic tagging + 9d budget) #3597).
  • Dockerfile.relay COPY line for the new helper (caught by tests/dockerfile-relay-imports.test.mjs).

Why 7 days

Carbon Brief, Guardian Environment, NASA EO, UNEP, Phys.org, Copernicus, Inside Climate News, Climate Central, and ReliefWeb collectively publish at multiple-times-per-day cadence. A 7-day budget tolerates a major holiday weekend across all sources without false-positive paging, and trips on a real upstream-aggregator outage (e.g. every feed's RSS regex stops matching because a publisher rebuilt their site).

Why no synthetic-tagging needed (unlike disease-outbreaks Sprint 2)

seed-climate-news.mjs:76 and :132 already drop items with publishedAt: 0 at parse time — `if (!title || !url || !publishedAt) return;`. So contentMeta reads item.publishedAt directly. No _publishedAtIsSynthetic/_originalPublishedMs helpers, no publishTransform stripping required.

Test plan

  • node --test tests/climate-news-content-age.test.mjs → 10/10 pass
  • node --test tests/dockerfile-relay-imports.test.mjs → 3/3 pass (closure satisfied after Dockerfile.relay update)
  • Full content-age stack: 59/59 across Sprint 1+2+3a
  • npm run typecheck:api clean
  • npm run lint clean (pre-existing warnings only)
  • Post-merge: confirm /api/health snapshot shows climateNewsIntelligence with contentAge block populated and status: OK
  • Post-merge: simulate stale (manually set newest item to 8d ago in a preview env) → confirm STALE_CONTENT

@vercel

vercel Bot commented May 4, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
worldmonitor Ready Ready Preview, Comment May 8, 2026 9:06am

Request Review

@greptile-apps

greptile-apps Bot commented May 4, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

Adds a 7-day content-age contract for the climate-news seeder by extracting a pure climateNewsContentMeta helper into scripts/_climate-news-helpers.mjs, wiring it into runSeed via contentMeta/maxContentAgeMin, and covering it with 10 deterministic unit tests. The pattern is a clean match to the Sprint 2 disease-outbreaks approach, and the data-shape reasoning (seeder already rejects publishedAt: 0 items at parse time) is sound.

Confidence Score: 4/5

Safe to merge; only finding is a wrong year in a test comment that has no impact on test correctness or runtime behavior.

All P2, no P1 or P0 findings. The single issue is a misleading year in a FIXED_NOW comment (2025 vs 2023) — functionally harmless but worth correcting before the codebase drifts further.

tests/climate-news-content-age.test.mjs — wrong year in FIXED_NOW comment on line 15.

Important Files Changed

Filename Overview
scripts/_climate-news-helpers.mjs New pure-helper module exporting climateNewsContentMeta and CLIMATE_NEWS_MAX_CONTENT_AGE_MIN; matches Sprint 2 disease-outbreaks pattern. Logic is correct: skips ts ≤ 0, non-finite, and future items beyond 1h clock-skew tolerance, returns null on zero valid items.
tests/climate-news-content-age.test.mjs 10 deterministic tests with injectable nowMs. One comment has the wrong year for FIXED_NOW (says 2025-11-14, is actually 2023-11-14); test logic itself is correct.
scripts/seed-climate-news.mjs Adds contentMeta and maxContentAgeMin to runSeed call using the extracted helper; import is correct and data shape (data.items) matches what fetchClimateNews returns.
Dockerfile.relay Adds COPY scripts/_climate-news-helpers.mjs before the seeder line; ordering is correct and satisfies the dockerfile-relay-imports test.

Flowchart

%%{init: {'theme': 'neutral'}}%%
flowchart TD
    A[runSeed called] --> B[fetchClimateNews returns data.items]
    B --> C[climateNewsContentMeta]
    C --> D[iterate data.items]
    D --> E{publishedAt valid?}
    E -- No --> F[skip item]
    E -- Yes --> G{ts beyond 1h skew?}
    G -- Yes --> F
    G -- No --> H[track newest and oldest]
    H --> I{validCount gt 0?}
    I -- No --> J[return null - STALE_CONTENT]
    I -- Yes --> K[return newestItemAt and oldestItemAt]
    K --> L{age gt maxContentAgeMin 7 days?}
    L -- Yes --> M[health status: STALE_CONTENT]
    L -- No --> N[health status: OK]
Loading

Reviews (1): Last reviewed commit: "feat(climate-news): Sprint 3a — content-..." | Re-trigger Greptile

Comment thread tests/climate-news-content-age.test.mjs Outdated
koala73 added a commit that referenced this pull request May 5, 2026
Unix timestamp 1700000000000 ms is 2023-11-14T22:13:20Z, not 2025-11-14.
Test correctness unaffected (FIXED_NOW is just an injected stable epoch),
but a reader reasoning about the skew-limit arithmetic would get the
mental date math wrong. Greptile P2 on PR #3598 (which copied the same
wrong comment from this file when Sprint 3a was branched off).
koala73 added a commit that referenced this pull request May 5, 2026
Greptile P2 on PR #3598: 1700000000000 ms is 2023-11-14T22:13:20Z, not
2025. Test correctness unaffected; comment-only fix so a reader reasoning
about skew-limit arithmetic gets the right mental date math.
@koala73
koala73 force-pushed the feat/health-readiness-content-age-sprint-3a-climate-news branch from ee2cd74 to e59d414 Compare May 5, 2026 06:48
koala73 added a commit that referenced this pull request May 8, 2026
Unix timestamp 1700000000000 ms is 2023-11-14T22:13:20Z, not 2025-11-14.
Test correctness unaffected (FIXED_NOW is just an injected stable epoch),
but a reader reasoning about the skew-limit arithmetic would get the
mental date math wrong. Greptile P2 on PR #3598 (which copied the same
wrong comment from this file when Sprint 3a was branched off).
@koala73
koala73 force-pushed the feat/health-readiness-content-age-sprint-2-disease-pilot branch from 1bcf7c5 to 287e3e6 Compare May 8, 2026 08:54
koala73 added a commit that referenced this pull request May 8, 2026
Sparse seeders sub-PR b/c of the 2026-05-04 health-readiness plan.
Branched off Sprint 1 (#3596) as a parallel sibling to Sprint 2 (#3597)
and Sprint 3a (#3598) per the plan's "Each PR is independently shippable"
note (line 498).

## Why this matters

IEA monthly oil stocks publish on an M+2 cadence — August data ships in
late October/early November. Without a content-age probe, a stalled
publication month is invisible to /api/health: the seeder runs fine on
its 6h cron, fetchedAt stays fresh, but data.dataMonth never advances. A
45-day budget trips STALE_CONTENT exactly when a month has been missed
(e.g. cache shows "2024-08" past Dec 1 when "2024-10" should have
landed).

## Shape contract — different from Sprint 2/3a

IEA is a SINGLE-SNAPSHOT seeder: every member shares one `dataMonth`
("YYYY-MM" string at the top level), there is no per-item published-at.
The new helper parses dataMonth → end-of-month UTC ms (the latest
possible observation date in the named period) and returns it as both
`newestItemAt` and `oldestItemAt`.

Defensive: contentMeta returns null when dataMonth is missing, malformed
("2024-13", "2024-8" single-digit), or future-dated beyond 1h clock-skew
tolerance (guards against upstream yearMonth garbage producing e.g. a
2099-12 dataMonth).

## Pattern parity with Sprint 2/3a

Following the established pattern: pure helpers in
`scripts/_iea-oil-stocks-helpers.mjs` (`dataMonthToEndOfMonthMs`,
`ieaOilStocksContentMeta`, `IEA_OIL_STOCKS_MAX_CONTENT_AGE_MIN`).
Seeder imports them; tests import them. No replicas.

`seed-iea-oil-stocks.mjs` is NOT in Dockerfile.relay (verified via
`grep`), so no COPY-line update needed (unlike Sprint 3a's
seed-climate-news which IS relay-COPY'd).

## Verification

- 15/15 iea content-age tests pass (incl. leap-year, month-rollover,
  invalid-shape rejection, M+2 lag realism, future-clock-skew defense)
- 78/78 across iea seed + Sprint 1 + Sprint 3b stack
- typecheck:api clean; lint clean (pre-existing warnings only)
- Dockerfile.relay closure test passes (no relay impact)
koala73 added a commit that referenced this pull request May 8, 2026
…ing + 9d budget) (#3597)

* feat(disease-outbreaks): Sprint 2 — content-age pilot (synthetic tagging + STALE_CONTENT @ 9d)

Implements Sprint 2 of the 2026-05-04 health-readiness plan
(docs/plans/2026-05-04-001-feat-health-readiness-probe-content-age-plan.md).
Stacked on Sprint 1 (#3596 — content-age probe infra).

Migrates disease-outbreaks as the proof-of-concept content-age consumer.
Pilot maxContentAgeMin=9 days chosen so the 2026-05-04 11d-old incident
would have correctly tripped STALE_CONTENT.

== Source-parser changes (3 sources, uniform shape) ==

scripts/seed-disease-outbreaks.mjs:

WHO DON parser (line ~117): tag synthetic timestamps when the upstream
omits PublicationDateAndTime. Carry _originalPublishedMs (parsed ms or
null) and _publishedAtIsSynthetic (boolean) alongside the existing
publishedMs (which keeps its Date.now() fallback for UI consumer compat).

RSS parser (line ~150, both CDC and Outbreak News Today): same pattern
when pubDate is missing/unparseable.

TGH parser (line ~211): always carries non-synthetic since the line-198
filter rejects undated items earlier. Migration is additive — every
TGH item gets _publishedAtIsSynthetic: false and _originalPublishedMs:
publishedMs so contentMeta + publishTransform apply uniformly.

mapItem (line ~244): carries _publishedAtIsSynthetic and
_originalPublishedMs through to the output shape so contentMeta can
read them at runSeed time.

== runSeed opts (Sprint 2 contract) ==

contentMeta: excludes _publishedAtIsSynthetic items + 1h clock-skew
tolerance + null when validCount === 0 (matches list-feed-digest's
FUTURE_DATE_TOLERANCE_MS pattern).

maxContentAgeMin: 9 * 24 * 60 = 12960 minutes (9 days) — chosen
deliberately so the production incident's 11d-old cache would have
flagged STALE_CONTENT. Tighter would page on normal WHO/CDC quiet
weeks; looser would have missed the incident.

publishTransform: strips _publishedAtIsSynthetic + _originalPublishedMs
from every item BEFORE atomicPublish so the helpers never reach:
  - the Redis canonical key (health:disease-outbreaks:v1)
  - /api/bootstrap response (data.diseaseOutbreaks)
  - list-disease-outbreaks RPC response
  - the DiseaseOutbreakItem proto-generated type

The Sprint 1 ordering contract (contentMeta runs BEFORE publishTransform)
guarantees contentMeta sees the helpers that publishTransform then strips.

== Anti-regression tests ==

tests/disease-outbreaks-seed.test.mjs (NEW) — 16 cases split by layer:

Pre-publish (in-memory) layer (5):
- WHO without PublicationDateAndTime → tagged synthetic
- WHO with valid PublicationDateAndTime → non-synthetic
- RSS without pubDate → tagged synthetic
- RSS with valid pubDate → non-synthetic
- TGH always non-synthetic

contentMeta behavior (5):
- All-synthetic → null (→ STALE_CONTENT)
- Mixed: synthetic with newer publishedAt does NOT win newest
- Picks newest+oldest from non-synthetic set
- Future-dated items beyond 1h tolerance excluded
- NEAR_FUTURE within 1h tolerance accepted

publishTransform strip (3):
- Both helper fields stripped from every item
- publishedAt remains non-null (UI/RPC consumer contract)
- Empty + missing outbreaks handled safely

End-to-end (1):
- contentMeta runs on raw data WITH helpers, publishTransform strips,
  canonical-shape JSON contains NEITHER _publishedAtIsSynthetic NOR
  _originalPublishedMs (combined-regex assertion per Codex round 4 P2)

Pilot threshold sanity (2):
- 11d-old items DO trip the 9d budget (anti-drift on the pilot threshold —
  any future change to 9d must update this test)
- 5d-old items DO NOT trip (no false positive on normal upstream rhythm)

Test totals: 95/95 pass across the seed-envelope, seed-contract,
seed-utils, content-age, health-content-age, and disease-outbreaks-seed
suites.

== Verification (post-deploy) ==

After Railway bundle redeploy:
1. /api/health.diseaseOutbreaks shows contentAgeMin and maxContentAgeMin.
2. Redis canonical health:disease-outbreaks:v1 contains NEITHER
   _publishedAtIsSynthetic NOR _originalPublishedMs (combined-regex
   grep returns 0).
3. /api/bootstrap?keys=diseaseOutbreaks response payload helper-free.
4. With current 11d-old WHO/CDC items + bug-pattern data, STALE_CONTENT
   surfaces in /api/health and ops can act on it.

* refactor(disease-outbreaks): extract helpers + inject nowMs to kill test drift and timing flake

Greptile P2s on PR #3597:

1. tests/disease-outbreaks-seed.test.mjs replicated parser/mapper/contentMeta
   logic locally — a drift in fetchWhoDonApi or contentMeta would not have
   failed any of the 16 tests because they asserted against their own copy
   of the logic, not the seeder's.

2. The "near-future ≤1h accepted" test relied on Date.now() being stable
   between test setup and the call into contentMeta. On a loaded CI runner
   the gap could exceed the (1h - 30min) margin and flake.

Fixes both at once:

- New scripts/_disease-outbreaks-helpers.mjs exports the pure functions
  (whoNormalizeItem, rssNormalizeItem, tghNormalizeItem, mapItem,
  diseaseContentMeta, diseasePublishTransform, DISEASE_MAX_CONTENT_AGE_MIN).
  diseaseContentMeta accepts an optional nowMs for deterministic skew tests.

- Seeder imports those helpers instead of inlining them. ~150 lines
  removed; behavior unchanged (verified by node -c + smoke test).

- Test file imports the real helpers (no replicas). All skew-limit tests
  inject FIXED_NOW=1700000000000 — no wall-clock dependence.

- Tightens the "within 1h tolerance" test from +30min to +5min ahead of
  injected NOW, well clear of the 1h boundary regardless of the timing fix.

Net: -265 lines across the two existing files; +200 in the new helpers
module. 17/17 disease tests pass; 49/49 across the full Sprint 1+2 stack.

* fix(test): correct FIXED_NOW comment year (2025→2023)

Unix timestamp 1700000000000 ms is 2023-11-14T22:13:20Z, not 2025-11-14.
Test correctness unaffected (FIXED_NOW is just an injected stable epoch),
but a reader reasoning about the skew-limit arithmetic would get the
mental date math wrong. Greptile P2 on PR #3598 (which copied the same
wrong comment from this file when Sprint 3a was branched off).
Base automatically changed from feat/health-readiness-content-age-sprint-2-disease-pilot to main May 8, 2026 09:00
koala73 added 2 commits May 8, 2026 13:01
Sparse seeders sub-PR a/c of the 2026-05-04 health-readiness plan. Adds a
content-age contract on seed-climate-news.mjs so /api/health surfaces
STALE_CONTENT when the freshest cached climate-news item is older than 7
days — covering the failure mode where every RSS parse silently breaks at
once (e.g. our regex stops matching because a feed bundle changed) and the
seeder keeps running clean while the cache fossilizes.

Why 7 days: Carbon Brief, Guardian Environment, NASA EO, UNEP, Phys.org,
Copernicus, Inside Climate News, Climate Central, and ReliefWeb publish
collectively at multiple-times-per-day cadence. A 7d budget tolerates a
major holiday weekend across all sources without false-positive paging,
and trips on a real upstream-aggregator outage.

Why no synthetic-tagging needed (unlike disease-outbreaks Sprint 2):
seed-climate-news.mjs:76 + :132 already drop items with publishedAt=0 at
parse time, so contentMeta reads item.publishedAt directly. No helper
fields, no publishTransform stripping required.

Following the Sprint 2 post-refactor pattern: pure helper lives in
scripts/_climate-news-helpers.mjs (climateNewsContentMeta with injectable
nowMs for deterministic tests + CLIMATE_NEWS_MAX_CONTENT_AGE_MIN constant).
The seeder imports it; the test imports it. No duplicated logic, no drift
surface.

Verification: 10/10 climate-news tests pass; 59/59 across the full
content-age stack (Sprint 1 infra + Sprint 2 disease + Sprint 3a climate).
typecheck:api clean; lint clean (pre-existing warnings only).
Greptile P2 on PR #3598: 1700000000000 ms is 2023-11-14T22:13:20Z, not
2025. Test correctness unaffected; comment-only fix so a reader reasoning
about skew-limit arithmetic gets the right mental date math.
@koala73
koala73 force-pushed the feat/health-readiness-content-age-sprint-3a-climate-news branch from e59d414 to dd3afa1 Compare May 8, 2026 09:02
@koala73
koala73 merged commit 4304b3c into main May 8, 2026
11 checks passed
@koala73
koala73 deleted the feat/health-readiness-content-age-sprint-3a-climate-news branch May 8, 2026 09:06
koala73 added a commit that referenced this pull request May 8, 2026
* feat(iea-oil-stocks): Sprint 3b — content-age probe (45d budget)

Sparse seeders sub-PR b/c of the 2026-05-04 health-readiness plan.
Branched off Sprint 1 (#3596) as a parallel sibling to Sprint 2 (#3597)
and Sprint 3a (#3598) per the plan's "Each PR is independently shippable"
note (line 498).

## Why this matters

IEA monthly oil stocks publish on an M+2 cadence — August data ships in
late October/early November. Without a content-age probe, a stalled
publication month is invisible to /api/health: the seeder runs fine on
its 6h cron, fetchedAt stays fresh, but data.dataMonth never advances. A
45-day budget trips STALE_CONTENT exactly when a month has been missed
(e.g. cache shows "2024-08" past Dec 1 when "2024-10" should have
landed).

## Shape contract — different from Sprint 2/3a

IEA is a SINGLE-SNAPSHOT seeder: every member shares one `dataMonth`
("YYYY-MM" string at the top level), there is no per-item published-at.
The new helper parses dataMonth → end-of-month UTC ms (the latest
possible observation date in the named period) and returns it as both
`newestItemAt` and `oldestItemAt`.

Defensive: contentMeta returns null when dataMonth is missing, malformed
("2024-13", "2024-8" single-digit), or future-dated beyond 1h clock-skew
tolerance (guards against upstream yearMonth garbage producing e.g. a
2099-12 dataMonth).

## Pattern parity with Sprint 2/3a

Following the established pattern: pure helpers in
`scripts/_iea-oil-stocks-helpers.mjs` (`dataMonthToEndOfMonthMs`,
`ieaOilStocksContentMeta`, `IEA_OIL_STOCKS_MAX_CONTENT_AGE_MIN`).
Seeder imports them; tests import them. No replicas.

`seed-iea-oil-stocks.mjs` is NOT in Dockerfile.relay (verified via
`grep`), so no COPY-line update needed (unlike Sprint 3a's
seed-climate-news which IS relay-COPY'd).

## Verification

- 15/15 iea content-age tests pass (incl. leap-year, month-rollover,
  invalid-shape rejection, M+2 lag realism, future-clock-skew defense)
- 78/78 across iea seed + Sprint 1 + Sprint 3b stack
- typecheck:api clean; lint clean (pre-existing warnings only)
- Dockerfile.relay closure test passes (no relay impact)

* fix(iea-oil-stocks): bump budget 45d→90d to cover M+2 natural lag

Greptile P1 on PR #3599: a 45-day budget contradicts the helper's own
M+2 cadence claim. End-of-observation-month (Aug 31) is ~60-65 days
BEFORE publication (~late Oct/early Nov), so fresh-arrival data is
already past the 45d threshold at the moment a successful seed run
writes it. STALE_CONTENT would have fired on every cron tick.

Corrected math: 90d = ~60d natural M+2 lag + ~30d missed-publication
slack. Trips only when a month is missed entirely (cache stuck at
"2024-08" past mid-Jan when "2024-10" should have landed).

Also addresses 3 P2 review nits in the same edit:

- Test "60 days old" → "fresh-arrival regression guard: ~60d-old fresh
  M+2 data does NOT trip" (the math was right, name was wrong; rewrote
  the test to actually pin the failure mode the P1 cited).
- Test "~30 days old" → "~14 days old" (the fixture was "2023-10" =
  ~14d before FIXED_NOW, not 30).
- M+2 lag scenario comment "Sept data published ~Oct 25" → "~late Nov
  (M+2 cadence)" — Oct 25 is M+1, not M+2.

Added: dedicated fresh-arrival regression guard test that asserts a
~75d-old fresh M+2 dataMonth is within budget. Without it, a future
budget tightening could re-introduce the immediate-page bug invisibly.

Verification: 16/16 iea content-age (was 15/15 — added regression guard);
79/79 across iea seed + Sprint 1 + Sprint 3b stack; typecheck:api clean.
fuleinist pushed a commit to fuleinist/worldmonitor that referenced this pull request May 9, 2026
…ing + 9d budget) (koala73#3597)

* feat(disease-outbreaks): Sprint 2 — content-age pilot (synthetic tagging + STALE_CONTENT @ 9d)

Implements Sprint 2 of the 2026-05-04 health-readiness plan
(docs/plans/2026-05-04-001-feat-health-readiness-probe-content-age-plan.md).
Stacked on Sprint 1 (koala73#3596 — content-age probe infra).

Migrates disease-outbreaks as the proof-of-concept content-age consumer.
Pilot maxContentAgeMin=9 days chosen so the 2026-05-04 11d-old incident
would have correctly tripped STALE_CONTENT.

== Source-parser changes (3 sources, uniform shape) ==

scripts/seed-disease-outbreaks.mjs:

WHO DON parser (line ~117): tag synthetic timestamps when the upstream
omits PublicationDateAndTime. Carry _originalPublishedMs (parsed ms or
null) and _publishedAtIsSynthetic (boolean) alongside the existing
publishedMs (which keeps its Date.now() fallback for UI consumer compat).

RSS parser (line ~150, both CDC and Outbreak News Today): same pattern
when pubDate is missing/unparseable.

TGH parser (line ~211): always carries non-synthetic since the line-198
filter rejects undated items earlier. Migration is additive — every
TGH item gets _publishedAtIsSynthetic: false and _originalPublishedMs:
publishedMs so contentMeta + publishTransform apply uniformly.

mapItem (line ~244): carries _publishedAtIsSynthetic and
_originalPublishedMs through to the output shape so contentMeta can
read them at runSeed time.

== runSeed opts (Sprint 2 contract) ==

contentMeta: excludes _publishedAtIsSynthetic items + 1h clock-skew
tolerance + null when validCount === 0 (matches list-feed-digest's
FUTURE_DATE_TOLERANCE_MS pattern).

maxContentAgeMin: 9 * 24 * 60 = 12960 minutes (9 days) — chosen
deliberately so the production incident's 11d-old cache would have
flagged STALE_CONTENT. Tighter would page on normal WHO/CDC quiet
weeks; looser would have missed the incident.

publishTransform: strips _publishedAtIsSynthetic + _originalPublishedMs
from every item BEFORE atomicPublish so the helpers never reach:
  - the Redis canonical key (health:disease-outbreaks:v1)
  - /api/bootstrap response (data.diseaseOutbreaks)
  - list-disease-outbreaks RPC response
  - the DiseaseOutbreakItem proto-generated type

The Sprint 1 ordering contract (contentMeta runs BEFORE publishTransform)
guarantees contentMeta sees the helpers that publishTransform then strips.

== Anti-regression tests ==

tests/disease-outbreaks-seed.test.mjs (NEW) — 16 cases split by layer:

Pre-publish (in-memory) layer (5):
- WHO without PublicationDateAndTime → tagged synthetic
- WHO with valid PublicationDateAndTime → non-synthetic
- RSS without pubDate → tagged synthetic
- RSS with valid pubDate → non-synthetic
- TGH always non-synthetic

contentMeta behavior (5):
- All-synthetic → null (→ STALE_CONTENT)
- Mixed: synthetic with newer publishedAt does NOT win newest
- Picks newest+oldest from non-synthetic set
- Future-dated items beyond 1h tolerance excluded
- NEAR_FUTURE within 1h tolerance accepted

publishTransform strip (3):
- Both helper fields stripped from every item
- publishedAt remains non-null (UI/RPC consumer contract)
- Empty + missing outbreaks handled safely

End-to-end (1):
- contentMeta runs on raw data WITH helpers, publishTransform strips,
  canonical-shape JSON contains NEITHER _publishedAtIsSynthetic NOR
  _originalPublishedMs (combined-regex assertion per Codex round 4 P2)

Pilot threshold sanity (2):
- 11d-old items DO trip the 9d budget (anti-drift on the pilot threshold —
  any future change to 9d must update this test)
- 5d-old items DO NOT trip (no false positive on normal upstream rhythm)

Test totals: 95/95 pass across the seed-envelope, seed-contract,
seed-utils, content-age, health-content-age, and disease-outbreaks-seed
suites.

== Verification (post-deploy) ==

After Railway bundle redeploy:
1. /api/health.diseaseOutbreaks shows contentAgeMin and maxContentAgeMin.
2. Redis canonical health:disease-outbreaks:v1 contains NEITHER
   _publishedAtIsSynthetic NOR _originalPublishedMs (combined-regex
   grep returns 0).
3. /api/bootstrap?keys=diseaseOutbreaks response payload helper-free.
4. With current 11d-old WHO/CDC items + bug-pattern data, STALE_CONTENT
   surfaces in /api/health and ops can act on it.

* refactor(disease-outbreaks): extract helpers + inject nowMs to kill test drift and timing flake

Greptile P2s on PR koala73#3597:

1. tests/disease-outbreaks-seed.test.mjs replicated parser/mapper/contentMeta
   logic locally — a drift in fetchWhoDonApi or contentMeta would not have
   failed any of the 16 tests because they asserted against their own copy
   of the logic, not the seeder's.

2. The "near-future ≤1h accepted" test relied on Date.now() being stable
   between test setup and the call into contentMeta. On a loaded CI runner
   the gap could exceed the (1h - 30min) margin and flake.

Fixes both at once:

- New scripts/_disease-outbreaks-helpers.mjs exports the pure functions
  (whoNormalizeItem, rssNormalizeItem, tghNormalizeItem, mapItem,
  diseaseContentMeta, diseasePublishTransform, DISEASE_MAX_CONTENT_AGE_MIN).
  diseaseContentMeta accepts an optional nowMs for deterministic skew tests.

- Seeder imports those helpers instead of inlining them. ~150 lines
  removed; behavior unchanged (verified by node -c + smoke test).

- Test file imports the real helpers (no replicas). All skew-limit tests
  inject FIXED_NOW=1700000000000 — no wall-clock dependence.

- Tightens the "within 1h tolerance" test from +30min to +5min ahead of
  injected NOW, well clear of the 1h boundary regardless of the timing fix.

Net: -265 lines across the two existing files; +200 in the new helpers
module. 17/17 disease tests pass; 49/49 across the full Sprint 1+2 stack.

* fix(test): correct FIXED_NOW comment year (2025→2023)

Unix timestamp 1700000000000 ms is 2023-11-14T22:13:20Z, not 2025-11-14.
Test correctness unaffected (FIXED_NOW is just an injected stable epoch),
but a reader reasoning about the skew-limit arithmetic would get the
mental date math wrong. Greptile P2 on PR koala73#3598 (which copied the same
wrong comment from this file when Sprint 3a was branched off).
fuleinist pushed a commit to fuleinist/worldmonitor that referenced this pull request May 9, 2026
…3#3598)

* feat(climate-news): Sprint 3a — content-age probe (7d budget)

Sparse seeders sub-PR a/c of the 2026-05-04 health-readiness plan. Adds a
content-age contract on seed-climate-news.mjs so /api/health surfaces
STALE_CONTENT when the freshest cached climate-news item is older than 7
days — covering the failure mode where every RSS parse silently breaks at
once (e.g. our regex stops matching because a feed bundle changed) and the
seeder keeps running clean while the cache fossilizes.

Why 7 days: Carbon Brief, Guardian Environment, NASA EO, UNEP, Phys.org,
Copernicus, Inside Climate News, Climate Central, and ReliefWeb publish
collectively at multiple-times-per-day cadence. A 7d budget tolerates a
major holiday weekend across all sources without false-positive paging,
and trips on a real upstream-aggregator outage.

Why no synthetic-tagging needed (unlike disease-outbreaks Sprint 2):
seed-climate-news.mjs:76 + :132 already drop items with publishedAt=0 at
parse time, so contentMeta reads item.publishedAt directly. No helper
fields, no publishTransform stripping required.

Following the Sprint 2 post-refactor pattern: pure helper lives in
scripts/_climate-news-helpers.mjs (climateNewsContentMeta with injectable
nowMs for deterministic tests + CLIMATE_NEWS_MAX_CONTENT_AGE_MIN constant).
The seeder imports it; the test imports it. No duplicated logic, no drift
surface.

Verification: 10/10 climate-news tests pass; 59/59 across the full
content-age stack (Sprint 1 infra + Sprint 2 disease + Sprint 3a climate).
typecheck:api clean; lint clean (pre-existing warnings only).

* fix(test): correct FIXED_NOW comment year (2025→2023)

Greptile P2 on PR koala73#3598: 1700000000000 ms is 2023-11-14T22:13:20Z, not
2025. Test correctness unaffected; comment-only fix so a reader reasoning
about skew-limit arithmetic gets the right mental date math.
fuleinist pushed a commit to fuleinist/worldmonitor that referenced this pull request May 9, 2026
…la73#3599)

* feat(iea-oil-stocks): Sprint 3b — content-age probe (45d budget)

Sparse seeders sub-PR b/c of the 2026-05-04 health-readiness plan.
Branched off Sprint 1 (koala73#3596) as a parallel sibling to Sprint 2 (koala73#3597)
and Sprint 3a (koala73#3598) per the plan's "Each PR is independently shippable"
note (line 498).

## Why this matters

IEA monthly oil stocks publish on an M+2 cadence — August data ships in
late October/early November. Without a content-age probe, a stalled
publication month is invisible to /api/health: the seeder runs fine on
its 6h cron, fetchedAt stays fresh, but data.dataMonth never advances. A
45-day budget trips STALE_CONTENT exactly when a month has been missed
(e.g. cache shows "2024-08" past Dec 1 when "2024-10" should have
landed).

## Shape contract — different from Sprint 2/3a

IEA is a SINGLE-SNAPSHOT seeder: every member shares one `dataMonth`
("YYYY-MM" string at the top level), there is no per-item published-at.
The new helper parses dataMonth → end-of-month UTC ms (the latest
possible observation date in the named period) and returns it as both
`newestItemAt` and `oldestItemAt`.

Defensive: contentMeta returns null when dataMonth is missing, malformed
("2024-13", "2024-8" single-digit), or future-dated beyond 1h clock-skew
tolerance (guards against upstream yearMonth garbage producing e.g. a
2099-12 dataMonth).

## Pattern parity with Sprint 2/3a

Following the established pattern: pure helpers in
`scripts/_iea-oil-stocks-helpers.mjs` (`dataMonthToEndOfMonthMs`,
`ieaOilStocksContentMeta`, `IEA_OIL_STOCKS_MAX_CONTENT_AGE_MIN`).
Seeder imports them; tests import them. No replicas.

`seed-iea-oil-stocks.mjs` is NOT in Dockerfile.relay (verified via
`grep`), so no COPY-line update needed (unlike Sprint 3a's
seed-climate-news which IS relay-COPY'd).

## Verification

- 15/15 iea content-age tests pass (incl. leap-year, month-rollover,
  invalid-shape rejection, M+2 lag realism, future-clock-skew defense)
- 78/78 across iea seed + Sprint 1 + Sprint 3b stack
- typecheck:api clean; lint clean (pre-existing warnings only)
- Dockerfile.relay closure test passes (no relay impact)

* fix(iea-oil-stocks): bump budget 45d→90d to cover M+2 natural lag

Greptile P1 on PR koala73#3599: a 45-day budget contradicts the helper's own
M+2 cadence claim. End-of-observation-month (Aug 31) is ~60-65 days
BEFORE publication (~late Oct/early Nov), so fresh-arrival data is
already past the 45d threshold at the moment a successful seed run
writes it. STALE_CONTENT would have fired on every cron tick.

Corrected math: 90d = ~60d natural M+2 lag + ~30d missed-publication
slack. Trips only when a month is missed entirely (cache stuck at
"2024-08" past mid-Jan when "2024-10" should have landed).

Also addresses 3 P2 review nits in the same edit:

- Test "60 days old" → "fresh-arrival regression guard: ~60d-old fresh
  M+2 data does NOT trip" (the math was right, name was wrong; rewrote
  the test to actually pin the failure mode the P1 cited).
- Test "~30 days old" → "~14 days old" (the fixture was "2023-10" =
  ~14d before FIXED_NOW, not 30).
- M+2 lag scenario comment "Sept data published ~Oct 25" → "~late Nov
  (M+2 cadence)" — Oct 25 is M+1, not M+2.

Added: dedicated fresh-arrival regression guard test that asserts a
~75d-old fresh M+2 dataMonth is within budget. Without it, a future
budget tightening could re-introduce the immediate-page bug invisibly.

Verification: 16/16 iea content-age (was 15/15 — added regression guard);
79/79 across iea seed + Sprint 1 + Sprint 3b stack; typecheck:api clean.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant