Infectious disease modelling in the age of AI

Sam Abbott

London School of Hygiene & Tropical Medicine

14 September 2026

Talk plan

  • Ten years of outbreak modelling
  • Use of AI methods in outbreak modelling
  • Agentic AI and outbreak modelling

samabbott.co.uk/AIandIDVUB/keynote

Combining Infectious Disease Modelling and AI, VUB, Brussels, Monday 14 September 2026, 09:00.

Ten years of outbreak modelling

Two and a half centuries of outbreak models

A timeline from 1760 to 2026. Grey dots for Bernoulli 1760, Ross 1911, Kermack and McKendrick 1927, HIV/AIDS back-calculation 1988, Anderson and May 1991, foot and mouth 2001, H1N1 2009 and Ebola in West Africa 2014. Teal dots for COVID-19 2020, mpox 2022 and Ebola disease caused by Bundibugyo virus 2026.

Bernoulli 1760, after Dietz and Heesterbeek 2002, Math Biosci · Ross 1911, The Prevention of Malaria · Kermack and McKendrick 1927, Proc R Soc A · Brookmeyer and Gail 1988, JASA · Anderson and May 1991, OUP · Ferguson, Donnelly and Anderson 2001, Science · Fraser et al. 2009, Science · WHO notified of Ebola in Guinea, 23 March 2014 · mpox, WHO, 23 July 2022 · Bundibugyo virus, WHO DON602, 17 May 2026

My last ten years

A timeline from 2014 to 2026. A grey bar for Ebola in West Africa 2014 to 2016, a teal dot for the Wuhan estimates in January 2020, a teal bar for the Rt dashboard and SPI-M-O from 2020 to 2022 with a slate segment for variants in 2021, a teal bar for mpox in 2022 and a brick bar for the Bundibugyo virus outbreak from May 2026.
  • Ebola, West Africa. Watched from a PhD on tuberculosis vaccination
  • COVID-19. Early Wuhan estimates, then \(R_t\), the reproduction number over time, daily for several thousand places. Variant transmissibility
  • mpox. Nowcasting for the UK Health Security Agency (UKHSA) and heavy-tailed sexual contact networks
  • Ebola disease caused by Bundibugyo virus. A joint model of the streams in the situation reports, refit on every push

Started at LSHTM in January 2020, 4 January 2023 post at samabbott.co.uk

Ebola, West Africa. CDC’s projection

  • Meltzer et al., US Centers for Disease Control and Prevention. Exponential growth without further intervention
  • 550,000 cases by 20 January 2015, 1.4 million corrected for under-reporting
  • About 21,000 had occurred by mid January

Two line graphs from the CDC EbolaResponse model, cumulative cases and daily beds in use in Liberia and Sierra Leone combined, March 2014 to January 2015. Each has an uncorrected curve reaching about 550,000 cases and a corrected curve reaching about 1.4 million cases by 20 January 2015, both rising slowly then very steeply.

Meltzer et al. 2014, MMWR 63 Suppl 3 · CDC 2016, MMWR 65 Suppl 3, on the 21,000 · figure: Appendix Figure 2, Meltzer et al. 2014, MMWR 63 Suppl 3, public domain, US government work

Ebola, West Africa. Real-time forecasts in Sierra Leone

  • Camacho, Funk and colleagues. Estimated \(R\), new infections per infected person, by district in Sierra Leone and forecast weekly with a semi-mechanistic model
  • Funk et al. scored the forecasts in 2019. Calibrated one and two weeks ahead
  • Three weeks or more ahead they overestimated, by more the further out they looked

Weekly Ebola incidence in Western Area, Sierra Leone, from September 2014 to February 2015 as a thick black line, with one-week-ahead forecasts as dots joined by a thin line and grey credible bands, peaking near 350 cases a week in December 2014.

Camacho et al. 2015, PLOS Currents Outbreaks · Funk et al. 2018, Epidemics · figure: panel B of figure 2, Funk et al. 2019, PLOS Comput Biol, doi:10.1371/journal.pcbi.1006785, CC BY

Epidemiological delays from the WHO Ebola Response Team

  • Case investigation forms for 4,507 probable and confirmed cases by 14 September 2014
  • Incubation period, serial interval and case fatality from the forms, by ad hoc methods. Park et al. showed in 2024 that these lead to biased estimates
  • Their forecast. More than 20,000 cases by 2 November if control did not change. 14,383 had been reported by 11 November
  • The models on the following slides lean on delay estimates like these

Three events on a line, infection, symptom onset and onset of the next case, with the incubation period of mean 11.4 days and the serial interval of mean 15.3 days marked beneath, and a box giving case fatality of 70.8 percent.

WHO Ebola Response Team 2014, NEJM 371, 1481, doi:10.1056/NEJMoa1411100 · Park et al. 2024, medRxiv, doi:10.1101/2024.01.12.24301247 · 14,383 cases from CDC 2014, MMWR 63(46), 1130

Early Wuhan nCoV-19 estimates

  • 12 January 2020. We started on 2019-nCoV, with sparse, delayed case reports from China and detected cases elsewhere
  • 3 February. \(R_0\) in Wuhan from the size and duration of the outbreak under a range of serial intervals
  • 28 February. Whether isolating cases and their contacts could control new outbreaks, by simulation
  • By March we were estimating \(R_t\) for China, Italy, the UK and a growing list of other countries

A grid of density plots of the reproduction number, rows for MERS-like, SARS-like and pre-intervention SARS-like serial intervals, columns for outbreak duration in days, ridges for transmission event sizes from 20 to 400. Estimates run from about 1 to 4.

Figure 2 of Abbott et al. 2020, Wellcome Open Research 5:17, doi:10.12688/wellcomeopenres.15718.1, CC BY · Hellewell et al. 2020, Lancet Global Health · the 4 January 2023 post at samabbott.co.uk

Real-time \(R_t\) for thousands of locations

  • \(R_t\) daily for several thousand national and subnational locations, on a public dashboard from spring 2020 to March 2022
  • EpiNow2. A renewal process, where today’s infections are \(R_t\) times a weighted sum of recent infections, with delays to report
  • Weekly to SPI-M-O, the UK government’s modelling advisers. Other governments used the dashboard directly or through their scientists
  • Software widely used by researchers and public health agencies worldwide
  • In 2022 I wrote: “we sometimes published nonsensical estimates”

Four panels of estimated reproduction number over time for Australia, France, Germany and South Korea, each a green ribbon around a dashed line at one, with an orange widening ribbon for the most recent days where the estimate is still uncertain.

Abbott et al. 2020, Wellcome Open Research, doi:10.12688/wellcomeopenres.16006.1 · epiforecasts.io/covid · figure and quote from the 25 March 2022 post at samabbott.co.uk · SPI-M-O consensus statements on GOV.UK

What we struggled to do in real-time

  • Variants. Alpha in late 2020, Delta in 2021. The routine model could not estimate their transmissibility on the timeline policy needed
  • Chains. So we went back to using the output of one model as the input to other, simple models
  • The data changed under us. Tests to hospitalisations in 2021, tests unreliable and sparse by 2022, the Office for National Statistics survey and sequences arriving
  • Omicron. Most of the specialised models of 2021 could not take its immune escape and shorter generation time. Chains again

Four panels from the ONS Community Infection Survey model. Prevalence, incident infections, antibody prevalence and the reproduction number from mid 2020 to November 2021, each a fitted line with a fan of posterior draws.

The pandemic years from the 4 January 2023 post at samabbott.co.uk · figure: Abbott and Funk 2022, medRxiv, doi:10.1101/2022.03.29.22273101, prevalence to incidence to antibodies to \(R_t\) fitted as one model

The tools that came out of the pandemic, 2020 to 2021

A horizontal timeline from 2020 to 2022 with lollipops for scoringutils, EpiNow and the Rt dashboard, EpiNow2, the European COVID-19 Forecast Hub and epinowcast, each at its first commit or repository creation month.

First commits from the local clones, repository creation dates from the GitHub API, 11 September 2026

mpox 2022. Heavy-tailed sexual contact networks

  • A global outbreak. A public health emergency of international concern on 23 July 2022
  • More than 16,000 cases in 75 countries by then
  • Endo et al. A heavy-tailed sexual partnership distribution explained the sustained growth among men who have sex with men
  • Why it turned over. Depletion of the high-activity core. The counter-argument was behaviour change. Both sides agreed heavy-tailed networks played a part
  • Reuse. Existing COVID-19 models lacked the flexibility to adapt. New, relatively simple, single-source models were built instead

Four panels. Risk of an outbreak of at least 728 cases against the secondary attack rate per sexually associated contact, for one to five initial cases among men who have sex with men with sexual or other exposure, for general index cases, and outbreak size in a non-MSM sexual network for a thousand initial cases.

Endo et al. 2022, Science 378, 90, doi:10.1126/science.add4507 · Murayama, Pearson, Abbott et al. 2024, J Infect Dis, doi:10.1093/infdis/jiad254 · WHO Director-General’s statement, 23 July 2022

Nowcasting for the UK Health Security Agency

  • Nowcasting. Estimating what has already happened but has not yet been reported. Cases by symptom onset arrive late, so the most recent days look like a fall
  • Overton, Abbott and colleagues nowcast the England outbreak for UKHSA
  • UKHSA could not install other nowcasting methods, so the models were built on standard regression packages

Two panels of daily mpox cases in England by symptom onset. Red dots are the counts known on the nowcast date, blue dots the counts eventually reported, a dashed black line the nowcast median and a grey band its interval. The red dots fall to zero in the most recent days while the nowcast and the later data do not.

Figure 7 of Overton, Abbott et al. 2023, PLOS Comput Biol, doi:10.1371/journal.pcbi.1011463, CC BY

The tools since, 2022 to 2026

A horizontal timeline from 2022 to 2026 with lollipops for epidist, EpiAware.jl, primarycensored, baselinenowcast and BVDOutbreakSize, each at its first commit or repository creation month.

First commits from the local clones, repository creation dates from the GitHub API, 11 September 2026

Ebola disease caused by Bundibugyo virus, DRC, 2026

  • WHO alerted on 5 May to a high-mortality illness in Mongbwalu, Ituri. The Democratic Republic of the Congo declared its 17th Ebola outbreak on 15 May. A PHEIC on 17 May
  • 6,757 confirmed cases and 3,267 deaths in DRC by 7 September. 20 cases and 2 deaths in Uganda
  • Six provinces and 61 health zones. The largest Ebola disease outbreak recorded in DRC
  • The data arrive as a daily PDF situation report, in French

The front page of INSP situation report 87 of 9 August 2026, in French, with headline tiles for 4,381 confirmed cases, 2,011 confirmed deaths, 45.9 percent case fatality, 704 patients in isolation and 869 recovered, and a table of cases and deaths by province.

WHO DON602, DON616 and DON617, the last published 10 September 2026 with data to 7 September · INSP SitRep 087, 9 August 2026, insp.cd

McCabe et al., and a replication the next day

  • McCabe et al., Imperial, 18 May. Outbreak size from the exports to Uganda and from the deaths. 400 to 900 cases in their 20 May update
  • We replicated it on 19 May, in a few hours
  • Then a joint renewal model. Suspected, confirmed, deaths, exports and bed occupancy from one infection process
  • Each stream alone gives a different size. The joint estimate sits inside their spread

Posterior densities of cumulative infections, one curve per data stream, cases, confirmed, deaths and exports, and one for the joint fit, which sits inside the spread of the single-stream fits.

McCabe et al., Imperial College London, 18 and 20 May 2026, doi:10.25560/130007 · BVDOutbreakSize, first commit 19 May 2026 · figure with data to 7 June 2026

Refit on every push

  • The streams in the situation reports. Suspected and confirmed cases and deaths, laboratory, isolation, recoveries, a digitised onset curve, Uganda exports and a genomic bound
  • Updated up to daily. 145 releases since 20 May, each with the draws
  • A one-week-ahead forecast scored against the data that came later
  • No other regularly updated public model of this outbreak that I know of

BVDOutbreakSize, live · release counts from the GitHub API, 11 September 2026

How an outbreak gets modelled

  • The question first. The estimand decides which streams matter
  • Catalogue the data. What each stream measures, and how late it arrives
  • One process for infections and an observation model per stream. Delays sit inside the joint model
  • Fit, check against what was known independently, report with a date. Then again tomorrow

What keeps going wrong

  • Methods built for one outbreak are not reused for the next
  • Methods not flexible enough to take the data sources that exist
  • The data change under the model. Streams start, stop and get redefined. Daily suspects stopped on 6 August this year
  • Evaluation. Knowing which of it worked

Step lines of confirmed cases by week of symptom onset as read from situation reports of 12 July, 26 July, 9 August, 23 August and 6 September 2026, rising from near zero in May to about 400 a week in August, with the reads for the same week differing by up to a few tens of cases between reports.

BVDOutbreakSize, read 11 September 2026

Use of AI in past, present and future outbreak modelling

What even is AI?

  • Machine learning (i.e. statistics)
  • Neural networks (i.e. statistics)
  • Physics informed neural networks (i.e. statistics and mathematical modelling)
  • Universal differential equations (i.e. statistics and mathematical modelling)
  • Reinforcement learning
  • Attention based neural networks
  • Large language models
  • My friendly agentic CLI tool (i.e. claude code)

My view

Many of these approaches are under-explored and not well integrated with the rapidly evolving, noisy, sparse data settings I have worked in. I am not aware of any real-world operational use so far, and I am mostly not convinced by the evaluations of proposed methods.

Sam Abbott, September 2026

And yet

I am very excited about their potential.

Sam Abbott, September 2026

AI methods by task, after Kraemer et al. 2025

A table with columns task, Kraemer et al. 2025 and in this talk. Eight tasks from the paper with the AI method named for each.

Table 1 of Kraemer et al. 2025, Nature 638, 623, doi:10.1038/s41586-024-08564-w

Universal differential equations

  • A universal differential equation (UDE) is a differential equation with one or more terms a neural network, fitted jointly with the mechanistic parameters. Rackauckas et al. 2020
  • Applied to epidemics by Dandekar, Rackauckas and Barbastathis 2020. A network for quarantine strength in an SIR (susceptible, infectious, recovered) model, fitted to case counts from 70 countries
  • Reporting is not modelled, and there is no comparison with a simple statistical or semi-mechanistic model, as Funk et al. made for Ebola. Philipps et al. 2025 found noise and limited data degrade the fit and the mechanistic parameters lose their meaning
  • Charlotte Holt, epiforecasts PhD student, is using UDEs to try and recover behavioural change in outbreaks and improve forecasts
  • Potentially most exciting where many hard to model data sets are fitted jointly to an otherwise well understood system

Left, an SEIR box model with the transmission rate a random walk and one arrow to a cases box. Right, the same SEIR boxes with the transmission rate a neural network fed by mobility, surveys, policy, weather and wastewater, and an arrow to cases, deaths and admissions.

Rackauckas et al. 2020, arXiv:2001.04385 · Dandekar, Rackauckas and Barbastathis 2020, Patterns, doi:10.1016/j.patter.2020.100145 · Funk et al. 2019, PLOS Comput Biol, doi:10.1371/journal.pcbi.1006785 · Philipps, Schmid and Hasenauer 2025, npj Syst Biol Appl, PMC12398592

Physics-informed neural networks

  • A physics-informed neural network (PINN) maps time to the model state, \(S(t)\), \(I(t)\), \(R(t)\). The loss adds the residual of the differential equation to the misfit with the data
  • Time-varying parameters such as \(\beta_t\) come out of the same fit. Kharazmi et al. 2021 on New York City, Rhode Island, Michigan and Italy
  • The equation constrains the fit where surveillance series are short or noisy. Millevoi et al. 2024 fitted Italian COVID-19 cases and admissions jointly, and found extrapolation past the training window can go far from the data
  • Rama et al. 2026 forecast Italian influenza seasons with a PINN, competitive on the weighted interval score
  • Deterministic, so no uncertainty quantification by default

A schematic. A time input feeds a neural network. The network outputs the model state S, I and R and the transmission rate. The outputs go to two loss boxes, a data misfit against reported cases and a physics residual of the SIR equations. The two losses sum to one training loss.

Raissi, Perdikaris and Karniadakis 2019, J Comput Phys, doi:10.1016/j.jcp.2018.10.045 · Kharazmi et al. 2021, Nat Comput Sci, doi:10.1038/s43588-021-00158-0 · Millevoi, Pasetto and Ferronato 2024, PLOS Comput Biol, doi:10.1371/journal.pcbi.1012387 · Rama et al. 2026, arXiv:2506.03897

Renewal processes as a neural network layer

  • A specialised form of physics-informed neural network
  • A renewal process is an autoregressive transformation of the reproduction number \(R_t\) and past infections to today’s infections
  • Can be thought of as a neural network layer
  • Can then be stacked with other layers. Inputs to \(R_t\), an \(R_t\) layer, observation layers, and multiple data streams
  • Against a PINN on a differential equation, time-varying parameters are natural to a renewal process, and its discrete time matches the data
  • Not seen this used in research yet. Proposed in Samuel Brand’s RenewalExamples repository

Five stacked boxes joined by downward arrows. Inputs to Rt, an Rt layer, a renewal layer as a recurrent cell with the renewal equation, observation layers as a convolution over the delay, and data streams. Side notes read swap a layer, keep the rest, and add a stream, add a layer.

Samuel Brand, github.com/SamuelBrand1/RenewalExamples, conv branch · Pervez, Locatello and Gavves 2024, Mechanistic Neural Networks

AI-driven inference

  • Infectious disease models are mostly fitted by Markov chain Monte Carlo (MCMC), sequential Monte Carlo (SMC) and similar Bayesian methods
  • Compute intensive, especially for stochastic models where a particle filter is needed to estimate the likelihood
  • AI inference methods offer an alternative. Simulate parameter and data pairs from any model you can run. Train a network to map data to a posterior
  • OutbreakFlow, Radev et al. 2021. Generation time, undetected fraction and reporting delays for early COVID-19 in Germany
  • Pinotti et al. 2025 fitted an SEIR model, SIR plus an exposed class, to the 2014 Sierra Leone Ebola data by neural posterior estimation. Jang, Candan and Chowell 2026 compare it with approximate Bayesian computation

The OutbreakFlow architecture. A training phase on the left with a prior, an SIRD model and simulated time series, a stack of filtering, summary and inference networks in the middle producing an approximate posterior, and an inference phase on the right taking observed time series.

Figure 1 of Radev et al. 2021, PLOS Comput Biol, doi:10.1371/journal.pcbi.1009472, CC BY · Pinotti et al. 2025, bioRxiv, doi:10.1101/2025.11.25.690436 · Jang, Candan and Chowell 2026, PLOS Comput Biol, doi:10.1371/journal.pcbi.1014364

Foundation time series models

  • Transformers, attention-based neural networks, pretrained on many series from other domains. TimesFM is decoder-only, Chronos uses T5
  • Your series in, a forecast out, no fitting
  • Kalahasti et al. 2025. Strong short-term accuracy on sparse or irregular data. Wang, Li and Perra 2026. TabPFN-TS, zero-shot, rivalled the RespiCast influenza-like illness ensemble of the European Centre for Disease Prevention and Control
  • Jafari et al. 2026. A mixture of pretrained forecasters did best on US influenza. CAPE pretrains on simulated epidemics
  • How do these methods hold up prospectively when the data change over time?

A stack of many grey time series labelled many series, many domains, feeds a slate box labelled one pretrained model, a transformer, which produces a brick forecast fan on a teal epidemic curve to the right, labelled your series in, a forecast out, no fitting.

Kalahasti et al. 2025, medRxiv, doi:10.1101/2025.02.24.25322795 · Wang, Li and Perra 2026, medRxiv, doi:10.64898/2026.05.11.26352889 · Jafari et al. 2026, arXiv:2606.19560 · CAPE, Liu et al., KDD 2026, PMC13464487

An outbreak-specific foundation model that handles data changing over time is a project I would like to be involved in.

Sam Abbott, September 2026

Reinforcement learning for control (I don’t know anything about this)

  • Reinforcement learning. A policy, a network from the state of the epidemic to an action, is learned by trial and error against a simulator to maximise a reward
  • Optimal control needs the model equations. Reinforcement learning only needs to run the simulator. Where does it beat optimal control on the same model? An open question for me
  • Libin et al. 2020. School closure policies learned with proximal policy optimisation (PPO) on a 379-district influenza model of Great Britain
  • Reymond et al. 2024. The Pareto front of multi-objective COVID-19 mitigation policies on a Belgian stochastic compartmental model
  • Bolshov and Chumachenko 2026, a scoping review of 13 studies. Limited prospective validation or real-world deployment testing

Two boxes, a policy network on the left and a simulator on the right, joined by an action arrow labelled close schools, allocate vaccine, and a return arrow labelled state and reward, with a line above reading learn a policy by trial and error against a model, then hope the model was right.

Libin et al. 2020, ECML PKDD, arXiv:2003.13676 · Reymond et al. 2024, Expert Syst Appl, doi:10.1016/j.eswa.2024.123686 · Bolshov and Chumachenko 2026, PLOS ONE, doi:10.1371/journal.pone.0351176

Agentic AI and outbreak modelling

What I mean by an agent

A language model with tools. It reads files, writes code, runs it, and loops until a check passes

A brick box on the left, you, a prompt, feeds a loop of four boxes, read the files and the docs, write or change code, run it, check passed. No returns to the top. Yes leads to a pull request and a brick box, you, read it, merge or send back.

How I am working at the moment

  • I was travelling for summer schools, teaching and workshops, so made a very focused go of working through agents only
  • All code and most text I have produced in the last six months has been agent written, pushed from a bot account, a second GitHub login
  • Agents write the tasks, from a spec or from a conversation with me
  • I am reviewing at different levels. Sometimes all the code, sometimes only the outcome
  • My agents also work with Sebastian Funk’s, who is working in a similar way

Bar chart of seabbs-bot pull requests opened per month in 2026: 17, 92, 78, 61, 124, 531, 744, 399, and 98 for September to the 11th shown in grey. Footer: 2,144 total, 85.5% merged, 11.9% closed unmerged.

GitHub search API total_count per month, 11 September 2026; seabbs-bot created 23 January 2026

The seabbs-bot GitHub avatar, a robot wearing a party hat.

The sbfnk-bot GitHub avatar, a robot in a red scarf.

@seabbs-bot and @sbfnk-bot · sbfnk-bot pull request #656 in BVDOutbreakSize, merged 8 September 2026

Agents building methods into software

  • Two approaches on epiaware.org. Composed distributions, an event tree of delays. Composable Turing models, transmission, latent process and observation as swappable parts
  • What kind of infectious disease modelling language works well for agents?
  • epidist. A meta-analysis model for published delay estimates, from scratch in about a week
  • EpiDelays. A UHasselt package rewired onto the primarycensored likelihood in two days
  • primarycensored. The generalised gamma integral solved analytically by an agent. Three earlier ones we solved by hand

The approaches section of epiaware.org. Two cards, composed distributions and composable Turing models, each with a short description and a read the approach link, above the line more approaches may join these as the ecosystem grows.

epiaware.org/approaches · epidist #620 · EpiDelays #3 · primarycensored #334

Agents in the Bundibugyo DRC outbreak model: the data

  • The situation reports from the Institut National de Santé Publique (INSP) are French PDFs. Agents read every stream from them
  • Where a public transcription has the same series, the two are cross-checked
  • Recent reports get two blind reads
  • Symptom onsets are reported only as a figure. An agent-written digitiser recovers the counts from pixel colour and axis tick geometry
  • These are cross-checked again by several agents, and by digitisers written in both Python and Julia
  • Onsets also move between reports because reporting is late. That nowcasting problem is in the model

Page 3 of INSP situation report 87, in French. A table of alerts by province, a bar chart of confirmed cases by date of symptom onset over 21 days, and a map of cases by health zone.

INSP SitRep 087, 9 August 2026, page 3, insp.cd · BVDOutbreakSize, read 12 September 2026 · the INRB-UMIE transcription mirror

Agents in the Bundibugyo DRC outbreak model: the model

  • Model development through issues and pull requests. 333 of 425 pull requests from the bot account
  • Model structure and approach set by me, at a high level
  • Automated performance work by agents, reviewed by agents. 19 perf commits, all from bot identities. One review found nothing worth shipping. One halved the gradient cost
  • Agents managed two versions of the model in tandem. The original integral model and its renewal process successor, kept in step until the renewal model merged on 9 June 2026

The agent loop again, now fed by a box where agents write the tasks from a spec or a conversation, with review agents sending findings back into the loop, me reading all the code or only the outcome, and a merge that refits in CI and sends new tasks back to the top.

BVDOutbreakSize: pull request counts from the GitHub search API, 11 September 2026 · git log --grep perf, issue #444, pull request #656 · the renewal model, pull request #155, opened 31 May, 100 commits, merged 9 June 2026

What agents made possible

  • Issues with the McCabe et al. Imperial College report were addressed, and every stream they had used was fitted jointly
  • Going wide. One posterior over the sitrep streams, the Uganda exports and a genomic bound
  • Checks on every rebuild, following the Bayesian workflow. Predictive checks per stream, forecasts scored against later data, the McCabe replication
  • That much debugging would not have happened by hand. Agents did not choose the model. Infrastructure and review kept them in check

The six-step workflow, question, data, model, fit, check and report, with a tag over each. People over question; agents read and digitise over data; people choose, agents write over model; agents, every push over fit; agents run, people judge over check; agents draft, people sign over report.

README and docs/src/news.md of BVDOutbreakSize, read 12 September 2026 · Gelman et al. 2020, Bayesian workflow, arXiv:2011.01808 · A workflow for infectious disease modelling, Abbott et al.

Digitising the onset curve

  • SitRep 081. JPEG compression left the axis ticks one pixel tall. About 10 of 19 ticks found, and a total 54% under the printed one
  • SitRep 087. The y-axis grid moved from steps of 20 to 25. Caught only because it broke the reporting-triangle invariant on dates that should have been stable
  • SitRep 098. A losslessly encoded image read about 7% high against both neighbours and was excluded
  • Every fix must reproduce every previously committed block unchanged

The raster image embedded in situation report 87. Confirmed cases by date of symptom onset from April to August 2026 as stacked bars, alive in light blue and died in red, a dashed red line at the first positive laboratory result, a y-axis in steps of 25 and a shaded band on the right for potentially incomplete data.

The embedded image from INSP SitRep 087, 9 August 2026, page 3 · data/README.md in BVDOutbreakSize, notes for SitReps 081, 087 and 098, and issue #594

What else went wrong

  • Forecasts. Bugs keep turning up. Five defects in the one-week-ahead forecast, all in the published report, fixed in August. Four weeks ahead, \(R_t\) ran from 0.245 to 14.9
  • Automatic differentiation. Enzyme cannot differentiate the joint model, open since July. Mooncake, the backend we use, paid for type-unstable closures until a fix in September halved the gradient
  • Observation models. Agents proposed models that did not match how the response works. Bed capacity as a walk that could only rise. One length of stay for patients who died and patients discharged
  • Found by looking at plots, by profiling, and by domain knowledge

Three cards. Forecasts: a fan in log R_t that widens linearly with the horizon as shipped against one that widens with its square root as fixed, caught by looking at a plot. Automatic differentiation: bars for one Mooncake gradient of the joint model, 21 ms then 10 ms, and a note that Enzyme cannot differentiate it at all, caught by profiling. Observation models: a staircase that only rises, as written, against a line with a trend and mean-reverting deviation, as fixed, caught by domain knowledge.

docs/src/news.md of BVDOutbreakSize, v1.15.0 and v1.18.0 · issues #445, #606, #297 and #299 · pull request #656

Agents that search over models

  • For the Bundibugyo model this was ad hoc. Agents proposed patches to the model as issues, one branch each. Every release scored the forecast against later data, and I reviewed the patches against domain knowledge
  • Where there is a known target, the systematic version is a tree search, or an inference step over models. An MCMC proposal in model space
  • An agent proposes candidate models, one loop each. Fit and score, keep the best, branch again
  • Open questions. How to enforce domain knowledge, and how to pick the evaluation target

A tree. The current model at the root branches into three candidates, add a stream, change the delay prior, and split Rt by province. Two are kept and branch again into three each; one child of each is kept in teal and the rest are grey, and the kept branches meet at a box reading keep the best, branch again.

Google’s tree search topped the CDC forecast hubs

  • Martinson et al. 2026. A large language model (LLM) guided tree search that writes forecasting code. An ensemble went to the CDC FluSight, COVID-19 and respiratory syncytial virus (RSV) hubs
  • Most candidates adapt or recombine models already in the hubs. 142 prompts in three families
  • Prospective, 2025/26 season. Matched or beat the CDC ensembles, and ranked first of 43 eligible FluSight, 12 COVID-19 and 4 RSV submissions
  • Bracher and Funk 2026. The 11% retrospective gain claimed earlier came from data revisions leaking into the evaluation
  • 207,500 candidates scored. My guess is the gap in compute and people against other teams is huge
  • The hubs, the data, the target and the review were still built by people

The arXiv abstract page for Prospective multi-pathogen disease forecasting using autonomous LLM-guided tree search, by Martinson, Brenner, Plomecka, Williams, Reich and Shamsi, submitted 15 May 2026.

Martinson et al. 2026, arXiv:2605.16238 · Bracher and Funk 2026, arXiv:2608.05883 · team folders Google_SAI-* in the three CDC hubs on GitHub

Who gets to do this

“The people modelling outbreaks where outbreaks happen are the least likely to have paid access.”

Sam Abbott, how I LLM, July 2026

  • A Claude Max subscription. £200 a month, from my own pocket
  • Who can do this work at all is narrowing. Agents make that worse
  • Real work is now possible on small open models
  • What happens when the venture capital money dries up?

Sam Abbott, how I LLM, July 2026

Thank you