Research exchange

VUB AI group and UH-UA SIMID group

Sam Abbott

London School of Hygiene & Tropical Medicine

15 September 2026

Talk plan

  • Who I am, what I work on, where the tools get used, and who I work with
  • Composable infectious disease modelling. The gap, what we want from any approach, two approaches, one replication
  • Where AI might plug in. A network in the slots of a composable model, physics-informed networks, renewal as a layer, a foundation model
  • Agentic modelling workflows. The workflow, agents building methods into software, search over models
  • Recent delays work. primarycensored, epidist, CensoredDistributions.jl

samabbott.co.uk/AIandIDVUB/exchange

Research exchange, AI Experience Centre, Pleinlaan 9, Brussels, Tuesday 15 September 2026, 09:45.

Who I am

  • London School of Hygiene & Tropical Medicine. Centre for the Mathematical Modelling of Infectious Diseases (CMMID)
  • A small group. One small internal grant held as principal investigator
  • Many collaborators. epiforecasts, CMMID, the epinowcast community, EpiAware, the US Centers for Disease Control and Prevention (CDC) and the UK Health Security Agency (UKHSA)
  • Tools in wide use. The use has not turned into funding

Sam Abbott

Package Since Stars Downloads
EpiNow2 2020 141 58,073
scoringutils 2020 64 63,648
epinowcast 2021 67 not on the archive
epidist 2023 16 not on the archive
primarycensored 2024 9 11,294
baselinenowcast 2025 10 1,678

Stars from the GitHub API; downloads are totals from the Comprehensive R Archive Network, August 2026

What I work on

  • Composable models. Components that can be reused across contexts and combined into joint models, with the uncertainty carried through
  • Improving outbreak modelling. Methods that survive from one outbreak to the next. Models that fit the data sources that exist
  • A workflow for modelling with multiple data sources. Data sources added one at a time, with the checks and feedback loops written down
  • Joint models give better evidence than chains of separate models. Building one during an outbreak is slow and needs several kinds of expertise
  • In R and Stan first. In Julia now

Four columns of building blocks. Process models: three renewal variants. Observation models: eight ways to link infections to counts, with delays and ascertainment. Inference and computation choices, from HMC to synthetic likelihood. Data integration choices, from ensembling to full joint modelling. The four snap together at the bottom right.

Figure 6 of Abbott et al., A workflow for infectious disease modelling

Where the tools get used

  • Consulting. Time with CDC collaborators, UKHSA, and state and local public health departments in the US
  • The aim is to add other countries’ public health departments
  • In outbreaks. epidist in the DRC response’s correspondence in the New England Journal of Medicine this August
  • Several teams used various tools in Exercise Pegasus, the UK’s 2025 national pandemic exercise
  • Next. Tools flexible and extensible enough to be easy for a large language model to use

A table of seven rows. US CDC's Center for Forecasting and Outbreak Analytics, EpiNow2 for nowcasts and R_t and EpiAware.jl built together, 2024 to 2026. UKHSA, mpox nowcasting, norovirus nowcasts and Exercise Pegasus, 2022 to 2025. DRC, INSP and INRB in Kinshasa, epidist delay estimates in the NEJM correspondence, August 2026. ECDC, the European COVID-19 Forecast Hub now RespiCast, 2021 on. Ottawa Public Health, EpiNow2 projections. The Philippines FASSSTER platform, EpiNow2 for regional R_t. South Africa's NICD and SACEMA, EpiNow2 on clinical and wastewater data, 2025.

CDC CFA, Behind the model: nowcasting, 2026 · Akilimali et al. 2026, NEJM, doi:10.1056/NEJMc2608070

epiforecasts and the epinowcast community

  • epiforecasts. Led by Sebastian Funk at LSHTM, and I am a member. Real-time modelling and evaluating it
  • EpiNow2, scoringutils, and the European COVID-19 Forecast Hub from 2021, now RespiCast at the European Centre for Disease Prevention and Control (ECDC)
  • 2020 to 2022, weekly estimates and forecasts for the UK government’s modelling advisers
  • epinowcast. A community I created. A forum of over 50 researchers and public health analysts
  • A monthly seminar since May 2023, 31 so far, organised with Clara Brigitta and Fanny Bergström
  • In 2023 I wrote: “community >> methods or models alone”

community.epinowcast.org · Abbott, IISA 2023 post, 4 January 2023

Who I work with

Real-time tools, delays and nowcasting

Six collaborators on real-time tools and hubs: Sebastian Funk, Samuel Brand, James Azam, Nikos Bosse, Kath Sherratt and Katelyn Gostic.

  • Five of these six are EpiNow2 co-authors. Nikos Bosse wrote scoringutils, the forecast scoring package
  • Sebastian Funk, Samuel Brand and Kath Sherratt are co-authors on the live DRC model. Katelyn Gostic is at CDC

Eight collaborators on delays and nowcasting: Adrian Lison, Kaitlyn Johnson, Adam Howes, Carl Pearson, Kelly Charniga, Sang Woo Park, Thomas Ward and Christopher Overton, the last four shown as initials tiles.

  • Adrian Lison and Adam Howes are second only to me on epinowcast and epidist. Kaitlyn Johnson built baselinenowcast
  • primarycensored is joint work with Thomas Ward, UKHSA, and Christopher Overton

Package author lists · photos from GitHub

Composable infectious disease modelling

Three outbreaks, the same problems

  • COVID-19, 2020. \(R_t\), the number of new infections per infected person, daily for thousands of locations
  • mpox, 2022. Nowcasting for UKHSA. Nowcasting is estimating what has already happened but has not yet been reported
  • Ebola disease caused by Bundibugyo virus, Democratic Republic of the Congo (DRC), 2026. One joint model across five data streams, refit on every push
  • Each time the methods from the last outbreak were not reusable
  • Each time the models did not fit the data sources that existed, and the data changed under us
A timeline from 1760 to 2026 in two panels with an axis break. Grey dots for Bernoulli 1760, Ross 1911, Kermack and McKendrick 1927, Anderson and May 1991, foot and mouth 2001, H1N1 2009 and Ebola in West Africa 2014. Teal dots, marked work I was part of, for COVID-19 in 2020, mpox in 2022 and Ebola disease caused by Bundibugyo virus in 2026.

Abbott et al. 2020, Wellcome Open Research · Overton et al. 2023, PLoS Comput Biol

Analysing data and processes separately, or all together

  • The Office for National Statistics (ONS) COVID-19 Infection Survey. 4M+ swabs, 150K households, £500M+, April 2020 to March 2023
  • External modellers accessed only summarised prevalence, not underlying observations
  • Lison et al. found biases at each step. Joint modelling of all steps together avoided these cascading biases
  • But models are difficult to develop within policy-relevant timelines

Lison et al. (2024), doi:10.1371/journal.pcbi.1012021

Four panels of the survey chain fitted as one model: prevalence, incidence, antibody prevalence and R_t, each with a posterior ribbon over the survey period.

Abbott and Funk (2022), doi:10.1101/2022.03.29.22273101

What we want from any approach

  • Uncertainty and inference. Joint and staged inference that avoids the cascading biases, and automatic differentiation for fitting
  • Model structure. Components separated by process type, and models nested inside models
  • Workflow. One specification that is useful for simulation and for inference

Warning

These design considerations were developed by the authors and have not yet received broader community input.

Table 1 of the paper. Twelve proposed requirements under six themes, each paired with the problem that motivates it, from probabilistic specification with full uncertainty propagation to programmatic generation of model components.

Table 1 of Abbott et al., Composable probabilistic models can lower barriers to rigorous infectious disease modelling · epiaware.org/approaches

What is composable modelling?

Components can be reused across contexts and combined in different configurations whilst maintaining statistical rigour.

Abbott et al., Composable probabilistic models can lower barriers to rigorous infectious disease modelling

Four applications drawn as graphs over a band of shared parts. The incubation period model appears in all four and the latent infection model in three.

Figure 1 of Abbott et al., Composable probabilistic models can lower barriers to rigorous infectious disease modelling, in CDC clearance

The Julia composable work

  • EpiAware. Loosely an extension of the epinowcast community, and in its initial stages
  • Top row. EpiAware, and the composable paper. Samuel Brand, Damon Bayer and Joseph Lemaitre built CensoredDistributions.jl with me. Hong Ge leads Turing.jl
  • Bottom row. JuliaBayes, formed April 2026, building the tools around Turing.jl. Chains, regression models, a Bayesian workflow package

Photos from GitHub

Twelve collaborators on the Julia work in two rows. Top: Samuel Brand, Damon Bayer, Michael DeWitt, Joseph Lemaitre, Hong Ge and Sandra Montes-Olivas. Bottom: Penelope Yong, Peter Thestrup Waade, Guillaume Dalle, Ryan Senne, Simon Steiger and Jessica Cox. Three shown as initials tiles.

ComposableTuringIDModels.jl

using ComposableTuringIDModels, Distributions, Turing

# Renewal infections with an autoregressive log-R_t prior
renewal = Renewal(;
    generation_time = Gamma(1.4, 1 / 0.38),
    rt = AR(; ϵ_t = HierarchicalNormal(std = HalfNormal(0.1))),
    initialisation = Normal(log(1.0), 1.0))

# Observation layers: delay → ascertainment → negative-binomial noise
obs = LatentDelay(
    Ascertainment(
        NegativeBinomialError(),
        Intercept(Normal(log(0.4), 0.1))),
    truncated(Normal(5.0, 2.0), 0.0, Inf))

model = IDModel(renewal, obs)

# Fit to case data with Turing's NUTS
chain = sample(as_turing_model(model, cases, length(cases)),
    NUTS(0.95), MCMCThreads(), 100, 2)
  • A library of parts over Turing.jl
  • Three kinds of part. An infection model, an observation model, and prior models nested in the slots of both
  • Parts carry their own priors
  • A model definition does not know about the data. One definition simulates and fits
  • Adding a new application means defining only the new parts

ComposedDistributions.jl

using ComposedDistributions, Distributions

# BDBV-inspired delay tree: parallel pathways
# with sequential, resolve, and compete nested together
tree = @uncertain compose((
    clinical = sequential(
        :onset_admit => Gamma(Normal(1.2, 0.2), 3.0),
        :admit_resolve => resolve(
            :death => (Gamma(2.0, 3.5), 0.3),
            :discharge => (Gamma(1.0, 8.0), 0.7))),
    surveillance = sequential(
        :onset_notif => Gamma(0.7, 20.0),
        :notif_compete => compete(
            :confirmed => Gamma(2.0, 1.0),
            :lost => Gamma(5.0, 0.5))),
))

# Draw a synthetic case, or evaluate a real one's likelihood
case = rand(tree)
logpdf(tree, case)
  • Write the epidemiology as Distributions.jl distributions. The probabilistic programming language becomes optional
  • A case’s event history is a tree. sequential chains delays, resolve picks an outcome with a probability, compete races two
  • The whole tree is itself a distribution. It slots in wherever a delay is needed inside a larger model
  • Parameter uncertainty is an ordinary leaf

Real-time nowcasting (EpiNow2)

  • Reuses the autoregressive prior from the paper’s first case study, swapping the Poisson observation model for a negative binomial one
  • broadcast_weekly wraps it to give piecewise constant weekly \(R_t\) values
  • ascertainment_dayofweek adds day-of-week reporting effects
  • Two LatentDelay wrappers compose the incubation and reporting delays by nesting structs

Important

No new components needed and took ~ 3 hours

One configuration of EpiNow2 replicated, not a benchmark. Abbott et al. (2020), doi:10.12688/wellcomeopenres.16006.2

Six panels. Prior checks on each part: time varying R_t draws, latent infection draws, observed cases against latent infections, and the incubation period and reporting delay distributions. Then the joint fit to daily COVID-19 cases from Italy, February to June 2020, and the posterior for time varying R_t.

Where AI might plug in

A neural network in the slots of a composable model

  • Today the slot holds a Gaussian process, or an autoregressive (AR) process. It is free to move and one data stream constrains it
  • A network in the same slot. This is what a universal differential equation does. It learns what the Gaussian process cannot say. Behaviour, season, policy
  • A network on the observation side. Several data streams fitted from the same latent infections
  • Most exciting where there are many hard to model data sets to fit jointly in an otherwise well understood system
  • How does this work with Bayesian modelling? Priors over weights, identifiability against a slot that already flexes, and what the posterior then means
  • No neural network code in any EpiAware package yet

Three boxes in a row: an infection model, Renewal with an empty rt slot, an observation model of LatentDelay, Ascertainment and NegativeBinomialError, and cases. Two candidates point at the R_t slot: an AR process, labelled today, and a small neural network of covariates, labelled a question, with a swap arrow between them. Under the observation model a third candidate, a network from the latent infections to cases, deaths and wastewater at once.

A physics-informed network as the mechanism arrives

  • A physics-informed neural network (PINN) adds the residual of the differential equation to the data misfit in its loss, so the mechanism enters as a penalty rather than as structure
  • Early in an outbreak we know almost nothing. Start with the network carrying most of the fit
  • Then dial the mechanism up. Add equations as they become defensible, inside the same fit, rather than switching model
  • Will it hold when the data are very sparse? That is the question for the first weeks, which is when we would want it
  • Deterministic, so no uncertainty quantification by default

A schematic. A time input feeds a neural network. The network outputs the model state S, I and R and the transmission rate. The state goes to a data loss, the misfit to reported cases; the state and the rate go to a physics loss, the residual of the differential equations, with automatic differentiation supplying the derivatives. The two losses sum to one training loss.

Raissi, Perdikaris and Karniadakis 2019, J Comput Phys, doi:10.1016/j.jcp.2018.10.045 · Millevoi, Pasetto and Ferronato 2024, PLOS Comput Biol, doi:10.1371/journal.pcbi.1012387

Renewal processes as a neural network layer

  • A renewal process is an autoregressive transformation of \(R_t\) and past infections to today’s infections. It can be thought of as a neural network layer
  • The layers compose. Inputs to \(R_t\), an \(R_t\) layer, delay and ascertainment layers, a likelihood on each stream. The same parts we already assemble by hand
  • The network can stay small, because the epidemiology is in the layers rather than learned. That is what makes it plausible on the data an outbreak actually gives you
  • Then build up. More streams and more structure by stacking, not by writing a new model
  • Not seen this used in research yet. Proposed in Samuel Brand’s RenewalExamples repository

Samuel Brand, github.com/SamuelBrand1/RenewalExamples · Pervez, Locatello and Gavves 2024, Mechanistic Neural Networks

On the left the renewal equation, I_t equals R_t times the sum over s of I_{t minus s} g_s, described as a convolution with a fixed kernel scaled by R_t. On the right a stack of six layers: inputs, an R_t layer marked learned, then a renewal layer, a delay layer and an ascertainment layer marked fixed by the epidemiology, then a negative binomial likelihood on each stream.

An outbreak-specific foundation model

  • Outbreak data change after the fact. On the DRC model the daily suspected stream stopped on 6 August. In May INSP, the DRC’s public health institute, reclassified suspected cases
  • So pretrain on outbreaks as they were seen at the time, not on the final curves. CAPE pretrains on simulated epidemics, but on the curves
  • We can simulate what it would need. Composable models generate early and mid-outbreak data, including complex surveillance streams, which could be the training set
  • How do these methods hold up when the data change under them, in prospective settings?

Kalahasti et al. 2025, medRxiv, doi:10.1101/2025.02.24.25322795 · Liu et al. 2026, CAPE, KDD

Left, the same outbreak curve as it looked in May, July and September, each vintage lower and shorter than the next, with a dashed line where the case definition changed, and below it bars showing data streams that start and stop. Middle, an arrow labelled every vintage, every stream, with its date, into a box labelled an outbreak foundation model pretrained on outbreaks as they were seen at the time. Right, a nowcast and forecast fan on the latest vintage.

“An outbreak-specific foundation model that handles data changing over time is a project I would like to be involved in.” Sam Abbott, September 2026

Agentic modelling workflows

A workflow for infectious disease modelling

  • The workflow. Nine stages, from the research question to data integration, with feedback loops running back to earlier ones
  • On BVDOutbreakSize agents wrote the code. The model structure and approach were specified by me at a high level. The judgements stayed with people
  • Agents do best with something to check against. A simulation with known parameters, an analytic case, another group’s estimate
  • 333 of 425 pull requests, the unit of change on GitHub, came from my bot account

The nine workflow stages as a column of numbered boxes. To the left, in brick, what stayed with people: choosing the question at stage 1, which model at stage 2, what the genetic lower bound means at stage 4, whether ascertainment holds at stage 8, which streams to trust at stage 9. To the right, in teal, what agents did: reading French PDF sitreps at stage 3, digitising the onset curve at stage 4, fitting pipeline, releases and docs at stage 7, tests and recovery checks at stage 8, sensitivity analyses and figures at stage 9.

Agents building methods into software

  • epidist. A meta-analysis model for published delay estimates, built from scratch in about a week. Onset to death for Bundibugyo virus disease, beside the study the estimates came from
  • EpiDelays. A UHasselt package rewired onto the primarycensored likelihood in two days
  • What kind of infectious disease modelling language works well for agents?

Agents building and checking the parts

  • Replication. The McCabe et al. Imperial report replicated in a few hours, the day after it appeared. Then a joint model across five situation report streams, the Uganda exports and a genomic bound
  • Data. Agents read the French PDF situation reports, two blind reads each. Onsets are only a figure, so an agent-written digitiser, cross-checked by a second one in another language
  • Checks on every rebuild. Predictive checks per stream, forecasts scored against later data, the replication. That much debugging would not have happened by hand
  • Infrastructure and review were still needed to keep the agents in check
  • Could agents build and check the simulators you learn against?

BVDOutbreakSize README.md · McCabe et al. report of 18 May 2026

Posterior densities for the cumulative size of the outbreak, one per data stream and one joint, with the confirmed-case stream far to the right of the others.

June 2026 fit; the live report is at epiforecasts.io/BVDOutbreakSize/stable

Agents that search over models

  • Could a language model be the proposal step? Inside the inference algorithm rather than wrapped around it, proposing models rather than parameters
  • Composability gives it a move set. If models are assembled from components, a proposal is a swap of one part, not a new program
  • Open questions. How to enforce domain knowledge, and how to pick the evaluation target

A tree. The root is a task, forecast admissions, scored on held-out weeks. Four proposed models below it: ARIMA and a GAM in grey, dropped, and two renewal models in teal, kept. Under the kept ones four more variants, adding a reporting delay, a spline, day of week effects and a wastewater stream, with the spline dropped. Arrows from the survivors lead to a red box, an ensemble of the survivors is what gets submitted. Side labels read propose, write the code, fit, score, expand the best, many rounds.

Recent delays work

An epidemiological delay

  • The time between two events in one case. Infection to symptom onset is the incubation period
  • We want the whole distribution over days across a population, not a mean
  • Nowcasts, forecasts and transmission models take one as an input and do not re-estimate it

A time axis with four marked events, infection, symptom onset, hospitalisation and death. Three double-headed arrows above the axis span pairs of them: the incubation period from infection to onset, onset to hospitalisation, and onset to death. The last two share a start, so the intervals overlap.

Unfortunately, biases

Both events are recorded to the day

  • The primary event falls somewhere in a window \([0, w_P]\)
  • Mix the delay over where in the window it sat, a convolution on the cumulative distribution function (CDF)
  • The secondary event is recorded to the day as well
  • The probability of a daily bin is a difference of the primary-censored CDF

A timeline. The primary event sits in a blue window of width w_P, the secondary event in a red window of width w_S, the delay T = S minus P spans them, and a dashed observation cutoff C sits to the right with longer delays not yet seen beyond it.

\[ F_{\mathrm{PEC}}(t) = \int_{0}^{w_P} g(\tau)\, F(t - \tau)\, d\tau \]

\[ \Pr\!\big(T \in [n, n + w_S]\big) = F_{\mathrm{PEC}}(n + w_S) - F_{\mathrm{PEC}}(n) \]

Maths after primarycensored

Note

In the past this was often written as a double integral, which is typically more complicated to evaluate.

Short delays are over-represented early in an epidemic

  • A pair enters the data only if the secondary event happened before the cutoff \(C\)
  • Condition on being observed. Divide by the CDF at the horizon
  • Move the cutoff and watch the mean of what you can see. Continuous LogNormal(1.6, 0.6) delays, every primary event at day zero

\[ L^{\mathrm{RT}} = \frac{F_{\mathrm{PEC}}(n + w_S) - F_{\mathrm{PEC}}(n)} {F_{\mathrm{PEC}}(C - P)} \]

A mean of 3.9 days instead of 5.9

  • At a cutoff seven days after the primary events, and 5.3 days at fourteen

LogNormal(1.6, 0.6) with daily windows, computed with CensoredDistributions.jl v0.2.22

Two panels. On the left, the delay distribution seen by day 7 and by day 14 sits well to the left of the true distribution. On the right, the mean estimated delay rises from under 2 days towards the true mean of 5.9 days as the cutoff moves away from now, passing 3.9 days at a cutoff of 7 and 5.3 days at a cutoff of 14.

Three packages, one likelihood

using CensoredDistributions, Distributions, Turing

@model function delay_model(values, weights)
    α ~ truncated(Normal(1, 2), 0, Inf)
    θ ~ truncated(Normal(1, 2), 0, Inf)
    d = double_interval_censored(
        Gamma(α, θ); upper = 15, interval = 1
    )
    values ~ weight(d, weights)
end

chain = sample(delay_model(values, weights),
               NUTS(), MCMCThreads(), 1000, 2)

CensoredDistributions.jl README.md, v0.2.22

  • primarycensored, R. d, p, q and r prefixes, as in base R
  • epidist, R. The same likelihood through brms. Delays vary by district, by age and by wave
  • CensoredDistributions.jl. Each adjustment is a wrapper on a Distributions.jl object. The stack is still a distribution, so Turing takes it as it is

I do not think AI plugs in here, bar the algebra

  • No crossover that I can see between this work and AI methods. It has taken a lot of my time
  • The exception is the algebra. Language models have become good at solving integrals analytically
  • The generalised gamma. Solved and added to primarycensored in R and Stan by an agent, merged 10 September 2026. The three before it we did by hand
  • Every closed form is one fewer numerical integral in every likelihood evaluation

Thank you