Science communication with and under AI

Sam Abbott

London School of Hygiene & Tropical Medicine

14 September 2026

Talk plan

  • 2020, three audiences. Government, the public, other governments, the forecasting hubs and how they score
  • Software is communication. What a package carries, defaults, the epinowcast community
  • Communication in the age of AI. Limitations first, a bot account, a prompt log, and a reviewer that is not a person
  • Under AI. Bots talking to bots, what a language model says about us, and what the person is for

samabbott.co.uk/AIandIDVUB/communication

Combining Infectious Disease Modelling and AI, VUB, Brussels, Monday 14 September 2026, 13:40.

2020, three audiences

Government. Weekly estimates in, a consensus statement out

  • \(R_t\), the reproduction number over time. How many people each case goes on to infect
  • Weekly \(R_t\) estimates and reports to SPI-M-O, the Scientific Pandemic Influenza Group on Modelling, Operational sub-group
  • SPI-M-O reported to SAGE, the Scientific Advisory Group for Emergencies
  • The public saw a consensus statement

We don’t know whether any of this gave any policy maker any useful information that helped them make better decisions…

Abbott and Funk, 2022

What the consensus process did

  • Many groups. Each used its own data sources and its own model
  • Evidence from several models was considered, discussed and combined
  • A range went out. The detail behind each estimate did not

If results are presented individually this can get lost and it incentivises being first out (and maybe therefore cutting corners).

Abbott, 2023

The public. A map, then a page for each place

  • epiforecasts.io/covid, April 2020 to March 2022
  • Daily \(R_t\) and nowcasts for several thousand locations. A nowcast estimates what has happened but has not yet been reported
  • The front page was a map. Five words, increasing to decreasing, and no numbers
  • Just over 500,000 unique users and 1.2 million page views by March 2022

The global page of the dashboard. A world map with countries coloured increasing, likely increasing, stable, likely decreasing or decreasing, with the data date and download links above it

epiforecasts.io/covid, global summary, frozen 30 March 2022, screenshot 12 September 2026 · counts from Abbott and Funk 2022

The caveats sat in the captions

  • Orange for estimates from partial data, adjusted for right truncation, the recent cases not yet reported
  • Each figure caption carried the same sentence. “These should be considered indicative only”
  • The limitations sat on a methods page that 20,000 of 1.2 million page views reached

UK page of the dashboard. Cases by report date, cases by infection date and the reproduction number, with the most recent days in orange

epiforecasts.io/covid, United Kingdom, frozen 26 March 2022 · page views from Abbott and Funk 2022

Other governments used it too

  • CSV files on GitHub, comma separated values, used directly by other governments or by scientists who set them beside other evidence
  • The estimates travelled. The limits travelled less well
  • Checking several thousand locations a day was beyond us
  • Could AI do that checking?

Of course this can cause issues if the limitations of the method are poorly communicated.

Abbott and Funk, 2022

The per-country table on the global page. A CSV button, then rows for Afghanistan, Albania and Algeria with new cases, an expected change in words, the reproduction number, growth rate and doubling time, each with a 90% interval

epiforecasts.io/covid, global summary, 214 countries and territories, 30 March 2022 · epiforecasts/covid-rt-estimates

What went wrong

Because we did not have the capacity to manually inspect the data and estimates on a daily basis we sometimes published nonsensical estimates.

Abbott and Funk, 2022

  • Wisconsin reported on a Sunday for the first time. The model read the catch-up as an outbreak
  • What I wrote down afterwards. Be as open as possible. State the limitations first. No point estimates
  • I left the press to people whose speciality it is. No training, no time

Collaborative forecasting hubs

  • Each team submits weekly. The hub combines them into an ensemble
  • The European hub ensemble beat 83% of teams’ case forecasts and 91% of death forecasts on relative WIS, the weighted interval score

However, they can be very hard to learn from, and though we have really tried to learn more about how to forecast it has been difficult.

Abbott, 2022

Agreeing how to score is communication too

  • Why. A forecast can only be read against the outcome and a score. Agreeing the score is agreeing what a good forecast is
  • A hub is an agreement. Same targets, same dates, same quantiles, same score. WIS is the hubs’ headline measure
  • Shared practice. scoringutils implements the scores. Pairwise comparison gives a relative skill that teams who miss weeks can still read
  • UKHSA, the UK Health Security Agency, asked us for guidance on evaluation. Practice still differs between agencies

Three models compared in pairs to give a relative skill score

Figure from the scoringutils manuscript, Bosse et al.

Where it mattered

Probably our most useful contributions from this work were in the UK where we knew the data and its limitations particularly well and interacted directly and frequently with policy makers.

Abbott and Funk, 2022

Software is communication

A package carries our assumptions to every analyst who runs it

  • Who runs ours. The US Centers for Disease Control and Prevention, UKHSA, the Robert Koch Institute and the World Health Organization
  • Users run the defaults, so the defaults are the advice
  • The rest is support. We meet analysts from health agencies and answer their questions
  • In some cases the package is too complex for anyone else to use
Package Since Stars CRAN downloads
EpiNow2 2020 141 58,073
epinowcast 2021 67 not on CRAN
epidist 2023 16 not on CRAN
primarycensored 2024 9 11,294
baselinenowcast 2025 10 1,678

Table from the JuliaCon 2026 roadmap deck, August 2026 · stars from the GitHub API, downloads from cranlogs, CRAN being the Comprehensive R Archive Network, to mid August 2026

Defaults are advice

  • EpiNow2 has been on CRAN since 1 September 2020
  • Every argument a user leaves alone is a decision they inherited from us
  • The generation time and the delays. The prior on \(R_t\). The sampler settings
  • The code on the right is the vignette. It is what gets copied
estimates <- epinow(
  data = reported_cases,
  generation_time = gt_opts(example_generation_time),
  delays = delay_opts(example_incubation_period + reporting_delay),
  rt = rt_opts(prior = LogNormal(mean = 2, sd = 0.2)),
  stan = stan_opts(cores = 4)
)

EpiNow2 vignette · CRAN archive, 1.1.0 on 1 September 2020

epinowcast. Communicating through a community

  • epinowcast. Researchers and public health analysts. 31 seminars from May 2023 to May 2026
  • A forum, a monthly meeting with a rotating chair, release notes that name everyone who reported a bug or joined a discussion
  • In 2023 I wrote: “I am also interested in why efforts in this direction keep failing”

Trying to build a community of practice as community >> methods or models alone.

Abbott, 2023

Communication in the age of AI

Limitations first. The Bundibugyo report, 2026

  • The live report on the Bundibugyo virus disease outbreak puts its limitations before the methods, in bullets, most consequential first
  • Drafted with agents. The report, and the limitations list with it
  • The feedback is that a list of limitations first reads as off-putting
  • How should we warn?

Every estimate is a model-based extrapolation under strong assumptions, not a measurement.

Abbott, Sherratt, Brand and Funk, 2026

The Limitations section of the live BVDOutbreakSize report. A heading, one sentence on how they are grouped, an expand toggle, and the first bullets under Data and what it can support

BVDOutbreakSize analysis page, v1.18.0, screenshot 12 September 2026

Saying where the agents were

  • A bot account. seabbs-bot, from this morning. Agent work goes out from it
  • A disclosure line in the Bundibugyo repository’s README and on the limitations slide of the WHO call about it
  • A prompts page with the brief and the steers
  • A bio drafted by the bot. My JuliaCon bio came from seabbs-bot, “who has a very suspiciously high opinion of me”

A published prompt log

  • The JuliaCon 2026 site. One brief, then 130 steers by 14 August 2026
  • Mid-build the prompts page counted 53. The number more than doubled before the talks
  • This site has one too. The brief, the steers and the notes these decks were rewritten from

These slides were written by an agent

What it gives me

  • Fast, and the result looks the part
  • I am dyslexic. Drafting is the part I am glad to hand over
  • More of this got built than I would have built by hand

What it costs

  • A layer between what I want to say and what gets said
  • The words arrive already formed, so I have to push to get mine back
  • Every steer on the prompts page is me pushing

I see the tells in a lot of talks now, and I catch myself wondering how much is the speaker and how much is the model.

Checking what it wrote

  • 114 claims pulled from one deck by a second agent, each checked against git and GitHub’s application programming interface (API)
  • 67 needed changing
  • “530 pull requests” was GitHub’s latest issue number. The count was
    1. A pull request is a change proposed for review
  • 712 of 1,371 merged bot pull requests carried no review, so “reviewed by me” became “about half”

A single bar of 114 claims, 67 marked changed and 47 marked stood

seabbs/how-I-llm commits 8c7e834 and 1fceb9a, August 2026

seabbs-review-bot. A reviewer that is not a person

  • A GitHub App. It reviews each pull request I or my bot open, and again when either of us asks
  • 213 pull requests reviewed since 17 August 2026. Each review ends “Not a human review”
  • I fix what it gets right, say which findings I reject and why, and check each against the source myself
  • It has caught real bugs. It has also been wrong from its own truncated greps

The GitHub App page for seabbs-review-bot. A robot avatar, the name, and the line Automated reviews of @seabbs and @seabbs-bot work

github.com/apps/seabbs-review-bot, created 17 August 2026 · count from the GitHub search API, 12 September 2026

Under AI

Who is accountable for a number an agent produced?

  • The live Bundibugyo virus disease model. Four named authors. Most of its 497 commits are from bot identities
  • One answer to how we warn. Say who drafted it and who answers for it. The report does, at the top of its README

The model code and analysis were drafted by a language model, then reviewed and revised under human oversight. The named authors are responsible for that oversight.

Abbott, Sherratt, Brand and Funk, 2026

$ git log --format='%an' | sort | uniq -c | sort -rn
 238 Sam Abbott (bot)
 144 seabbs-bot
  50 dependabot[bot]
  40 Sam Abbott
  11 Sam
   7 Sebastian Funk - robot edition
   3 Sebastian Funk
   3 Claude
   1 Samuel Brand

epiforecasts/BVDOutbreakSize, local clone, 11 September 2026 · the disclosure from the report’s README

Bots talking to bots

  • Sebastian Funk’s bot opened the issue on 20 July with a diagnosis, and came back the next day with a fitted prototype
  • My bot closed it on 4 August with the commit that fixed it
  • No person wrote in the thread. The issue before it went the same way
  • sbfnk-bot has opened 11 issues on EpiNow2 this year. One was a Jacobian error in the \(R_t\) prior. It was right

The end of BVDOutbreakSize issue 443. sbfnk-bot's prototype results, then a comment from seabbs-bot naming the commit that moved the capacity constraint onto the admissions likelihood, then seabbs-bot closed this as completed

BVDOutbreakSize#443 and #438 · EpiNow2#1475, July 2026 · gh search issues --author sbfnk-bot, 12 September 2026

Issues nobody asked for

  • 91 issues opened by seabbs-bot on repositories outside my own organisations
  • Upstream bug reports land well. Mooncake, DocumenterVitepress and AirspeedVelocity each closed theirs
  • Review requests on colleagues’ repositories land less well. This one got “Done”, from me, two days later
  • I do not like sending them. I do not love receiving them either

GitHub issue 30 on nfidd/sismid-nowcasting. seabbs-bot opened it on 1 July, QA pass across all nowcasting sessions for consistency, correctness, and flow. seabbs replied Done on 3 July and closed it

nfidd/sismid-nowcasting#30 · Mooncake.jl#1241, DocumenterVitepress.jl#375, AirspeedVelocity.jl#158 · gh search issues --author seabbs-bot, 12 September 2026

Community, in the age of robots

  • EpiAware, our Julia packages. This slide is from the JuliaCon 2026 talk about them
  • Nobody from outside the team has committed since February 2026, so the other party on an issue is usually one of my robots
  • Answering a robot is dispiriting. Do 🤖 shut down collaboration?
  • 🌍 We want contributors who are not Julia developers, and not in well resourced settings

🤖 🤖 🤖 🤖 🤖 🤖
🤖 🤖 🤖 🤖 🤖 🤖
🤖 🤖 🧑‍💻 👩‍💻 🤖 🤖
🤖 🤖 🤖 🧑‍🔬 🤖 🤖
🤖 🤖 🤖 🤖 🤖 🤖
🤖 🤖 🤖 🤖 🤖 🤖

Slide from the JuliaCon 2026 roadmap deck, August 2026 · most recent outside commit 3 February 2026 from git log on EpiAware main, checked 12 August 2026

What a language model says about EpiNow2

  • Two large language models, Claude Opus and Claude Sonnet, with no tools. “What is EpiNow2 and who maintains it?”
  • Both said I maintain it. I created and developed it
  • Opus named three contributors who are not in the author list, and a function removed in 1.9.0
  • Sonnet claimed less and got less wrong

Opus.Sam Abbott (LSHTM) is the original lead author and long-standing CRAN maintainer”

Opus. “Substantial contributions over time from Katharine Sherratt, Sophie Meakin, Nikos Bosse, James Azam, Sam Brand, Hugo Gruson, Pratik Gupte, Adam Howes …” not in the author list

Opus. “delays are specified through … Gamma(), LogNormal() and fix_dist()renamed in 1.6.0, removed in 1.9.0

Sonnet. “It’s part of the epiforecasts ecosystem of tools and is maintained primarily by Sam Abbott”

claude-opus-5 and claude-sonnet-5, claude CLI, no tools, 12 September 2026 · full answers in notes/llm-answers.md

What a language model says about me

  • “Who is Sam Abbott, the infectious disease modeller?”
  • Right. The PhD, the supervisors, my ORCID researcher identifier and the packages
  • Wrong. “Helped run the European Forecast Hub” is a co-authorship. “Senior Research Fellow” is not my job
  • Both models had read my email address and said so

Opus. “PhD at the University of Bristol … Supervised by Ellen Brooks‑Pollock and Hannah Christensen” right

Opus. “latterly a Senior Research FellowAssistant Professor

Opus.Helped run the European COVID‑19 Forecast Hub (with Katharine Sherratt, Johannes Bracher, Sebastian Funk)” co-author on its evaluation, Sherratt et al. 2023

Opus. “which, given the address this message is coming from, is presumably you”

Sonnet. “He was based at the Centre for Mathematical Modelling of Infectious Diseases (CMMID) … where he worked as a research fellow” still there

When the writer and the reader both have a language model

What is the person for?

What value are you adding?

Thank you