Coding agents in a research workflow
London School of Hygiene & Tropical Medicine
30 July 2026
EpiNow2, epinowcast, scoringutils, primarycensored, and now EpiAware in JuliaSo: someone who writes a lot of research code, maintains it for other people, and has always worked in teams. Scholar
I wrote the argument. An agent wrote the words, found the numbers, and drew the charts. It took about two hours, and you can read the prompts at the end.
Note
Everything that follows is a description of a working setup, including the parts that do not work.
A grant, this week
Three of us each wrote a brainstorm file. Those became a 2,000-word expression of interest over 76 commits in two days.
pi is the tier belowocto.nvim for pull requests, tuicr for agent diffs, both a pane away from the agents that wrote the codegithub.com/seabbs/skills · what I maintain is prompts and config, not services
Model: BVDOutbreakSize (Abbott, Sherratt, Brand, Funk); diagram this talk
Figure: fit to 29 June, when the late estimate was still rising. The 17 July fit revises that tail down. Estimates are 30% credible intervals.
A Workflow for Infectious Disease Modelling, Abbott et al., 20 authors. doi:10.5281/zenodo.19097427
seabbs-bot opened 271 and dependabot 45. The 21 under my own account were agent work toogit log mostly tells you who did what, and almost every bot pull request says a bot opened itepiaware.org — component distributions combining into one joint distribution
EpiAwarePackageTools.jl holds the shared test and quality utilities, used by 9 of themOut of my own pocket. My institution does not cover it, and neither do most.
deepseek-v4-flash does the mechanical tier of my own setupThe question
If output per researcher now scales with spend, what does fair look like, and who is supposed to pay?
If a task is well-specified, bounded, and checkable, assume it is already automated. What is left is work where the hard part is deciding what the task is.
Every boundary I have drawn in two years has moved. “Agents cannot do long-horizon work” lasted about six months.
The big question
What is genuinely ours, rather than merely not-yet-automated? I do not have a good answer, and betting a career on a specific capability gap looks unwise.
A Quarto revealjs deck for a 15-minute talk on coding agents in research. Match my
how-to-serial-intervaldeck.Themed use cases: code, review, bullets to prose, focused edits. Tooling: Claude Code,
piwith DeepSeek, tmux,tuicr. The DRC Ebola work, which was very agent-driven.seabbs-botfor review at scale.epiaware.orgfor standards across repos.Cheap LLMs are ending. I pay £200 a month myself, institutions do not cover it, and access is unequal. We need work LLMs cannot trivially do — what is it?
Sparse slides, images where you can. Subagents to scan my recent work for material. A review loop against the brief and my old slides. Serve the preview over Tailscale. Put this prompt on the last slide.
What happened
Note
I wrote the argument and almost none of the words.
That is 36 slides and there are not enough pictures.
Make it more fun.
Check that what the slides say is actually true against my git history. Use a workflow.
TDD is right, leave it.
Add the always-on box with nothing confidential on it.
octo.nvimin neovim for review. I dropped the heavy LLM tools for parallelism and to stay near the terminal. Partly preference.Open weights are getting better. And £200 is the floor, $50 a day is easy.
Put these steers on a final slide.
This is the actual skill
Important
The prompt is not the work. The steering is the work.
Important
Tell me what you think agents cannot do.