Listen on your favorite platform:

Christopher Krapu was a self-described terrible math and physics student in high school who now spends his days building agentic AI systems at Nvidia, after a PhD at Duke and years of Bayesian hydrology research. He's also a longtime PyMC contributor!

The thread running through this conversation is a kind of intellectual restlessness that keeps landing, somehow, on the same underlying question: given sparse, noisy, or structurally weird data, what's the minimal piece of prior knowledge that makes an impossible problem tractable?

Chris has asked that question about flooding wetlands, gold deposits, thirty years of missing rainfall, and now about agentic AI systems built on top of large language models.

One Fact About Plants Fixes an Unfixable Problem

Chris' favorite illustration of Bayesian thinking, the one he uses to convert skeptical frequentists, involves wetlands. He modeled the water depth of ponds and wetlands using satellite imagery so coarse that a single wetland might be represented by a handful of noisy pixels, with the actual depth never observed directly.

The model was hopelessly underdetermined until he added one small piece of biology: the plants visible in the image don't grow in water deeper than a certain depth. That single constraint took an estimate that could plausibly range from less water than a drinking glass to more than Lake Superior and collapsed it into a tight, physically sensible range.

A related trick shows up in his work on Gaussian processes for mineral prospecting. In geostatistics, the classic model (known in mining as kriging) assumes you know exactly where each sample was drilled. Chris' project broke that assumption deliberately: the recorded coordinates were only accurate to within a rough radius.

By treating the true sample locations as latent variables under a Gaussian process, jointly with the measurements themselves, the model could still reconstruct the underlying gold-concentration field, even without ever knowing precisely where the samples came from.

In both cases, the lesson is the same: Bayesian models can absorb far more structural uncertainty than intuition suggests, as long as you're honest about what you don't know and inject the one fact you do.

The GPU Decision Tree

Chris has a practical decision tree for when GPU acceleration is worth the trouble: if your model's computation graph is dominated by large matrix operations, even ragged ones you can pad and mask, GPU acceleration tends to pay off handsomely. If it's dominated by sequential, recurrent operations with small dimension at each step, less so.

He argues the software has largely caught up: PyMC's JAX backend and NumPyro make GPU-accelerated modeling work out of the box today, a sharp contrast to the driver and compatibility headaches of the early Theano-on-GPU era.

What's missing, in his view, isn't capability but shared knowledge: companies clearly run large Bayesian models in production, but the results stay behind corporate firewalls, and he's floated starting a community benchmark effort to document what actually works on which hardware.

Agentic AI Through a Bayesian Lens

Formal Bayesian models aren't part of Chris' daily work now, but the intuition keeps showing up. Speculative decoding, a technique that uses rejection sampling to align a small model's outputs with a larger one, makes immediate sense if you already think in terms of priors and acceptance ratios.

The same training helps with reasoning about the high-dimensional geometry of modern embedding spaces, and with building hierarchical models to evaluate agentic workflows, where a single interaction now involves a whole sequence of tool calls rather than one chat output, without a corresponding increase in the amount of evaluation data available.

Looking Ahead

Chris wants to see someone try speed-running the training of a very large Bayesian model, in the spirit of the community that competitively speed-runs GPT-3-scale language model training, to find out how big and how fast is actually achievable with today's tooling.

He's also drawn to the sheer volume of data now available about how machine intelligence behaves: model weights, outputs, long agentic traces. He sees an opening for genuinely Bayesian analysis of that data, work adjacent to mechanistic interpretability, and for building out evaluation frameworks for long-running LLM agent processes that treat a full trajectory, not just a final answer, as the object of study.

Check out the full episode above, and the show notes for links to Chris' papers and blog posts.

You can also interact with the episode on NotebookLM! Ask questions, generate flashcards, and more.

Hope you enjoyed it, and see you in two weeks, my dear Bayesians!

Chapters:

22:57 When does GPU acceleration actually pay off for a Bayesian model?

26:33 What did it cost to train a two-million-parameter model on Modal?

30:36 What happened when Chris asked 200 different LLMs to flip a coin?

34:50 Where do Bayesian ideas show up in the agentic AI systems Chris builds at Nvidia?

40:16 Are statisticians being made obsolete by large language models?

41:19 How does putting a Gaussian process on unknown coordinates fix noisy data in mineral prospecting?

58:05 What is Chris looking forward to working on next?

My guest today is Christopher Krapu, who spent years building Bayesian hydrology models before landing at NVIDIA, where he now builds retrieval and agentic AI tools.

Christopher is also a longtime PyMC contributor and one of the rare guests on this show.

I had actually met in person before we recorded.

Spent years applying Bayesian methods to hydrology and wetland ecology, including a clever trick where a single fact about how deep certain plants can grow was enough to stabilize a

badly underdetermined model.

We also get into his experiments training a 2 million parameter Bayesian model on model for pocket change and how the same Bayesian instincts that shaped his hydrology work now

show up in his day job.

Building agentic AI systems at NVIDIA.

This is Learning Basion Statistics, episode 162, recorded June 5, 2026.

Let me show you how to be a good basian.

Change your predictions after taking information.

And if you think it now be less than amazing, let's adjust those expectations.

What's a Bayesian?

It's someone who cares about evidence.

Welcome to Learning Basion Statistics.

A podcast about Bayesian inference, the methods, the projects, and the people who make it possible.

I'm your host, Alexandre.

You can follow me on Twitter at Alex underscore andorra, like the country, for any info about the show.

LearnBayesStats.com is Lamplace to be.

Show notes, becoming a corporate sponsor, unlocking Bayesian merch, supporting the show on Patreon, everything is in there.

That's LearnBayesStats.com.

If you're interested in one-on-one mentorship,

online courses or statistical consulting, feel free to reach out and book a call at topmate.io slash Alex underscore and Dora.

See you around, folks, and best Patient wishes to you all.

Chris Craypoo.

Welcome to Learning Vasion Statistics.

Well, thanks for having me on.

Really glad to chance to get a chance to talk to you.

Huge fan of the work you've done, both in terms of the podcast as well as on IMC.

Really excited to have a discussion.

Yeah, thanks a lot.

And I'm super happy also to to have you on the show.

interestingly, you're one of the only guests I think I've met in person before having on the show.

Um usually it goes the other way around.

but but yeah, well I had the the chance to to meet you a few weeks ago now now that we are both in the Bay Area, so so it was really fun.

and I'm very happy to have you on the show because you also have done a lot of things on the PMC side.

I knew you from um your contributions to the the examples and and really fun models you've been working on, which is what we're gonna focus on a lot uh today.

But first, you know, as usual, we'll go through your origin story because uh nowadays you're working at NVIDIA, um, but you've done a lot of P uh a lot of Bayesian stuff and

you've also done a PhD at Duke.

So how did that happen?

You know, how what's your origin story?

Yeah, well it it starts way back when I was in high school and I was actually a terrible, terrible student in math, physics.

I did take statistics done, hated it.

And actually this is not an uncommon story for people that end up

Loving the Beijing paradigm for modeling.

Um but at the time I was also very interested in environmental issues related to the area where I was from.

We had a lot of flooding.

We were kind of unique amongst areas with um, you know, an ecosystem that had a lot of lakes and ponds, not a lot of rivers.

And so I was always thinking about water.

Water was a central topic for agriculture, for disasters.

It dominated a lot of the political discourse too.

And this is, you know, the uh late nineties, early two thousands.

Now that was the back of my mind when I went to college and you know, I got to college and for some reason I took one class with a a wonderful, wonderful professor of physics there,

Jim Doyle, McAllister College.

And um I was just hooked.

And I didn't really care what it was.

I just wanted to keep doing more of that, taking more physics.

And I didn't really have a plan after that.

Nothing.

Um but then came I think junior year, and I was taking my sixth class with this professor, Professor Doyle, and he was teaching statistical physics.

And

For those of you who don't know, there's usually four classes you have to take as a physics major.

It's some combination of electromagnetism, quantum mechanics, classical mechanics, and statistical mechanics.

And people hate stat mech, which is statistical mechanics.

It's more probabilities, it's less deterministic rules about balls rolling down ramps and about charges moving in a force field.

And it really chips a lot of people up because they're not used to dealing that.

but that was the one subject that I liked the most.

And

I was kind of sad at the time when I was graduating because I knew I didn't want to go to grad school for theory.

I didn't think I was perhaps competitive enough because people had usually built strong pedigrees by the time they graduated, or at least they'd sh sh shown strong signals.

And I also uh wanted to have a more applied career at the time.

And so um, you know, I was looking at things to do.

And engineering grad school is always a good choice if you has a if you have a physics major, because a lot of the skills transfer.

But I had no idea when I got to grad school that um when I took the first few classes in Bayesian statistics at Duke.

That it was the exact same topic as StatMek from a formalism point of view.

And um Duke has always been a hotbed of Bayesian research in America, um, you know, prior to the 90s, in fact, is one of the few places where there were many faculty who were

working on it.

And ever since the advent of MCMC, um, there's been a proliferation of applied work there too, because in fields like ecology, political science, health sciences, all of that um

Bayesian modeling has been transformative for them.

And so I got to this intro series of classes at Duke, uh intro for grad students, but like, you know, you take in your first, second, third year of a PhD.

And it was the exact same math.

You ha you just rename everything.

And so in stat mech, you have something called the normalizing constant or partition function.

It just becomes the evidence or like log like or uh marginal likelihood and Bayesian stats.

Algorithms are the same.

You approach things the same way, and it that was awesome.

I loved that.

and I was also extremely fortunate at the time because this was about twenty fourteen, twenty fifteen.

That was also when PyMC3 was starting to come to the fore.

And it was the first time I'd seen such, you know, easy turnkey plug-and-play um software in Python for like arbitrary probabilistic models.

And you know, I think of this kind of like the golden age in my career when I was playing around with this so much.

And one one experience I remember that very, very strongly influenced me was um when you could hook up packages depending on Theano.

And so Theano is an older version of, you know, uh

Tensor machine learning framework for those who don't know, like TensorFlow or PyTorch.

And you could hook it up to PyMC.

You could use packages that use Theano in a million other ways, like Keras, for example.

And then you would build Bayesian models from scratch.

And then most of my applications in grad school for that were mostly taking uh problems in hydrology or ecology, environmental science and engineering broadly, and then trying to

fit them into Bayesian models.

Okay, yeah.

So um a journey of uh a lot of curiosity.

I hear.

So so yeah, really love that.

Where you were really curious about the um the application per se.

And and Bayesian stats were really something that helped you solve your problems.

Um actually, is there anything in particular or somebody in particular that drew you to Bayesian stats specifically?

How did that happen?

Yeah, I think there were a couple of classes that I took.

Um there were several influential faculty at Duke who are

excellent teacher.

So I remember sitting in on lectures from uh someone who was a postdoc at the time, his name was Jeff Miller.

I think he ended up becoming faculty um I think at Harvard or somewhere else.

And, you know, he was an excellent teacher and he gave the first course where I got to see sort of building up from prior to Nykahoods, MCMC, turning the crank, and then um as they

say, getting to crack the Bayesian egg and enjoying the omelette.

Um and then I think

Later on, I got to take a class with Alan Gelfan.

So he was very influential for writing one of the first papers um that helped disseminate MCMC from the physics community into the statistics community.

And he was also extremely active in research for ecology and environmental science, which was also like my core application area.

And so that that was wonderful because there you got to learn from someone who had seen the breadth of both Bayesian methods as well as that methodology applied to problems.

I deeply cared.

And I was very fortunate.

And so those were some influential people I remember from my time there.

But to be frank, almost all the faculty I engage with in stats and also those who did stats for um environmental science.

So one example would be Jim Clark at Duke.

And so he was faculty in both stats and environmental science, which is, you know, fairly rare at universities.

Um, well known for his work in uh primarily for looking at studies related to trees and vegetation and climate change.

You know, great guy to work with as well.

Okay.

Yeah, yeah.

Yeah.

So

a collection of of influences you got.

I I love that.

And actually you went I think it's it's very it's gonna be very interesting to listener to hear about that because you went from Oakridge's geospatial group to now NVIDIA where you

build retrieval and agentic AI tools.

So from the outside, that looks like a sharp pivot.

Can you tell us what actually changed in your day to day and

Where does spaith in sync Bayesian thinking still show up or not, maybe in in industry AI work?

Yeah, that's a great observation because these topics still look kind of disconnected, but having lived through it, to me it seemed like a smooth continuous progression.

And I'll maybe explain how that happened.

And so um I was working at Oak Ridge starting in twenty twenty.

I was primarily focused on a project we where we were taking uh models of buildings.

And so the basic

use cases, you want to know what the internal structure or properties of buildings are.

So in the event there's a disaster like an earthquake or nuclear detonation, that you can assess roughly how many buildings are likely destroyed, where do you send the ambu

ambulances and so on.

And this is especially important for low resource areas where you don't have much data.

And so that's why we had a model that was primarily working with a lot of subjective data, like buildings in um this part of the world will primarily not have a strong steel

superstructure, or they'll be made out of brick or mud.

Now, one of the main challenges we had at the time was that when you have a little bit of information, that's great, because you can do a lot with even just very small subjective

pieces of information in Bayesian modeling.

but there were cases where we truly had nothing that was not qualitative.

Qualitative in the sense that, you know, an anthropologist might have made an observation about how people build certain types of huts and mechanical properties and what might

happen to them if there's an earthquake, but we didn't have more than that.

And so always at the back of my mind was the problem of how do you construct a good prior.

in this problem for saying what is it likely to what is likely to be the case for this building or this object um when you have virtually nothing to go on.

And how do you take all the information that is um summarized by the world's text and literature and engineering reports and so on and apply it to our problem.

And this is not easy because uh you know there's machine learning models in NLP where you can, you know, take text and transfer them into some sort of numerical representation.

But back in like twenty nineteen, twenty twenty, a lot of that was was still not widely spread outside of

primarily the NLP literature.

but I remember around um the end of twenty twenty one, I'd finally gone off the wait list for a GPT three API key and I I don't remember why I signed up for it.

I think it was just seemed kind of neat at the time.

And the very first thing I did when I got that API key from OpenAI was I made a bunch of fake news articles from The Onion.

So it's a s satirical newspaper, uh you know, humorous titles made up things, just to see, you know, could it reproduce something that had a spark of human

humor to it and it's mediocre at that.

The second thing I did with it was that I took a prompt which was meant to mimic an engineering report about a building and might say this is a building in Knoxville.

It was built in 1895, it has a brick exterior, it has a sloped roof, the roof is copper, um, and the internal superstructure is made out of blank.

And this is in the pre instruct model era, when really to get a GPT class model to work, you had to give it the data as expected to see it in the real world.

not instruct tuned formats.

And then they would just fill in.

this probably has an internal wooden superstructure.

It's going to have steel beams.

And I remember taking that into work, you know, the next time we had a group meeting and saying, you know, we've got there's this powerful new model that's there's a tech report

from OpenAI.

You know, I did a little bit of playing around and I look at the distributions it puts out.

And it's not perfectly calibrated, but it's for sure better than nothing.

And uh the reaction I got wasn't necessarily negative or mean, but it was just

Kind of a blank store, what is this?

and you know, I acknowledged it was a little bit of a out of left field at the time, but looking back at it now, that was like the first time that I thought, this could be really

useful for for hard problems and really neat too.

And so just pulling on that thread of trying to fit LLMs into more and more interesting problems and often using them as sort of a reference prior, uh, a very, very informed

prior, perhaps, not at all objective, but some sort of background prior you can use in the absence of anything else.

That was also a theme for some of my earlier work as I transitioned into primarily Gen AI stuff, with some also with also some more Bayesian work on the side.

I'd love to turn a bit more into your technical work now because you have a lot of that and you still have an active uh blog which I recommend people to to check out.

You have, in particular, a few papers on, well, you know, uh environmental data and environmental models that I found really interesting, especially one

In twenty nineteen.

In this paper you introduced probabilistic programming to environmental models, which as a statistician who tries to make statistical education more affordable, I definitely

resonate with.

So what gap were you trying to close with that paper and then your subsequent work and now a few years later

Has the field actually moved?

the short answer is happily yes in some limited ways.

Some of the big problems that we were trying to address at the time, um, are still open.

And so the problem that we were really laser focused on at the time, which was kind of a big, big problem for that field, was um, you know, in the subfield of civil engineering

focused on water resources.

So that's hydrology.

People who care about is there going to be a flood?

is this basement gonna gonna be saturated with water?

Are we gonna have enough groundwater to feed um

The irrigation pumps for the next decade, things like this.

There's a whole range of modeling approaches.

So there are very, very finely detailed physics models where you're simulating fluid dynamics in a porous medium on a centimeter scale, and you're trying to model where all

the roots are, where all the plants are, because the plants are using the water and where there's cracks where the water can flow down deep into the soil.

And that's pretty hard because you have to characterize the entire subsurface.

Generally not feasible in most places.

And so there's a lot of intermediate.

Buxty approaches where you maybe split up your domain, your spatial domain, like maybe a hillside or even a city, into a bunch of units, maybe more than 10, less than a thousand.

You assign them some sort of high-level parameterization that on average the water is going to infiltrate like this, there's going to be this much runoff, and so on.

Um and then there's very, very simple models where you just treat the whole thing as a bucket and a sponge conceptually.

Water comes in, water comes out.

It's very simple.

Now, the problem for that intermediate range of models is that

Generally, there isn't a good way to calibrate them from the hydrology literature.

A lot of it was just guess and check.

And actually the hydrology literature was um they were coming up with ad hoc versions that were very, very close to variants of sampling algorithms like rejection sampling, and even

some things that kind of looked like MCMC if you made some slight tweaks, but almost completely ad hoc and without the good theoretical guarantees because those literature

those sub literatures were disconnected that they didn't really know about Bayesian inference, and a lot of this is happening in the 70s, 80s, 90s, they're not just simply

not talking.

So

Um I didn't run really want to deal with that because I'd seen some nasty cases where models were um under constrained by the data sets.

The calibration results were unreliable.

It was bad practice and bad science, in my opinion, to keep doing that.

And so we were thinking, okay, how do you take a whole field, in this case hydrology, and get them to do the right thing, methodologically speaking, which is to use grounded

methods for fitting the parameters of your models in a way that guarantees uh decent or good performance subject to whatever constraints you have.

And essentially the best thing we could find, and that's how I got into Bayesian, applied Bayesian methods, was that if you just take your intermediate complexity model, so it's

not a completely pure physics model, but it's also not trivial either.

And you wrap the whole thing with Bayesian priors, you use the model internals to help get the likelihood, you know, you're out and you use some sort of reasonable likelihood on the

data, like a a normal or uh a positive value distribution for that.

And then you use MCMC, you often can get performance in terms of calibration that's

better, but you also get better guarantees and better uncertainty quantification.

Now you can do this without a probabilistic programming language, what you know, destroying everything yourself.

And some people did that, you know, 20, 30 years ago.

But that research did take off because you had to be both an expert at hydrology and you also had to be a good programmer and you also had to be a good statistician.

And you know the people in this world that were good at all of those, that's a very small set.

And so the problem of transferring that skill set and making the whole toolkit easy to use

was something that I thought uh PyMC, but also Stan and other peers at the time were doing a very, very good job doing.

And so we're trying to push people to think about putting priors on your models, thinking about the likelihoods, uh being aware of representation issues.

So in Bayesian modeling, we often face things like the funnel or, you know, noncentered versus centered representations, um, or uniqueness and uh exchangeability.

And to get those terms in that literature as

And so that paper, which talked about some examples of probabilistic programming and how they translate to environmental modeling, was broadly focused at tackling that big problem

in hydrology, which is that the model calibration processes were f were in some ways severely deficient, methodologically flawed, ad hoc.

Um but with a little bit of adjustment, many, many things could be solved.

Yeah.

Okay.

Yeah, I think VCs like with the you know the

the diagram, the Venn diagram of of people who could do that work that you were saying.

Yeah, it's like it's it's getting it's definitely getting small s and smaller as you as you describe it and add uh skill that people need to have.

So adding that to these kind of packages is very important.

And it's actually stuff that you did, right?

Because I know that you you did some work uh in 2016, which was originally how you were

First thing for you to PyMC.

Um so you wanna you wanna talk a bit about that?

Because I think it's uh it's very interesting to to mention on the show.

Yeah.

One of the very first papers I ever did using PyMC was is trying to be a shot across the bat to the hydrologists.

And basically we took a relatively simple model, um, so not that many parameters, maybe less than 10.

Um but then we gave it an extremely hard problem from a hydrology point of view, which is

You're gonna have, you know, many years worth of rainfall data, let's say 30 years of daily rainfall.

So 30 times 365, you know, you're talking 10,000 observations.

Um, and you're also going to assume that the rainfall inputs are unknown.

You don't know what the rainfall is.

And we're gonna we're gonna do an inverse estimation of the input rainfall.

And the hydrology literature at the time, this is kind of a toy problem.

It's not like something we care about a lot because it's easier to measure rain than it is to measure anything else.

Um

But it was an interesting problem because it was a way that they would kind of probe the aspects of the model to see if it would map from um, you know, a high dimensional input

space to, you know, a reasonable set of outputs.

Anyway.

So we were trying to estimate the rainfall inputs of which there would be daily rainfall.

So that's, you know, 10,000 inputs and do that um in a reasonable amount of time.

And the state of the art back then was like, we can do this for like monthly rainfall for like two years.

So like, you know, 10 to 20 data points.

And even then they did

very poor uncertainty qu uncertainty quantification.

So basically we just wrote a simple hydrology model in PyMC in Theano at the time.

We used PyMC to place the priors on it.

And then we used both Hamiltonian Monte Carlo and the variational inference to get to get estimates.

And they're both pretty good.

I think they do the variational inference was much faster because we were also running on GPU at the time.

em And we were able to accurately recover you know within the bounds of what the model could tell us about the incoming rainfall given the output

um river levels.

And we could do that for the entire 30 years.

And I'm not entirely sure why this paper didn't get more discussion.

I think perhaps maybe it was uh the Bayesian stuff was still a little bit too new to the hydrologists.

But to me it was a pretty strong validation of this way of attacking problems.

So you just wrap a simple model, you know, with this framework and using methods that we know s uh scale better than either simple MCMC, which doesn't have the gradients in it.

Um

Or ad hoc methods, which are just guess and check more or less.

And that it worked beautifully and it it beat the state of the art by, you know, four orders of magnitude here.

So that was a major coup for us.

I also want to have your thoughts about at like as a longtime Pime C contributor, what do you think the current state of Bayesian at scale tooling is and and what's the bottleneck

that still needs to fall?

Yeah, it's a wonderful question.

I think about this actually on a weekly basis, probably for the last couple of years.

So for those who haven't ever tried to do big Bayesian models, and I'm saying big, like you have more than 10,000 data points or it takes you more than like two days to run a

sampler.

Um it's actually quite good if you use any of the existing platforms that have GPU support.

And I'm not saying this simply because I work for NVIDIA, although doing MCMC on GPUs was how I got into doing this work originally for NVIDIA.

Um

But now, you know, there's excellent options in PyMC obviously, where if you use the Jax backend by and large for almost anything, it just works out of the box.

And I know it's taken a large amount of work both to do the transpilation of Jax and as well as address the edge cases where um Theano or Aesara or PyTensor maybe had slight gaps

and then to fix those up.

Um, you know, or if you want you're more experienced with Pyro or NumPyro, it's just use that, you know, use that back end directly.

And I've been using that quite a bit just for side projects and finishing up papers in grad school.

And it's been excellent.

And I think that the compatibility also is is actually probably easier than when I was first starting out and there was direct GPU support in Theano for PyMC because there were

still issues with like uh CUDA and drivers and compatibility.

And frankly, I spent more time back then, even though Theano supported GPU as a first class compute backend as compared to today.

So I think that the overall

the software has uh the world is catching up to capabilities of the software.

Um I also think that there is a lot of excellent work that's being done for large scale models, but a lot of it's siloed behind corporate firewalls where people aren't talking

about it.

So, you know, Uber released Pyro, they clearly were thinking about doing b large, large Bayesian models.

And you look at their business model, you can understand why.

Um, you know, there's financial firms and other firms that do this as well, but it's not publicly disseminated.

I I think um, you know, maybe one thing that does need to happen a little bit more is that we just have a bit of a community effort.

And I'll volunteer to start this or do something like it, where you just have benchmarks for what can be done easily in which um programming languages or packages in terms of, you

know, I want to fit a Markov random field model with a million parameters on my GPU.

Which platforms does it work in, which frameworks does it work in out of the box?

And this is a pretty poor thing for package maintainers to do because

You're you're testing using expensive hardware um and things that take a long time to finish.

So it's a very, very bad match for a CI CD workflow, but I think it's still something valuable that community should probably pick up.

Yeah.

Yeah, yeah.

Um completely agree with your with your lay of the land here.

And actually you have a blog post if I remember correctly, where you do some

some experiments that you called poverty base.

that was quite provocative.

you fit so if I remember when I was preparing for the show, I looked at it and um if my memory's correct you fit a like a huge model, something like two million parameter patient

model on model for about fifty cents.

can you walk us through that setup and like what's the actual stack?

What does that price point?

Open up for people without institutional compute budgets, which I think is is also a lot of people.

Yeah.

Yeah, I'd love to talk about that.

And the title was a little hyperbogged.

The original title was going to be Big Mac phase because I thought it was going to cost me five bucks.

But I made two mistakes here.

First off, Big Mac isn't is much more expensive now.

And second, I overestimated by an order of magnitude how much it would cost.

And um

When I was writing it, I was thinking back to grad school where I'd gotten this $2,000 Titan X P GPU, free from NVIDIA, fortunately, but it was still quite expensive to produce.

And I was running chains on that and I was thinking, gosh, you know, that was kind of a like a high barrier way to get started.

And so if I were to redo this now as a fresh grad student with no budget starting in 2026, here's how I would do it.

And um I I do have to advocate a little bit here for Modal just because I enjoy the experience so much.

Essentially you just write a Python script, you deploy it on their servers, GPU included.

not very much build time it deploys quickly.

And you're right.

It was a hierarchical logistic regression model.

I think there was close to two million parameters, um, and even more observations.

the kind of model that if you tried to use a Gibb sampler on, you'd be watching it go for weeks and weeks and weeks and weeks and might not ever converge just because it was too

big.

And because it had because PyMC has that Hamiltonian Monte Carlo, which scales much, much better in terms of mix and time, then that's not really an issue.

Your issue is rather if you can just execute the likelihood function

or the gradient of the likelihood function fast enough.

And that's what the GPU is for.

um And I think that there's again, there's massively under underutilized hardware out there for this kind of work.

And that was on like just one A100, if I recall.

You know, there's uh newer ones, there's the Hopper and there's the Blackwell GPU.

And the jumps for those in terms of floating point 32 performance isn't as stark as it is for things that are useful for LOMs like floating point eight.

floating point sixteen or BF sixteen or even the NBFP4 formats that are popular nowadays.

Um but there's still big, big jumps for us in what we do.

And I think that like a lot of work can be expedited.

And I think there's also a lot of business value that can be unlocked um by having people prototype on these cloud platforms, especially because you know for us, big data is big in

the terms of models that take a long time to sample and for the chains to mix, but they're not so big that like

need to have a whole data infrastructure team and platform to get your data onto the m onto the modeling container and run the MCMC.

And so some things that might be harder using other approaches are kind of easy for Bayes because we don't need that much data compared to like, you deep neural net.

Hmm.

Okay.

How would you like do you have an a tree decision tree in your head that you can help people use to take these kind of decisions depending on the

on the projects they work on?

Yeah, yeah, really good question.

I do.

A really, really well formed one.

the first one is if your model has a lot of recurrency or you find yourself using operators like the for loop or scan in in PyTensor, um forget what the equivalent is in

JAX.

Um where the compute graph or the likelihood or something like it is going to be a bunch of chained operations.

That's, you know, not always a great fit.

Like if each of those operations does have a large dimension in it.

Where you're doing a large matrix multiply, then okay.

But really like the the excellent applications for this kind of GPU accelerated Bayesian computing is where you do have big matrices or big arrays.

Or even if you don't see how it comes out of one variable, maybe you can sort of shoehorn things into uh you know, a large array of some sort, like even if you have ragged data

where not all the op not all the groups have the same number of observations, well, you can just do some padded computations.

You mask out the ones that don't contribute so on and play this sort of game.

And then still find a way to do uh efficient computing using the GPU.

ah I think that's it's a pretty simple decision tree.

I think also there are edge cases for PyMC and umpiro as well, where when you use certain exotic types of model components, like for example, uh certain types of Gaussian processes

or maybe certain new kinds of um distributions that might have specialized mathematical computations in them, those might not all either translate cleanly to the GPU in terms of

operators.

Or um they might not offer much extra uh efficiency for compute just because of the nature of how that thing's done.

Okay.

Yeah, that's that's very helpful.

I mean I'm gonna iste that myself because you know much more than myself about uh about when the the GPS are gonna be most helpful or not.

So these actually um very interesting to to hear you talk about that.

another another question on that topic I had actually while preparing for the show, I saw that

You had a post where um you asked two hundred different L and Ms to flip a fair coin, uh so that immediately got me hooked, and then to analyze the bias with a bi Bayesian

binomial regression.

So that was super fun.

Um to to do that and to read that.

Um what was the motivation behind that and what did you find?

You know, I I like yeah, I'm very curious.

Like it did what does

What does that say about the different models that you've tried?

If if that suggests about, you know, what we're optimizing for when for when we we fine-tune these models.

You know, what what did all of that tell you?

Yeah.

So I'll you know, I'll be honest, that was a little while ago and around the time when I was doing that, that was also when our second child was born.

And so like the intervening periods of sleeplessness.

or lack of sleep, uh cloud my recollection of exactly what I was trying to do at that time.

But I do recall my main motivation for doing something like this.

And that was related to a series of posts I'd seen from other people in the community who would just sort of like, look at the outputs from a bunch of LLMs.

And now that we have enough models coming out and people are doing interesting things to fine-tune them, to get different styles or different capabilities, you can almost start to

think about doing a large, large sample analysis of, you know, LLMs.

Because of the diversity of approaches people have tried.

And I remember being particularly inspired by um, and I forget exactly who did it, I'm sorry, but um analysis where you have an LLM and you do this for each LLM that you're

considering.

You the ones from Microsoft, the ones from OpenAI, the ones from Google, the ones that some random person fine-tuned on their desk.

And you ask it, you you give it um a coordinate.

pair for each point in the world, like a latitude longitude, and you ask it to predict is it water or is it land?

And you do this for every latitude-longitude coordinate on a grid.

And you just draw a picture, which for you know a perfect model is going to draw a map of the Earth with the coastlines and the land and all of that.

And then sometimes they don't do that correctly.

And it's really interesting to see the patterns that come up across different model families.

And I'd like to do more work in this.

And this is kind of a first uh step in that direction and I'm going to hopefully try to follow up where

Just simply asking the LM to reproduce things based on representations that aren't very ergonomic to us.

Like, you know, if you ask me Lat Longs, there's maybe like a handful that actually could meaningfully tell you for sure if they're land or water based around places I grew up or

places I had to memorize.

but then for LMs, especially the big ones, it's like no problem.

They know that this coordinate pair is in the middle of the Pacific Ocean or that it's in the Great Lakes.

And so probing them that way, which is definitely my blog post, uh, was always of interest to me.

And so um when I was kind of working on this blog post about flipping coins, you know, it it was a little disappointing in the sense that I was I was hoping there'd be strong,

consistent variations across model families.

And and it seems like the a lot of it might have been just the uh overall, um, you know, the models are pretty similar.

I forget exactly what the results were.

and there wasn't like a clear, it's the small models that are always more biased or so on.

Um, but I've definitely got more thought experiments like that that hopefully we'll try out in the future.

And if there's any that you uh have thought about but you don't wanna write the code or you don't want to pay for the open router API key credits, I'm happy to do it.

Yeah, that was that was a really fun one.

And I know that was mostly, you know, for a thought experiment, nothing very serious.

but I think it's also great to have that kind of of uh content out there.

So so yeah, like thanks for doing that.

Even even

why in a sleep deprived state uh having your your second child um well done.

Like yeah, I'm not sure we'll be able to do that because this was still a lot of thinking that went into that that experiment.

Um something that will be much closer to home for you is what you do every day.

You know, you you build adjating systems uh at NVIDIA that retrieval, tool use, multi-step reasoning.

So from a Bayesian point of view

Where do uncertainty and priors actually fit into that kind of stack?

And I'm also curious about the failure modes in agents that a Bayesian eye like yours sputs first.

Yeah.

Um so with the disclaimer that a lot of this work there isn't really a strong Bayesian tie, and I think that a lot of it was made considerably easier to get into because I

understood because a lot of the research literature was framed in terms of Bayesian computation, at least at

terms of things like priors and structural priors and likelihoods and temperature and sampling and all of this.

Um, that makes it a lot faster to keep up with.

Uh there aren't strong notions of uncertainty in the core work that I do in most of my day-to-day, I still consult a little bit internally for forecasting where Bayesian models

are important.

And that's rewarding and fun.

Um, but that's not my core uh responsibility.

What I do see really helpful is that when research literature comes in, like a new paper says, this is

speculative decoding and there is an algorithm where we use rejection sampling and a lot of people look at that and they're like, oh, like I have zero idea why that would work.

That you sample and maybe you throw away and somehow this is gonna match the performance between the small bond and the big model.

whereas if you're coming at it from a Bayesian point of view, you know, rejection sampling is an older technique and it's not strictly Bayesian, but it's something that often gets

brought up.

And so it's for me it's really fun when that happens because you can tell a lot of the folks who do cutting edge work.

For LMs nowadays, we're deeply, deeply instilled with some of those principles of probabilities for modeling, uh sampling to do wonderful things, um, and that toolkit for

attacking problems.

and then also like thinking here's one more thing also that's quite important.

Um Bayesians are also especially well equipped to handle problems where there's high-dimensional geometry involved, because a lot of the motivation for doing a Bayesian

modeling in the first place was that you need to regularize some priors for problems where your data is

Insufficiently large to constrain the problem relative to parameterization.

And so this sort of thing comes up all the time.

For example, one of the first things you might learn in grad school is like the Stein's estimator, where the best estimator of a high-dimensional mean isn't actually just a

sample mean, but it's some sort of thing that actually kind of looks Bayesian in the sense that it turns out to be equivalent to a regular IZE prior.

And from there, you get a lot of faculty in working with high-dimensional objects, numerical objects.

And a lot of that wasn't super useful from a practical point of view until recently, where now you have embedding models where they're, you know, 4,000, 8,000 dimensions.

And then some things you know you can do, some things you maybe shouldn't do.

For example, like can you interpolate between two points in that space and have a high guarantee that um the point that you pick is also meaningful in a semantic sense.

And that's one point also where the Bayesian training was invaluable.

Um I think also

evaluations are becoming harder and harder.

And so like early in the era of LMs, you know, if you're a data scientist working with outputs from AI, you know, you might have like A B testing to say this was preferred one

thing was preferred over the other and so on.

But it's not just about one chat output anymore.

It's usually a whole sequence of tool calls.

And this goes for not just like enterprise AI workflows where you might edit a spreadsheet or you might write read an email and then you might um update a record and database.

But even for

Like consumer-facing applications, like for example, there's a codex and cowork desktop app, and these things will do agentic workflows for the mass market.

And those evaluations typically will involve many different elements.

And as your evaluation gets more and more sophisticated, but the amount of data doesn't really grow, in the sense you have the same amounts of data as in the pre when workflows

are simpler, then you start thinking, okay, maybe I use hierarchical models, models of structure to figure out if

you know, approaches using this tool call or using this variant of model or using this hyperparameter are going to be useful.

And then, you know, the strength of the Bayesian approach for evaluations is that it's it's okay if you have a large number of factors in your analysis.

You know, it's not going to lead to numerical instability.

And if it's under constrain, then um sensible priors can often help you learn even from small data sets relative to the size of the model you want to fit.

Mm-hmm.

Mm-hmm.

Okay.

Yeah.

Yeah, yeah, I love that.

This is this is super practical.

Um thanks a lot for for working through that with us.

and actually now I'd like I think we've covered the part of what you do and how it it you know, it fits the vision and and and AI background that you have.

Um anything you wanna add on that, by the way, before we dive a bit more into the the different papers that you worked on during

uh the years?

I think at a high level, I just uh feel very, very grateful that a lot of the explosion and interest in compute for machine learning has spilled over so beautifully into Bayesian

modeling.

And it feels like, you know, there's just achievement after achievement um in terms of software frameworks, model size, um insights also into how high dimensional uh parameter

sets and objects behave.

Um that's made

I think a lot of interesting work possible on the statistical side.

And so, you know, people often say that statistics is is at risk in the Arab LM.

So I have two things to say that.

First off, the LLM as a model of intelligence is a statistical paradigm.

And so I would say the statisticians won with all due respect to computer scientists, many of whom are better statisticians than me.

you know, expert systems, deterministic logic, we tried a lot of that, but it looks like just learning from data.

Um, train against the likelihood sampling algorithms.

This is the way forward.

And so, um, you know, give statisticians to give themselves a pat on the back too.

And second, I think there's a bright future for all of us.

And I think that these innovations will continue to spill over and support the work that Bayesian statisticians do, and it's going to be uh another golden age to go along with the

advent of MCMC and Hamiltonian Monte Carlo and now proliferation of uh compute for very, very large models.

Yeah.

Yeah, yeah, yeah.

Definitely.

Um

Talking about large models, you've worked on you've worked on Gaussian processes.

And so you know me, I could not resist asking you about that.

So also because it's a very cool use case.

I I really loved it.

So you've used Gaussian processes to model unknown coordinates.

Uh if I remember correctly, that was mineral prospecting, where the require the recorded locations were noisy.

So

Can you tell us about that project?

Like what does putting a GP on top of latent locations actually fix?

And when do you start to care about that kind of correction?

Mm-hmm.

Yeah, I'm really glad you brought that up because that analysis was an example of a trick that um I've often used in introducing Bayesian statistics to people who are familiar with

stats and in this case coming from the literature where there was a lot of

analysis for uh environmental data or mining data where it's spatial and the core problem is that you sample a few points in space, you look at and maybe you're drilling a core and

you look and see if there's enough gold in the dirt.

And if there's more gold, good.

If there's not enough gold, bad.

Then you keep trying to figure out where there's um the best place to drill.

And when you have people that come from a certain statistical training and you need to, you know, get them exposed to a new idea, in this case the Bayes Bayesian paradigm,

quickly from a practical point of view.

What I've found is one of the easiest things to do is you take a model they're quite familiar with, in this case, the what the the Gaussian process model, which was also

called geostatistical model and in the mining literature is called kriging, after a prominent South African analyst.

And you take some core assumptions of the model that you always assume are fixed, and you just throw them out.

And this always throws people for a loop when you do it the first time.

And that's kind of what we did with that hydrology paper I mentioned before.

That you take something that you think, I need this.

piece of data as the input to the model and anything else.

So you simply can't work with it.

And so you know you take your model, you break some sort of core assumption in the Gaussian process model, you assume that you have data, it has fixed coordinates, you know

what the coordinates are, and you model the correlation structure.

You just say, let's throw that out.

We don't know where the coordinates are.

I mean we know to a certain very, very rough degree, we know they're roughly in some radius around the recorded um the measurement that's written down.

But but actually they could be in a big, big um radius around that.

And this is I I've seen this in person where people see this uh application that that simply doesn't make any sense.

It can't work, your problem is under constraint, this is not gonna work.

I'm not even gonna bother dealing with these.

And then you work through the computations, you explain how sometimes you need a very, very small amount of information to constrain the problem space.

And then in the end, you know, show them these slides where, okay, like we didn't really know where the observations were.

We did have the labels for them, but we can actually reconstruct the

Original the parameters of the data generating process, in this case, the true field that tells you where is the gold, even though you didn't really know where the gold

measurements were from, not to a high degree of precision.

And for those who are more mathematically sophisticated, you also point out that the computations are numerically stable, that your chains mix, and that you have all of the

guarantees for good inference that you tend to get from the other Bayesian models that people have studied.

And then it's eye-opening to see how those assumptions get broken.

And then

Keep building on that, and then you bring more and more cases where you break more and more assumptions, and you can also test these assumptions too.

Because one of the main reasons that Bayesian modeling had strong, strong uptake across people in certain experimental paradigms, physics, psychology, and so on, is that it

allows them to test more assumptions than conventional statistical models under the frequentist paradigm, under the inference paradigms they were used to, could easily be

addressed.

oh The differences between the frequentist and Bayesian approaches are not as great as reported by many people, but there are important subtle differences in how easy it is.

Especially for people that are not um, you know, the very, very top level of um expertise in those respective fields do.

And for this, being able to relax those assumptions on knowing where your data actually come from and still doing a good job at modeling, subject to that.

That's always been a fascinating case study, and that's why I want to write about it for that blog post.

Yeah, completely agree.

And what I find really fascinating about the about Gussian processes is how how coupled

complex of a mathematical object they are, but in the end, how intuitive and easy to understand and use they end up when you've been exposed to them and b mainly use them.

That's the most important.

Just use them and you'll see that they are actually extremely helpful.

And yeah, like I've I've used them for so many models, you know, sports analytics, electoral forecasting

any anything, really, you name it, they they can come and help you.

And there is just so much beauty about doing what you've done, which is like, well, we've got that latent coordinate that's noisy.

We can actually squeeze even more information from the data if we use a Gaussian process um in a smart way here.

And I think this is extremely interesting to show to practitioners because also now

It's uh it's much more doable with the current uh the current technical technological tooling, whether that's hardware or software.

100% agree.

Yep.

And so, yeah, as I was saying, a lot of your applied work was in wetlands, hydrology, spatial ecology, and that's also related to, you know, where you come from.

I think it's it's awesome to see that, you know, what what you've seen uh

growing up is also something that you love doing as as a job.

and these were are all domains where the data is sparse, it's messy, it's structured by geography.

So I would say this is really a domain where Bayesian thinking and mostly Bayesian stats actually let you do their um things that frequentist or pure ML tools can't.

But

Yeah, how good is my prior here?

what has been what has your experience been in that matter?

Yeah.

It is funny that I got into this.

And actually when I was growing up, and so for reference, my my father was a federal scientist for fish and wildlife and then for USGS.

And he worked uh he was an avian biologist ecologist, and so he studied how uh generally how migratory waterfall behave and in particular um focus on breeding as well as foraging.

And when I was growing up, I you know, I actually spent a lot of time around his research center and knew his colleagues, but I never, ever, ever wanted to do that.

And and first because there is a little, you know, running dialogue back in my head that's saying, you know, do the the same thing as your father in this day and age is slightly

dishonorable.

And also because there's a wide world of things out there.

What is the coincidence that I and my father are gonna want to work on the same, you know, broad subject area?

And then when I came to the first couple of years of grad school, I worked on some problems to hydrology and they're like kind of satisfying, but not super interesting.

I didn't really like them.

And then I was writing applications for federal fellowships.

And the f the fellowships were asking for things, you know, where there is a clear need to serve the American public.

Um in addition, uh the primary need was good research, but if it can also serve the public, that's great.

And I remember thinking, um, I also want something that's a good Bayesian problem and something that's kind of rich from a scientific point of view.

It's and it has not a lot of people working in it where there might be, you know, area to sort of spread my wings to junior scientists.

And at the time I'd also just been having chats with my dad about

You know, what's going on at the farm?

how high are the water levels lately?

You know, is it gonna flood over the road?

Um, is there gonna be a flood in our hometown and all this?

And then it just sort clicked like this is actually a really interesting problem from a modeling point of view, because you have a high degree of um measurement noise, you have a

state variable which is also very nonlinear, it's the amount of water in a pond, for example, and then you have you have a read on all the important inputs and drivers, like

The temperature, the water, the rainfall, and so on, but you never see anything with high precision.

At least not at a continental scale.

And so the specific problem that I spent the most time on in grad school is that if you think about a pond or a wetland with focus to focus, the distinction being a wetland

occasionally dries out.

Ponds don't dry out.

Um, you know, it's it's like a bowl.

It's got some water.

You know that the amount of water, if you can see from the outside, it

You know, and you know roughly how deep it is.

You know it's going to be kind of like a bowl.

So you use the geometry to parameterize the volume of the water.

It's not precise.

And also we can't see under the water easily with satellites, so we can't measure every single little wetland in North America.

But we know roughly they're going to be bowl-shaped.

There's going to be water in it, it's going to evaporate out.

Some amount of it might infiltrate into the soil, some amount of it might be taken out for farming, and then primarily they are replenished because water flows in uh from rainfall

and then also from melting snow that runs into it.

And so you have a a reasonable, straightforward physical model, flux in, flux out, but then nothing is observed nicely.

And versions of modeling efforts where we took these assumptions and didn't use a Bayesian approach were really, really rough because we didn't have enough data to constrain.

But then I and this is one paper that my biggest regret from grad school is not getting it out at the time and not picking it up and finishing the publication process.

was that if you take this approach where you model the amount of water in um

in this pond or this wetland.

And your observations are very, very noisy satellite data where each individual wetland or pond might be, from the perspective of a satellite picture, like ten pixels or two pixels.

And even in the largest case, maybe a hundred pixels.

And each of those are going to be very noisy as well.

But even with that observation, and then one more point, which is that you know that the depth is never greater than a certain amount because you can see plants growing, you know

that those plants don't grow in water deeper than X.

That single

fact was enough to constrain the model.

That little injection of extra information that you know there's plants there, the plants never like it in more than five feet of water.

And all of a sudden you go from highly unstable estimates, terribly ill condition poorly conditioned numerical operations, and then um enormous uncertainty bars where your

forecast could be, you know, the wetland is uh as shallow as a th is less water than a glass, as much water as Lake Superior.

Um

Down and constrain them with that little piece of extra information to a very, very useful and um physically consistent representation.

And that was also if I didn't need more validation at that point that this was the right approach, then that would have been it too.

That just injecting a little bit of extra prior knowledge, scientific information that really was more biology flavored, that, hey, these plants don't like deep water.

We see the plant there from a picture.

What can we say?

And then that constrains the whole model and lets us do very useful things.

And now it's hard to unsee that experience when I go through other problems where I think, hmm, what is the minimal amount of information that would constrain this and make it

easily, easily tractable?

Which I think is something that comes up quite a bit if you work with Bayesian models.

Hmm.

Yeah.

Damn.

This is this is so fascinating.

Um and I I really love that you're doing that and and hopefully you'll you'll be able to to keep doing that because this is obviously something you you really care about and and

and are making um a lot of useful work.

on and actually have a twenty twenty-two paper that combines uh hydrology model with time varying parameters and differentiable inference.

So that was super fun to look at.

I put it in the show notes.

Um can you tell us in simple terms what's the modeling problem?

Um and why does differentiating through the physics actually matter?

Great question.

And so going back to the modeling problem, uh the issue that we're trying to address with that paper isn't unique to hydrology, but rather it's it's a problem that comes up in

virtually all modeling fields where you're not working with the underlying direct physics down to the, you know, the cla the scale of classical mechanics to figure out what's going

on.

And so this is pretty much any discipline except for physics in some fields of engineering.

And the problem is that um, you know, you have uh variations in the model output, you see that the dynamics of your system are changing.

But you're not really sure if it's because the model itself isn't valid, because your assumptions are bad now in this new regime, or because the controlling parameters of the

system are changing.

And sometimes it's obvious.

Like sometimes you're simulating um the size of a wake behind a or let's give a list of physics boring to an example.

Maybe you're you're s simulating uh voter turnout and um voter turnout is is is a function of the weather.

The weather is changing, and so you know obviously voter turnout's gonna change in this week versus another week because the weather's bad.

but then sometimes it's a structural thing, like you make the assumption in the past that people vote because they're patriotic, whereas now they vote primarily because of their

pocketbook, right?

And that affects turnout.

And so giving people the toolkit to figure out whether it's their modeling assumptions which are bad or whether it's the underlying parameterization uh is based on assumptions

which should evolve over time.

That's the core problem which I see come up in a lot of fields.

And that was the genesis of our thinking on this work in hydrology.

And so the approach that we took for that was was kind of similar to the earlier 2016 paper I mentioned, which was about in inverse modeling for rainfall, where you you see

using a a model for um how rain becomes river flow, um trying to recover the rainfall, only observing the river.

And instead doing something else, which is you assume that the parameters of the model, so your inputs are fixed, you know that the rainfall is a certain amount, but you assume that

the parameters are fixed.

And in this case, one of the parameters for the model was the porosity.

Which is how much water can your soil absorb uh absorb before it starts overflowing?

And this actually does change over time in watersheds for a variety of reasons.

Um, if it's compacted, so people drive over a field, their tires will compress the earth and then it will not be very porous.

If it's broken up or if plants are freely growing, their roots will um induce more pores.

If animals are burrowing in it, they'll increase the porosity.

So it changes for a variety of reasons.

but that use case isn't

especially permanent here.

The main point is that you take this model and then you you look at its parameterization, you say, I'm not going to assume just one fixed value of this parameter, the porosity,

soil porosity.

I'm going to assume it varies over time.

Now, this seems like a small assumption to make, but in practice it means that what once was one scalar parameter you estimate all of a sudden has potential to become some

something which is many, many thousands um of elements large in dimension.

And so if we let the parameter vary, say every week for

30 years, I think it was a the use case was kind of similar to that again.

Then you have again a high dimensional estimation pro problem.

Now, when you're working with these problems where the quantity you're trying to estimate, in this case the parameter of a simple rainfall runoff model in hydrology, when that's

high dimensional, then a lot of the conventional methods you have for estimating it, like uh essentially guess and check sampling, simplest sampling doesn't work.

You need something which understands the high dimensional geometry of the post of the uh inference problem.

And that's why, um

we're using Hamiltonian, Monte Carlo, and PyMC to do it.

And for that to work, you need to have the gradient of the model likelihood, or more appropriately, the log likelihood, with regard to the parameters you care about.

And so if our log likelihood here is just one up, it's a scalar, okay, that's fine.

But then the input value that we're calculating the gradient with respect to, that might be 30,000 dimensions, or in this case 10,000 dimensions.

And so then you have to calculate that.

If you look at the methods that we had before automatic differentiation, which is what Pi Mc and similar frameworks are built on, um those generally scale uh linearly in terms of

each iteration.

They scale as linearly with the dimension of the problem to calculate the gradient.

And they can also be horribly unstable.

So these are things like symbolic differentiation and numerical differentiation.

And automatic differentiation gets around that because there's some clever tricks in how you keep running tallies of the gradients as you look at the function you're trying to

differentiate.

Um and so as long as you have those gradients, which in this case means that your model parameters connect to the log likelihood via this intermediate physics models, that means

that to get the gradient, you have to backpropagate those gradients through the physics model to get them in the first place, application of the chain rule.

But then once you have that, then the inference problem of figuring out what that high dimensional problem, what those parameters were in the first place, is much, much easier

because the methods using that gradient can navigate through that high dimensional space much more efficiently because they know which way is downhill.

And that's at a high level why we did that work um with that specific case study.

Damn.

Yeah.

super interesting.

I really encourage listeners to to go into the show notes and and look at that paper because it's very interesting.

Something I'm really curious to hear about you about though is looking ahead.

You know what?

You're a very curious person.

You're always learning, you're always trying new things.

So what are you excited to learn and work on in the next few months?

Yeah.

that's a great question.

Um, I think uh in terms of for my own personal curiosity, doing it just seeing how easy it is, how simple you can make it, to do these large modeling efforts and how fast you can

get it.

For example, to make an analogy, some people like to follow the progress for the community on speed running LM training to try to train like a GPT three style model as fast as you

can.

And I want to do the same thing for a large Bayesian model.

See how big you can make it and how fast it can be trained.

And I think there's so much low hanging fruit there.

In terms of either optimization or modeling forms or simply just choosing the right compute provider.

The other thing is that we're getting we're getting reams of information now about what machine intelligence does, how it thinks.

I think that there's a lot of room to do statistical analysis and Bayesian work just simply on the outputs and the weights of the models.

uh

And for me, I think, you know, if I were a neuroscientist, now I have like actual neurons where I can pull every single one of them out and I can do models on the whole thing, on

the whole system.

Um and that's very, very closely related to mechanistic interoperability, which I don't work in, but I read false in the literature.

And that's also an area that I think is is something I'd like to spend more time in.

and also building out the toolkit for um Bayesian evaluation of like long-running LM processes.

So if you want an LM that's going to do something very, very complex and then modeling success.

That is an intrinsically rich data object.

And I'm thinking that there's so many models I'd like to apply to it to be able to understand like, you know, what is the utility of taking a certain action at this point?

You know, when do we see it going off the rails?

Can we model that?

Can we get signals that we can provide up to research for model training?

And doing things in that vein as well.

That's what most excites me, I think, over the next, you know, six months, year, two years.

Well, awesome.

Chris, I think it's it's time to to call it a show, of course.

I'm gonna ask you the last two questions.

Ask every guest.

At the end of the show, right?

You you know it, so I guess you're prepared.

So first one, if you had unlimited time and resources, which problem would you try to solve?

That's a great question.

Um there was a blog post I had, I think maybe a year or two ago, and it might seem a little out there at first glance, but basically the idea w was inspired by something I

read and it was more physics flavored or engineering flavored and and the basic idea was um if you look at broadly the problem of climate change, one of the biggest problems it's

facing us right now.

like what are the approaches that actually work to move the needle?

And think you know, I have a strong interest in biology and ecology, and I think, you know, the the most advanced machinery on earth is biological machinery by far.

Like how would you harness biology to d deal with this problem?

And, you know, a simple way of looking at this is okay, you you know the carbon change is primarily a function of c of carbon in the atmosphere, and what is the most efficient way

to carbon out of atmosphere?

It's it's plants.

But then how do you keep those plants from returning it to the atmosphere when they die?

And there is a paper in the Nash proceedings of the National Academy of Sciences by a fellow from Berkeley about five years ago where he proposed you take all the plants and

you put them in the desert.

You dried them out.

And they never released their carbon back to the atmosphere.

And um I was dissatisfied with that for a number of reasons.

And so I wrote a short analysis saying, okay, if you can take all the plants, you put them in high latitudes, you freeze them essentially.

There's clever ways you can manipulate the physics of the situation so that they essentially freeze and they never freeze out during the wintertime.

Up to some loss fraction, probably less than one percent.

And then that's a vehicle to sequester carbon indefinitely so long as the so long as it stays cold at the high latitudes.

And that's not a problem I'm really equipped to handle on my own.

And especially right now, it's not my job.

But if I had infinite time and money, that's probably what I'd want to work on.

Just seeing if that would work.

Yeah.

Yeah.

Um I've added that uh blog post to the show notes for people who are curious.

I think it's it's an interesting idea.

and if you could have dinner.

with any great scientific mind, dead, alive, or fictional, who would he be?

Uh probably von Neumann.

So aside from being preeminent mathematician of his generation, uh, you know, he's also known for just having a good time.

And, you know, I often like wish that I could get away with the things he did.

He was known for like reading books while he was driving and there's a famous thought, I don't know if it's true that he was paid for the time he spent thinking on things when he

was shaving or brushing his teeth.

And um

You know, just being so engaged in the work that's just part of your life and you can't get away from it, even for things that most sane people would think you really should

focus on.

I bet he'd be a lot of fun to have dinner with and maybe a few drinks afterwards.

Yeah.

Yeah.

That sounds uh that sounds like a very interesting dinner.

Very very fun.

I I'm definitely down for that one.

Um Well, Chris, fantastic.

We're um we're on time and um I was

Super happy to to have you on the show.

We covered so much ground and you're doing uh very very interesting work, so please continue and well I'll be I'll be very happy to to welcome you back here on the show when

you have new things to share.

Thank you so much for taking the time and being on this show.

I was thrilled to be on and if folks want to contact me to discuss this related things, please reach out.

This has been another episode of Learning Bayesian Statistics.

Be sure to rate, review, and follow the show on your favorite podcatcher and visit learnbayesstats.com for more resources about today's topics, as well as access to more

episodes to help you reach true Bayesian state of mind.

That's LearnBayesStats.com.

Our theme music is Good Bayesian by Baba Brinkman.

feat. MC Lars and Mega Ran.

Check out his awesome work at bababrinkman.com.

I'm your host.

Alexandre.

You can follow me on Twitter at Alex underscore Andorra like the country.

You can support the show and unlock exclusive benefits by visiting patreon.com slash learnbayesstats.

Thank you so much for listening and for your support.

You're truly a good baby and change your predictions after taking information.

And if you're thinking I'll be less than amazing, let's adjust those expectations.

Let me show you how to be a good daisy.

Change calculations after taking fresh data.

Those predictions that your brain is making.

Let's get them on a solid foundation.

Key Takeaways

In mining and geostatistics, the classic Gaussian process model, known there as kriging, assumes you know exactly where each sample was taken. Chris’ project broke that assumption on purpose: the recorded coordinates for each core sample were only accurate to within a rough radius. By treating the true locations as latent variables and putting a Gaussian process over them jointly with the measurements, the model could still reconstruct the underlying gold-concentration field, even though the exact sampling locations were never known precisely. It's a demonstration that Gaussian processes can absorb structural uncertainty that looks, at first glance, like it should make the problem impossible.

Poverty Bayes was Chris’ experiment in seeing how cheaply a large Bayesian model could be trained using modern cloud infrastructure. He fit a hierarchical logistic regression with close to two million parameters, using PyMC's Hamiltonian Monte Carlo on a single A100 GPU rented through Modal, a serverless platform that deploys a Python script straight to GPU hardware with almost no setup. He'd originally guessed it would cost around five dollars, the price of a Big Mac, but the real bill came in an order of magnitude lower. A model that would take a Gibbs sampler weeks to run, and that once required a research lab's dedicated GPU, now costs pocket change and a few minutes of setup.

Chris argues the software has largely caught up: PyMC's JAX backend and NumPyro make GPU-accelerated Bayesian modeling work out of the box for most problems. What's missing is common knowledge. Companies are clearly running large Bayesian models in production, but the results stay behind corporate firewalls. Chris’ proposal is a community benchmark effort: which frameworks handle a million-parameter Markov random field on a given GPU out of the box, since this kind of expensive, slow-running benchmark is a poor fit for standard CI pipelines but valuable for the field to know.

Techniques like speculative decoding, which uses rejection sampling to match a small model's output distribution to a larger one, make immediate sense if you already think in terms of priors and likelihoods. That training is also useful for reasoning about the high-dimensional geometry of modern embedding spaces, and for building hierarchical models to evaluate agentic workflows, where a single interaction now involves a whole sequence of tool calls rather than one chat output, without a corresponding increase in the amount of evaluation data available.

Chris pushes back on the idea that statisticians are being sidelined by the LLM era. A large language model is fundamentally a statistical paradigm, trained against a likelihood with sampling algorithms doing the heavy lifting, which makes it a vindication of the field rather than a threat to it. He's optimistic that the massive investment in compute for machine learning keeps spilling over into better tools, faster hardware, and deeper insight into how high-dimensional models behave for Bayesian statisticians too, comparing the current moment to the golden age that followed the advent of MCMC and Hamiltonian Monte Carlo.

Related Episodes