#165 Hierarchical Sequential Sampling Modeling, with Alex Fengler
Alex Fengler is a postdoc in Michael Frank's lab at Brown University, and the lead architect of HSSM, hierarchical sequential sampling models, a toolbox that grew out of the older HDDM (hierarchical drift diffusion model) project and rebuilt it from the ground up around neural networks trained to approximate likelihoods directly from simulation.
The thread running through the conversation is a specific bet: rather than training a network to spit out a posterior directly, the way most simulation-based inference tools do, HSSM amortizes the likelihood itself. That choice costs more upfront and pays off differently, and understanding why is the fastest way into what makes HSSM, and the ecosystem growing around it, worth paying attention to.
The canonical cognitive science paradigm here is simple: a subject watches a field of moving dots and decides whether the dominant motion is up or down, over and over, and you record their choice and reaction time.
The standard model for this, the basic drift diffusion model, treats each decision as a random walk that accumulates evidence until it crosses one of two boundaries. A handful of parameters (the distance between those boundaries, a starting bias, an underlying drift rate) generate the whole distribution of choices and reaction times, and because this particular model has a closed-form likelihood, it's shown up in thousands of published papers.
That's the thing Alex set out to change: models with a convenient closed-form solution get adopted en masse, while models that are theoretically just as interesting, and just as easy to simulate from, get almost no uptake once their likelihood function turns analytically hairy. Choice end up being about which model is cheaper to compute, not about which one serves the science better.
That's where likelihood approximation networks (LAN) enter the scene, by learning the likelihood! You feed the network a set of parameters and a trial's outcome, and it learns to output how likely that outcome was, purely from repeated simulation. Once you have that trial-by-trial likelihood as a reusable plug-in, you can stack it into essentially any downstream Bayesian model, add a hierarchical prior, add a regression backend on any parameter, without ever retraining the network itself.
This is a different bet than the one most simulation-based inference tools make, which target the posterior directly. Actually, this reminds you of BayesFlow, doesn't it? And I can see why: BayesFlow's networks do give you near-instant inference once trained, but the network is scoped to the exact setting it was trained for: change the data size or the prior and you need to retrain.
Amortizing the likelihood instead costs more upfront, but the resulting network stays reusable across arbitrarily different models built on top of it.
The ecosystem Alex built around HSSM is a pretty neat piece of interwoven software: you've got ssm-simulators to generate fast, ready-to-train data from these cognitive process models, LAN Factory to train the likelihood networks themselves, and, to finish, HSSM itself as the inference layer, built on Bambi and PyMC, sampling via NumPyro, nutpie, or BlackJax depending on what you want.
Alex demonstrated that workflow live in the episode, so I recommend tuning in for that part, but in a nutshell, he trained a neural ratio estimator using BayesFlow and dropped it into an HSSM model's log-likelihood slot with no glue code beyond that conversion.
The only real requirement for amortizing a likelihood to pay off is a simulator fast enough to make the upfront training worthwhile, which describes plenty of problems beyong cognitive science, like computational biology or physics, both fields already leaning on simulation-based inference.
Alex is also pushing the model catalog further within it: diffusion models whose drift is informed live by eye-movement data, reinforcement-learning backends on any drift-diffusion parameter, and hidden Markov model backends on the parameters themselves to catch regime switches like attention lapsing mid-experiment. All pretty exciting, isn't it??
One last cool thing Alex is working on: Bayesify, built with Jerry Huang and friend-of-the-show Stefan Radev. Upload a paper, and Bayesify checks whether it's following [Gelman et al.'s Bayesian Workflow]()! Told you it was cool ;)
Check out the full episode above, and the show notes for links to the HSSM docs, the LAN and HSSM papers, Bayesify, and the related episodes on amortized inference.
You can also interact with the episode on NotebookLM! Ask questions, generate flashcards, and more.
Hope you enjoyed it, and see you in two weeks, my dear Bayesians!
00:00 What is HSSM and how does it fit into the Bayesian inference landscape?
12:09 How did HSSM evolve from HDDM, and what does it apply to?
30:25 How do neural networks learn likelihoods for Bayesian inference?
37:01 What makes amortized Bayesian inference so flexible?
41:04 What are the real computational costs of amortized inference?
55:16 How does HSSM integrate with libraries like BayesFlow?
58:57 What does a live demo of HSSM and BayesFlow look like?
01:18:33 What is Bayesify and how does it score a paper's Bayesian workflow?
01:23:12 What new model classes are coming to the HSSM ecosystem?
01:30:12 How is AI reshaping development in the HSSM ecosystem?
01:38:42 How should society incentivize keeping hard cognitive skills alive?
My guest today is Alex Fengler, not only a friend of mine, but most importantly, a brilliant postdoc in Michael Frank's lab at Brown University and the lead developer behind
HSSM, a toolbox for hierarchical sequential sampling models that's become one of the most interesting corners of simulation-based inference.
HSSM is built around neural networks.
That learn the likelihood of a cognitive process directly from simulation.
Instead of chasing a closed-form solution every time someone processes a small variation on a decision-making model, you train one likelihood network and reuse it for practically any
downstream Bayesian model you want to build on top.
We get into what these likelihood approximation networks actually learn, why Alex bet on amortizing the likelihood rather.
And the posterior, and how HSSM stays agnostic to where a trained network came from.
Alex even shows live how a network trained with BayesFlow slots straight into an HSSM model.
We also talk about two things Alex and his collaborators are building for the community: a personalized learning tool to onboard new contributors and basifying.
A tool that scores whether a research paper.
Actually followed a proper Bayesian workflow.
I'm sure you're gonna love that.
This is Learning Bayesian Statistics, episode 165, recorded July 1st, 2026.
Let me show you how to be a good basian change your predictions after taking information.
And if you think it now be less than amazing, let's adjust those expectations.
What's a Bayesian is someone who cares about evidence.
Welcome to Learning Bayesian Statistics, a podcast about Bayesian inference, the methods, the projects, and the people who make it possible.
I'm your host, Alex Andora.
You can follow me on Twitter at Alex underscore Andora, like the country, for any info about the show.
LearnBaseStats.com is Lab Plus2B.
Show notes, becoming a corporate sponsor, unlocking Bayesian merch, supporting the show on Patreon, everything is in there.
That's LearnBaseStats.com.
If you're interested in one-on-one mentorship, online courses, or statistical consulting, feel free to reach out and book a call at topmate.io slash alex underscore andora.
See you around, folks, and best Bayesian wishes to you all.
Alex Fengler, willkommen nah Learning Bayesian Statistics.
Honored to be here.
Mm-hmm.
Is it na or zur?
Learning Bayesian Statistics.
I always I'm I'm always confused.
Of dem Podcast?
To be honest with you, I've never said it in German.
I Zum Podcast.
Maybe Zum yeah.
Also okay.
Yeah.
Like prepositions in German are just like it's always been my nightmare.
I think even worse than uh you know, like according agreement of the of the names and so on.
It's just like prepositions are just terrible for me.
I guess it's the same in French, you know.
Yeah.
Yeah, it's the same in French.
It's like, you know, why is that feminine?
Why is that masculine?
But I don't know.
That uh the genderiness is horrible actually.
It took me like like learning other languages and seeing people struggle with it.
Yeah.
And especially when people ask you to explain why and you constantly come up empty is is also very embarrassing.
Yeah, yeah.
I mean there is I think I guess it's just random, you know.
And and it's terrible because e even between the Latin languages, um it's just like it depends.
Uh something will be feminine in French and masculine in Italian and then
Masculine in Spanish is you never know.
So yeah, let's not enter that that rabbit hole.
let's talk about you, actually.
I think that's what you're that's why you're on the show.
That's why I invited you.
Um yeah.
So yeah, I'm very happy to have you on because it's been long overdue episode, but finally we we made it.
Uh ironically you are not in the US right now, so it's not the the easiest the easiest timing uh that we did, but you know, we made it.
And so yeah, to start, um let's let's start with your origin story, you know, um what are you doing nowadays and how did you end up doing that?
Yeah, my origin story.
Uh I come from a smaller town close to Cologne and Germany.
that's also where I did high school and everything.
And I started studying in Maastricht, actually, which is uh it's a city in the Netherlands which is very close to the German border.
And um how I ended up there is uh one of the you know, that's where probability starts.
Like that was one of those dice rolls in my life.
Um I actually meant to go study machine um mechanical engineering actually in Aachen.
Which is a city that is also close to the Dutch border.
and my girlfriend at the time had a Dutch passport and she was really keen on also checking out some Dutch universities and she really wanted to study psychology there.
And I ended up going over with her, looked at the open day at Maastricht University and ended up getting stuck and then I just enrolled in a business degree there and got going.
Um
And then eventually I ended up doing a master's in something which was called neuroeconomics.
It turns out that Maastricht is actually one of the Maastricht Universities is one of the origin universities of that entire discipline.
There was a there was a course coordinator, Arno Riedel, an economist, who was really keen to um bridge the gap with neuroscience and psychology.
And I did that for a while.
So that's sort of mixed in in business I was like focused on uh finance actually.
Um and then I started I just wanted to sort of branch out, um, some sort of like I don't know, partly regrettable, partly nice yearning for other horizons, right?
Um yeah, and then through neuroeconomics I ended up doing getting into like studying let's say micro level decision making is how you would talk about it maybe from the equine
perspective.
And uh from the psychology perspective it's more in the canon of um you know, studying individuals' decision making on simple choice scenarios.
And these simple choice scenarios from for what's interesting for economics were then about preference based choices.
And somehow I started doing this, you do some microeconomic theory, this and that, and I ended up at uh Caltech, one of the pioneers of um that discipline.
is someone it's called Antonio Rangel.
And I ended up doing an RA ship there and write my master's thesis there.
Um and that was actually on a class an extension really of a class of models that then I'll later return back to, even today.
Um we were interested in how do people decide over snack items on a on a choice screen.
And it turns out that the one of the canonical paradigms was always about two alternative forced choice.
Um, so you get two items and you choose left or right.
We'll maybe see something later along those lines.
There are canonical experiments in psychology that do this.
And that professor applied a class of models that came out of cognitive science, switched it over to preference-based choices, and it ended in experimental paradigms that concerned
snack items.
So choose over Reese's and Haribo or whatever.
Um
And we ended up then extending that to more items on the screen and dealing with sort of the fi the process of fixations.
So the attention process across time until you reach a decision, right?
That was sort of the the game uh there.
And I ended up writing my master's thesis on this and you learn how to start applying to PhD programs in the US.
And it turns out that eventually my profile after all this was a good fit for
Cognitive science departments.
And I applied to to Brown.
Um again, I wanted to uh branch out, so I wanted to kind of leave all that behind and move into something which is called Bayesian non-parametrics.
So that professor was actually um Joseph Ostov was his name.
That professor was really interested in human category learning and was trying to use um Bayesian non-parametric approaches for that.
Um
And eventually he left Brown one year into my PhD.
And um at that point uh I got really hooked on um sort of learning much more about statistical methodology back then.
And on the occasion of him leaving the university, I then decided to actually do a detour.
So I went to Bocconi in Italy, went a PhD and stats there.
I was like fairly underprepared for that, uh especially in hindsight.
And um took out the master from that degree and uh was then able to continue uh my degree at Brown.
And that's under a different prof.
Um his name is Michael Frank, that's also still the the principal investigator in my lab now.
Um and under him I ended up drifting towards simulation based inference and applying simulation based inference to
variants of these models that I had started um studying during my masters.
Um he was always at Brown and he was actually always a big player in that discipline, but I didn't even know that when I applied to to Brown.
And eventually we found uh our path then converged and I was able to graduate under him uh at Brown and that's where I'm still now uh postdoc.
So much of my PhD research was finally some intersection of neural networks.
and um likelihood free inference, which today people uh really call it simulation based inference.
Um ABC, approximate Bayesian computation was an acronym that was used in the past much more commonly.
and we started building, you know, some sort of software ecosystem that then became the centerpiece of my postdoc, which is what I'm still doing.
Yeah, that's that's uh that's a very interesting um
Background.
I didn't know you had so many different experiences in in Europe first actually.
I didn't know you had lived a bit in in in Milan and uh if I if I infer it correctly.
So that's cool.
Eighteen months, yeah.
Yeah, yeah, yeah.
That's a nice city.
I've I've lived there also for a few months.
Really, really loved it.
Really cool.
Yeah.
No, I was actually it was it was great.
It was partly because of um
The prospect for um the job market after after graduation that I decided to go back to the university.
Yeah, yeah.
I mean that unfortunately that makes sense.
Yeah.
and so so yeah, now um you're working a lot on on on these topics still.
what's what's one of the the main things you have in mind right now that's that's you know taking
Most of your working days.
So in terms of the the academic research, um right now we're we're pretty much at sort of like a pivotal point.
So we just um brought to preprint uh our research on an on a software ecosystem, H S is what's called.
So it's yeah.
Yeah.
Um Yeah, nice t shirt.
So yeah, maybe talk to us a bit more about that.
What what does even H S mean?
Um what
Yeah, what what is that about and why and when would it be useful?
So HSSM is for hierarchical sequential sampling models, which is generalization and name to its predecessor toolbox, which is called HDDM for hierarchical drift diffusion models.
and why we can go from hierarchical drift diffusion models to hierarchical sequential sampling models is because
The HSSM toolbox is like from the get-go powered uh by deep integration with uh simulation based inference approaches.
I can show some of this.
Uh I have a I have a short slide deck.
We can either do this now or later.
Yeah, no, for sure.
Yeah, yeah.
Yeah, I should register that now.
Yeah, I mean if if that helps the the explanation for sure, uh feel free to to share a screen and hopefully at least wrong.
Um so just to introduce this uh
The the underlying like canonical paradigm real quick.
So I hope you can see these moving dots here on the screen.
Yeah, it's perfect.
This kind of uh screen here uh refers to a very canonical decision-making paradigm in cognitive science.
It's called uh random dots motion task.
Ultimately as a as a subject in such an experiment, you would decide if the dominant direction of motion is up or down, or in some you know, left or right, for example.
And
This particular choice screen is very hard.
So this is a very random direction, so it's very hard to decide.
But the point is that you would uh this is like a psychophysics experiment.
So uh people would do this like very many times, and you collect reaction times and choices.
And a very canonical modeling paradigm that has been applied to this rests on this kind of model here.
So this is called the the basic drift diffusion model.
By the way, um you cut out for me, so no, I'm here.
It's just like since I'm not uh talking, I uh I just have you in the uh in the video so that people can focus mostly on your uh on your presentation.
Because you were you weren't moving, so I started to start to get bored.
Um so what's the strift diffusion model?
Um so you basically assume that choices and reaction times
When people decide, for example, for such um motion directions, they derive from such a random walk process that eventually crosses a boundary.
And then such boundary crossings decide which choice was taken and when it was taken.
Right?
That's like the conceptual mechanistic underlying model here.
And these models have a few parameters.
So here I'm using a very simple one that has a parameter A for the the distance between these bounds.
That's a measure of caution.
How much evidence do I accumulate basically until I take a decision?
An a priori bias.
Where does this random walk start here?
This is not always relevant, but depending on the experiment it can be.
And then an underlying drift.
So for example, if um the random dot so if the choice screen is a little less random than what I show what I've shown here, then you would assume that people over time
accumulate some um fixed portion of evidence that leads them towards one of those choices.
So if it's very clear which direction the um these dots are moving, then you would have a stronger drift.
So you would with more accuracy take faster decisions because it's uh it's easier.
So I'll have here like one example.
Um
So here I'm not sure if that comes through on the stream, but um these are three three such conditions that um get a bit easier from left to right.
So you can see on the right side that it's it's much easier now to detect that the random um that the motion direction tends to cluster towards down.
And usually the way this is then modeled is that um you assume that this drift rate is affected and people will be able to take choice down, right?
more often and faster.
There's there's a whole conceptual background here of um thinking about speed accuracy trade-offs in in decision making.
But for for purposes of this pr um podcast here, this is just I'm just trying to introduce a modeling paradigm to get to the the actual crux, which is that
If a friend of ours comes around and says, That motto is nice, but honestly, I think that people will actually not have a constant evidence criterion.
They are, for example, impatient usually when taking such kind of choices.
So as time moves on, they may be willing to decide on much less evidence to get get it over with, right?
Get the trial over with.
You can motivate this from uh many different directions.
But the point is that if you go from this model to this model, this is just a placeholder example, there are many such variations.
What can happen is that the the mathematics um for you to compute what are called first passage distributions also, these these um histograms here.
So the first time the distribution over exit um
Times.
Mm-hmm.
These distributions, they can in this particular case, they have uh an easy solution.
So you can have a so-called likelihood function.
So if if I give you parameters and a data point, which is which which boundary was crossed and when, you can tell me how likely it was under the parameters.
Right.
For this model, there exists some closed form function.
For this model,
In this very particular case it still exists, but that's beside the point.
For any type of uh variation of these models, which is which are as simple to simulate from, simulate data from as it may be for for the canonical one, the mathematics to
describe these um likelihood functions can become really hairy.
Right.
And this distinction between how easy it may be.
To generate data from these processes and how hard it may be to evaluate likelihood functions, which is really what you what you need for any sort of standard um Bayesian
paradigm for parameter inference, for that matter, also maximum likelihood estimation.
This distinction here, so if I have a very small variation of a canonical model that it suddenly can become way harder to do parameter inference, is sort of the
original motivation of the the research agenda that led to something like uh HSSM.
So there are two key observations.
This is really from you know a few years ago when we started working on this.
Yeah.
Observation one is that these these models that have mathematical shortcuts so that you can actually conveniently do likelihood based inference are very sparse in the space of
models that people might be interested in or would like to propose.
And observation two is that that has an outsized impact or it's uh it had in um as far as far as our observation goes in cognitive science at least, it has a completely outsized
impact on how many downstream publications by experimentalists will make use of a particular model.
Because of the analytical convenience over the actual theoretical interest in a particular model.
So this simple model here was like it has literally uh if you're outside of cognitive science, this might be shocking, right?
But inside cognitive science, this is just such a massive uh paradigm that is applied.
Um thousands and thousands of papers um have been published using this.
And there's
Lots of software infrastructure that helps you do parameter inference for these types of models.
It's really distributed.
Uh one there's uh there are things in Python, in R, in MATLAB, right?
But the software infrastructure, especially a couple years ago, even for the simple variations, was basically non-existent.
And the kinds of experimental papers um that dealt with such variations on on these models.
were very, very sparse, very, very few publications, even though there were quite a few theoretical papers that clearly established uh interest in those models.
And we through our eyes it looked like this was all just coming from the lack of analytical convenience.
So you needed to do much more specialised things that demanded a lot more from um even a computational experimentalist.
that just created a very sharp cliff between canonical paradigms and even slight variations.
And we started from here.
and one thing that you might uh recognize is that well, you know, just because I'm changing the boundary here, or I for example put a parameterized function on the drift, um
it's not necessarily the case that these first passage distributions
become highly complex as a result.
They may still be very regular, uh they behave very regularly.
It just might be hard to define a mathematical representation to compute them extremely quickly.
And that's really the key for you to to perform um likelihood based inference.
If you can't do that, your inference procedure is just going to take forever.
So the core idea was that okay, well, if things are regular, um just hard to represent, maybe we can find a way to um learn just by repeated simulation from the process, maybe we
can actually learn what the shape of these likelihoods look like.
And then we'll just have a neural network that takes in the parameters of the process and spits out log likelihoods um for particular observations.
That's the
That's the key here.
And um this is sort of the flip side of nowadays, this is just the flip side of what has um arisen as one of the dominant paradigms in simulation based inference, which is focused
on directly learning posteriors.
What we're doing here is um trying to learn the likelihood of a process so that you can layer have a surrogate function that you can just plug into base rule, right?
So what we finally have is these these networks here that we trained.
We call them LANs for likelihood approximation networks.
At this point, they are um different versions of this paradigm.
Um we may see some of them later.
And you can literally just go in, plug this network here at the point of for this podcast, the beloved base rule.
And you can if you for example
like collect a thousand data points from such an experiment as we had seen in the beginning.
Then you can just you have your parameters coming from a prior, add like a thousand data points here, batch the computation, right?
And then you just evaluate the network and if you have a larger GPU, you can batch more trials.
So you can eventually you can run this on like massive experimental data sets without too much trouble.
And one of the key things that which originally made us decide to go for um approximating likelihoods over posteriors.
In the meanwhile, the research landscape of course also has evolved.
But um especially when we developed this, the flexibility of posterior based methods was much less.
So one of the things you get for free when you focus on learning the likelihood is that any type of um
hierarchical superstructure or choice of prior, etcetera, etcetera.
Is a post hoc decision where you just reuse your surrogate likelihood, for example, to build a completely different model.
So here I just um signified in in sort of magenta.
Um for example, if you want to have a hierarchical model that has a group distribution over parameters instead of just a basic prior likelihood model here, right?
Um you can just reuse the same network in here.
No problem.
Right.
So you never have to once you have a good version of such a network, then you can just reuse it downstream to also test all sorts of particular scientific um hypothesis about
your data set.
Right.
So that we we started sort of from this idea of what an experimentalist would usually try to do before like in terms of their modeling approach.
Right, until they bring something to publication.
And usually it's it's a large sequence of trying very many different kinds of model formulations that rest on some fundamental cognitive process model.
So we kind of amortize the fundamental cognitive process model, but then people can play around with all kinds of um, for example, regression backends that um allow them to add in
neural covariates, for example.
things that are also collected during experiments in the broadest sense.
Any kind of covariate that you might collect, eye tracking, skin conductance, there are all kinds of other ducal dilation, there are all kinds of other covariates of interest for
people.
And you can just put a regression back end for example on on a parameter and build build a bigger model around it and just keep reusing the same network.
So that was kind of the core idea here.
That kind of summarizes my PhD research really.
And then from here we said, okay, so if you
Can now ship these networks to some let's say some sort of centralized database, right?
And you have an ecosystem that makes it easy for people to go from I think I have a great idea about a variation of a model to train a network and then build a model around it,
right?
That that should sort of democratize this entire process here.
Um to really free experimentalists to be much more aggressive in the kinds of hypotheses that they uh can test.
on their empirical data.
And if you you know, if you then have a process to quote unquote productionise the likelihood networks, then the entire community can basically benefit from people testing
things, uploading them, and immediately during the ecosystem, immediately uh the entire community can start testing a new proposed model, right?
So that's eventually the core of my uh postdoc activity was to see this sort of bring this thing here to to completion.
And that's at this point like a it's a collection of three smaller uh Python packages.
One is to create fast simulators, especially for these variations of cognitive process models.
And one is a small neural network library that is then hooked up to to hugging phase and
HSSM, the the name on my chest is really the in this ecosystem context is the library that then focuses on making inference convenient and um
catering somewhat specifically to the needs of people that are doing cognitive process modeling, even though it is a priori conceptualized to be even more general.
But we're scoping this around, making it very nice to work with variations of these um sequential sampling models.
So that's like the this slide here is sort of my postdoc basically, wrapping this up into um into something that can be
very conveniently used by the community.
Yeah, yeah.
Yeah, for sure, damn.
Well done and all of these, that's that's really amazing.
Um before you continue, uh yeah, I have a few questions which is okay, so it sounds like basically having the neural network the if I understood correctly, the neural network in
there is used to learn the likelihood.
Is that correct?
Yeah.
So here we you would even you would say trial by trial likelihood.
So really um what makes this nice is that once I have the to find my mouse.
Um once I have the single trial likelihood represented as a neural net, then I can just start like stacking basically.
Right.
So just keep stacking data points or
I can also represent of course the parameters trial wise and put any kind of like modeling backend on it, right?
So if you maybe later uh we'll see you can put a process on the parameters connected to other modeling paradigms, etcetera.
Like what by the time you have the the single trial likelihood, then you can just flexibly compose it with anything.
Hmm.
Okay, yeah, yeah.
Yeah.
So basically the neural network here is is super flexible by definition, so it helps you learn that kind of
of of problem.
Um so I can definitely see the connection with uh amortized Bayesian inference.
So yeah, definitely understand why you're also working with uh Stefan Radev and all the good folks um on the BayesFlow side.
I'll put this eps these episodes in the in the related episodes of the show notes.
for sure also um feel free to add the
The links to at least these three packages and any of the papers or tutorials or things like that you think are are gonna be useful to people.
Um one of the main and the main questions, follow up questions I have net is what is the actual difference with amortized Bayesian inference, which I think now people on the show
are familiar with.
And also uh the difference with simulation based inference, so SBI, which Jonah Saruda came on the show to explain also.
And uh another one would be related to computational cost of these methods, because while you have to train a normal network, um you said that uh that uh lens
the the Python pack Python packet is is talking to Hugging Face, so yeah I'm guessing that that's uh some costly training pro shader in procedures in there.
Maybe not.
Yeah what's the what's the lay of the land here?
How how does it work?
And then I'll have further questions on the HSSM part in particular.
But let's start with that first.
Yeah.
So amortized
Bayesian inference.
So maybe yeah, we can try we can try to decompose the words, right?
So in some sense, what um so simulation-based inference um I think is on some sense the broader term.
the key is that you start from just having access to a simulator, but you do not have access to a likelihood function.
And to a degree
You may have access to a likelihood function, but maybe that's very expensive even to compute because it's not represented in a computational convenient way, right?
But the the key simulation based inference always just starts from someone gives you a function that you can evaluate, right?
That is a simulator.
So if you give me parameters, I can spit you out data, right?
Um
That's just the the the beginning setting.
And then
From there you can you can take many different routes.
So the the key is just I start from a simulator and in the end I want to have posteriors, however I get them.
And um one approach is to use the simulator and try to um directly target the posterior with the neural network.
There are historical approaches that um they don't go through any machine learning route.
There are some famous like early
simulation based inference back then it was really called approximate basion uh computation, ABC is what it was called back then.
there's some famous algorithms that sort of show how you can, you know, directly work with uh forward simulation inside base rule.
Some actually very simple.
There's a very famous rejection sampler.
I really suggest people to to look into that just for the fun of it.
It's kind of magical that it would work.
Um but the key is that these earlier approaches
They use simulation directly at the point of doing inference and that's what so you you keep the expensive part.
You bridge the gap from having only simulations to getting posteriors, but all these methods become super expensive when you when you're trying to do anything uh more
complicated, in particular hierarchical inference or all that kind of stuff.
Because the fundamentally you're bottlenecked by how unlikely real data is under your simulator.
And those early approaches in particular were very bottlenecked by that.
So in some sense, like if you didn't have a great model for your data, then systematically your inference procedure would also become extremely slow or you trade off with accuracy.
Um now eventually machine learning came in and many different kinds of like amortization schemes were built.
Um and it's you can
you know, in a in a simple way you can really branch it between what does the neural network target?
Is the neural network trying to go from uh data to a posterior?
Or is the neural network trying to go from uh is the is the neural network targeting the likelihood of the simulator?
And today, I think both approaches have merit.
Obviously in in our uh group we and like I could do some sort of forensics on this.
what motivated us to do it this way is because we are really in a we're coming from a group that is pretty heavy also directly on the experimentation side.
So we were pretty aware of how experimenters usually go about their data analysis.
Right.
So we were quite motivated by this a priori.
And then you put a lot of emphasis on reusing the same nucleus across many different types of um finally sort of hierarchical models or
I tend to call these like Bayesian superstructures, right?
Like any kind of like particular hypothesis you want to test, etcetera.
You do you didn't necessarily want to retrain your uh model for every one of those occasions.
So we were willing to pay a large price to train the likelihood to then make it feasible to do MCMC, right?
So we are sort we are we are amortizing the likelihood and then um inference goes back to all our normal uh machinery, right?
So we can do inference, we are not basic MCMC, right?
If you want variational inference, don't stream.
Um and you're able to just test a lot of different things while having amortized like one thing, right?
But we're willing to pay a pretty hefty price to amortize that one thing, right?
And the posterior route is quite motivated by having nearly instant inference by the time you have have it amortized.
But then especially initially, it means that you're amortizing a very specific scenario.
Right.
So the likelihood route gives you basically all the flexibility downstream for free, but you pay a larger price per inference.
The uh posterior amortization route um makes inference instant but locked you into particular scenarios, otherwise you had to retrain.
Right.
And uh
Sort of in the history of things now, of course, the the people that were focused on posterior amortization have tried to generalize that more and more across dimensions.
Like how many you want to be flexible with respect to the amount of data you have, with respect to the priors you choose, et cetera, et cetera.
Right.
And there's also, you know, there are some major successes on that front.
We just we are just flexible a priori.
So downstream flexibility is just not a problem once you have once you have light yields.
But inference can still take time.
So from the simulation based inference terminology, uh both um focusing on likelihoods or focusing on um posteriors uh essentially within terminology because the the key is that
you're just starting from a simulator.
Now if we talk about amortized Bayesian inference, that term I think is
personally think, right?
People use it the way they want, but that's like more tied to the posterior route.
So you're trying to amortize the inference uh step and you can even do that for models for which you a priori have likelihoods, right?
You can do that.
You can just still train and get instant inference.
As we you know, know from any of our applied work, even if you have um likelihoods doesn't mean that inference is gonna be fast.
via MCMC, right?
So um if you can amortize the inference step, that's kind of irrespective of whether you operate have likelihoods or not, you might sometimes benefit a lot from getting instant
inference downstream, right?
Where we're where you care about.
Yeah.
But the the key concept for me is the reusability.
So um you're kind of in general willing to pay a large price for the initial training.
If you can find ways to reuse the same object many times.
Yes, okay.
Yeah, yeah.
And
What about the the the concrete computational costs when people use these kind of methods?
What should they expect so that it gives them an idea of when it's worth it and when it might be too too much of a too much of a hassle for what they are looking for?
So one is making things possible at all.
There's always this um
Without any of these neural network based approaches, certain scenarios are just not possible at all.
So those are less interesting because there's only one thing uh you you should really be doing here.
Um
Of the approach I have taken is really the pressure for reuse downstream is what motivates the the um willingness to to pay a price for a priori um amortization.
So we for example the the networks we are using for production, so the ones that are coming shipped with the ecosystem, in contrast to um the user's ability to train networks
ad hoc and play around with things.
Those are sort of an overkill on training.
And they're you know, you you run a ton of simulations and run them on a cluster and whatever, and we're doing this because we kind of don't even care.
Like it's just those things are there to be reused uh at infinitum.
So we don't really care that much at all.
Right.
But for example, um if you want to if you have a a little model proposal and
You want to for example check if the model is broadly identifiable, right?
Broadly identifiable over like tons and tons of data sets, right?
That's like for ex
I'm backtracking because I I'm trying to answer the question when you would go for amortization at all.
Um trying to make a distinction between that question and the question of when you would go for likelihoods and when you would go for posteriors.
Right, yeah.
If the question is just about whether or not to amortize at all, I really overall um for me the key concept there is just the the ability to reuse.
Like what for ex um now in that
Framing, what does for example reuse mean?
So if I have a new model that I'm proposing, right, and I want to know if that model is really broadly speaking identifiable, right?
So identifiability is a concept that kind of depends a bit, right?
Like for some parameter settings, you may be able to get clean postures for some other parameter settings, maybe trade offs emerge, right?
So
For you to like broadly investigate that for a given given model that you just came up with, right?
To then justify that that model is actually used for science downstream.
You really want to be able to do like in you wanna do inference many, many times.
Right.
That's the same thing for these simulation based calibration methods.
You would want to to be able to do um inference many, many times to corroborate that um
uh in your model or the network that you trained is reliable, right?
And here I'm focusing on the model itself, right?
So there if you know that you want to let's say run thousand, two thousand inference settings or more, really depending on the model, um that is one such reuse phenomenon.
So you you're proposing a new model and then someone else comes and says, well, should I really apply this model to by to my experimental data?
You wanna give them some sort of wholesale claim on, well, if you apply it to your experimental data, at least I can tell you that for a pretty broad range of uh parameter
settings, the model is at least identifiable.
So you can use it and you're not gonna have too many pitfalls, right?
Yeah.
And this initial validation, for example, um, where you just have a little simulator, that is it very quickly makes sense to amortize um
that into a neural net where the neural net is amortizing across a large parameter space.
And then you can just condition on data many, many different times, right?
And get posteriors and investigate where are the corners in the parameter space where my model is, for example, not identifiable.
Or another way of saying would be where pr severe parameter trade offs emerge.
So it's over parameterized in some spaces of the parameter space, for example.
So that's for me it's it's one clean example where in this case posterior amortization in particular is very useful because you'll be able to just do inference thousand, two
thousand if you want a million times downstream for free once you have the network.
And it really allows you very broad investigation of uh properties of your model as inference is concerned.
Yeah, yeah, yeah.
And so if we go back a bit to um to HS
SM What I'm wondering is that so it it sounds like it's basically a a layer of specification on top of the first two layers you talked about SS, MS and Lance, and so
this is gonna specialize a bit more the kind of models that people can do with this package.
however, it does sound like it's quite
January, right?
So you are you're focusing it on cognitive science, but it does sound like it's it could be useful to a broad range of application, doesn't it?
And if yes, well what kind of applications do you think these kind of of models would be uh would be very interesting for?
Yeah.
So it's true that um there's a really
Um nothing a priori tying us to cognitive science with the approach.
Um it is also the case that these kind of uh models are just a great test bed.
They're very useful, they're widely applied, um they they yield also good science, right?
Um so that as a as an anchor for our ecosystem, they made sense and it's a natural generalization of the um original HD DM toolbox.
Um
So there's partly a positioning and initial scoping aspect that goes into exactly what we did, right?
that's really dissociated from where these things can be useful.
And um in principle, I know, for example, I know that other disciplines that really care about simulation based inference today, right?
I know that computational biology in particular is a is a strong has a stronghold on simulation based inference.
And um
Quite a few papers that make use of simulation based inference are coming out of uh physics.
But really, from my perspective, the only thing you need is a model that is in the space of fast enough to simulate that it makes sense to amortize at all.
There are some simulators where you really have to rethink your approach a bit.
If the simulator is extremely expensive, then amortization itself
can become completely intractable.
So those kind of settings, they are of course less immediately attackable.
Right.
There are certain approaches that don't work with as much of a view towards global parameter space amortization and really specialized.
It's kind of like a earlier versions of this neural network approach were really trying to speed up inference for single data sets with some
amortization idea, but they didn't really have this ambition to globally amortize across the parameter space.
Um and that was really motivated by having very expensive simulators.
But there's a large sweet spot where simulators are reasonably fast, but still way too slow to use them live during inference and use these legacy um simulation based inference
um algorithms.
Uh-huh.
But they're really fast enough to allow you to um
train a neural network and get very, very large gains uh downstream, especially if it's for models that have interest um by a broader community.
And um
We are generalizing here towards all kinds of different classes of models in cognitive science.
but it's actually in fact um on my overall to do list to find more great examples from other disciplines that one can address.
and like let me give you one simple example.
Um the representation let me just go up here.
So for this basic um drift diffusion model, the actual representation of these likelihoods, which is uh used for fast computation of an analytical version here, is an
infinite series representation, right?
Where like the you have an algorithm that sort of decides how ma how much to truncate, etcetera.
And there are other models, um I don't want to give the wrong name now, but there are other um I'm not sure if it was
negative binomial or whatever.
Um I think not negative binomial to be fair.
Um that show up for example in media mix modeling where you have likelihoods that are just you you may not want them exact.
You just want a fast way to approximate them.
And that's another use case where you could go and say, well I have this like exact algorithm to compute a likelihood, but that goes through a pretty expensive process.
So maybe I can have a slightly um approximated version of that that ends up being faster to compute.
Sometimes even because it's easier once you have the neural net, given the entire computer infrastructure, to place it on a GPU and make use of massive parallelization, right?
And you can speed up um inference for models that even have just expensive to compute likelihoods, for example.
I'm aware that there are quite a few of those examples.
Um but we are growing out and our research agenda towards it and so far we are we are still staying in the uh cognitive science adjacent space to grow the toolbox and be useful
to like be clearly useful to a particular community.
Yeah.
And is that an ongoing effort?
Are you fielding for some uh people to come to you with with use cases and things like that?
Because if you are, I think this is one of the best places to to call for it.
Yeah.
So uh we are definitely thinking broadly about the database aspect.
So um having a database of amortized likelihoods and making those then also
harvestable through uh different ecosystems.
We are of course like sort of very naturally integrated with on the HSSM side.
But in general, like the the contribution pipeline to also contribute likelihoods and uh new networks, etc., something where uh we would really be generically speaking looking
forward to contributions.
And if one grocks the handshaking mechanism between HSSM and um
th this LAN factory package here.
you can use this for your purposes and could live completely outside um of cognitive science, that's for sure.
Um and in principle, that's also a broader point, we're really putting quite an effort into designing things so that eventually community contributions can take over uh the
majority of um contributions and pivot away from uh us upholding the the entire contribution.
landscape, right?
So we are really very actively working towards making it as easy as possible um to go from I have a random simulator to okay I can now do inference via HSSM.
Right.
And then hopefully um if we have a few anchor examples outside of the discipline, um
that would spur some innovation coming from outside.
Yeah.
Yeah.
Yeah.
You know how you know how it goes with uh open source, it's uh Yeah.
It can take a very long time and at random moments in time, uh suddenly there is maybe this is such a non random moment in time, but there um there is suddenly some interest
from people outside the field just because someone outside decided to dig, a pioneer from elsewhere.
They decided to dig and then suddenly that opens up a uh pocket of contributions.
Yeah.
Yeah, yeah.
No for sure.
And
How does that so how do your packages communicate with BayesFlow?
Is that even something that's in the works right now?
Because it sounds like there is synergies to be had.
So yeah, what's what's the lay of lam the land here?
So one of the
So what we're doing here with these likelihoods is um that's already an a priori um sort of generalization we had we had in mind is that this the LAN factory package itself is
designed to use PyTorch or JAX.
That that's a decision that pre like goes a goes quite a while back.
Um and we had always thought about okay, so how do I make sure that by the time things are reaching HSSM we are sort of agnostic
To the real origin of these networks, right?
And we have few mechanisms.
One is that JAX native passes through, no problem.
But we also have this intermediate layer, um, and our production system is really around that.
Um that's called Onyx, ONNX.
And that's basically a translation mechanism between uh these different uh deep learning or graph libraries, right?
So um Onyx has
essentially its own mechanism to describe the neural network.
And you can feed in a PyTorch model, translate it into Onyx.
Right.
You can feed in the Jax network, translate it into Onyx.
And then from there we can also branch back out to things we want.
So we have our own Onyx to PyTensor converter.
But we also have we can also go back from Onyx to Jax, for example.
And the networks that that we upload here, they are just Onyx files in the end.
Onyx also has its own runtime.
We are not really exploring that downstream, but just for completeness.
You could stay in the Onyx framework and run the network through their own runtime, et cetera, if you want.
But the key here is so the initial contribution mechanism can go multiple ways, but our initial approach is that well, BayesFlow comes out with the network, right?
So we train, we get a network out, we translate that to Onyx.
And then it's also just an Onyx file and HSSM will know how to interface with that.
And downstream you can do all kinds of batching, etc., um as you wish.
You just need the template of the of the um batch one um version of the network.
So this this approach is can show later um how easy that looks, actually.
Um
But this approach we can take.
There's another library here in the broader ecosystem, it's called SBI.
That's uh it's another library that just encodes quite a few of the simulation based inference and also likelihood facing um networks.
We can just take those networks, train them, put them into Onyx, and the rest of the handshake comes really from um this simulation package, SSM simulators, and we can
construct HSSM models from there.
So we're already well on our way here to uh generalize this away from LAN factory into more or less right in terms of the relevant broad eco as uh simulation based inference
ecosystem.
Super cool.
Yeah.
Yeah, this is super exciting.
Um anyway, that uh I have other questions obviously still, but I know you still have some uh some information from us in the in the current slides.
Do you wanna go into that or do you wanna go into something else?
Um yes.
Let me I'll come back to this at the at the end here, I guess, just like things that are going on.
Or if you wish I can also do that now, like whatever future work that is that is coming in.
Um let me try to share with you
Notebook here quick.
Yeah, let's do that part at the end, exactly if we have time.
And and right now in the meantime if you wanna show a bit of concrete notebooks and stuff like that.
So you've understood folks that uh this episode is gonna be is an interactive one, so I recommend uh going over to the YouTube channel and and watching the video for most of the
parts.
Otherwise if you
If you mainly just enjoy listening, you're also welcome to do it for sure.
I am trying to really narrate as much as I can, but that is uh Yeah, I mean appropriate warning.
Yeah.
Yeah, exactly.
That means You can only do so much.
Okay, this should be
This awesome.
Yeah.
So what are we what are we looking at now?
Okay, so let me um show you this notebook.
this is now the notebook that concerns like what would what would it look like, right, to to start from BayesFlow and end in HSSM.
Okay.
So that's like without
It's more general than that.
So it has two examples, the LAN factory, the native example, and the BayesFlow example, just to illustrate that um like we have a convergence mechanism.
So skipping like the boring part here, right?
Some initial setup.
Um
Initially we can use HSSM here just to simulate a dataset.
So HSSM uh links back to the let me actually zoom in a little bit maybe.
So HSSM links back to this SSM simulators um library and then allows you to like generate synthetic data sets very easily here and you just choose the model you want from that
library.
You pass theta, I always use generically here as um parameters, right?
Um you pass like a vector vector matrix or um a list of parameters here.
You choose how many samples you want.
Um and you can get like a basic data set here of reaction times and choices.
So it's reaction times and responses.
Right.
So this is going to be just the data set that downstream we'll do inference on with the networks that that we have trained.
That's just a repertory here.
Now, first we can take the LAN road.
Um we have so the LAN road is really directly going using SSM simulators and LAN Factory, the two packages, so the simulation package and the little neural net library.
SSM simulators has utilities to generate training data directly instead of just basic simulations from the process.
So it generates training data with the right handshake.
And then in LAN Factory we can
More or less here with two lines of code, right?
We can instantiate a neural net, train it, and it's uh automatically going to save you this Onyx file that we had just talked about.
So here it's even simpler because for uh models that are already included, we don't even need to do this, right?
So um the only thing we need to do on the HSSM side, so this is like the the most trivial HSSM example here.
So I'm just simply gonna say I want to instantiate an HSSM model.
What's my observed data?
What's the name of the cognitive process model, right, that I want to choose.
And then for models for which I have multiple types of likelihoods, right?
I can choose here which type of likelihood I want.
So this is also useful for numerical experiments, etc.
So if I um have a LAN and I also have an analytical likelihood.
Right, I can I can choose here.
Um if you don't choose it it's chosen for you.
If there's only one, it's gonna be that one.
Um let me forget about this parameter here, that's not that interesting now.
But this is this is something that is an overlay over the um cognitive process model.
It's another box of Pandora that would highlight the flexibility of using likelihoods, training likelihoods.
Um okay, so once we have this, my HSSM model is then ultimately a PyMC model, right?
Um we may
Or may not do a little detour to highlight the contribution of Bambi in constructing these models.
So Bambi is um I want to do justice to all the main contributors there.
Um is it um Osvaldo Martin and Tomi Capretto?
Yeah, that's a lot of uh lot of Tomi's work, lots of uh Osvaldo's work, also uh Gabriel Stechschulte.
Um of course I'm not saying his uh name right, but he was on the show.
I'll I'll
I'll uh put the related episodes.
Okay, perfect.
Yeah.
So I I wanna highlight this.
I I didn't know when exactly to do it, but I want to make sure that's clear that um for the regression backend, um even if we might end up not showing this properly during the
podcast now, but we allow the construction of hierarchical regressions for every single one of those parameters of the of the cognitive process model.
And the process of constructing all the design matrices, etcetera.
is going through Bambi, right?
So HSM is deeply integrated really with Bambi on that front.
So if you're familiar with Bambi, you're also going to be able to uh come out guns blazing here pretty quickly with um yeah the kinds of Yeah and Bambi is very fast to uh to uh to
ramp up on by definition it's a bit like uh VRMS but for Python it's it's the idea so
Uh helping you specify your models as fast as possible using Wilkinson notation and and then taking care of the heavy lifting for you.
Exactly.
Yeah.
And then once you have your Bumbi slash uh PyMC model, right, then you can just call dot sample, which is a thin wrapper using the same semantics, right?
Um around the the PyMC version, right?
Of um
running MCMC, right, on your um constructed model.
So in this particular case we are using the NumPyro sampler.
We could be using many of the other favorites, right?
The basic PyMC sampler, not Py, uh send it out to blackjacks, etc.
NumPy is sort of for NumPyro is sort of for us the most trodden path, but anything goes.
Okay, so now fine.
This is like a the simplest example basically of using HSSM.
Just construct it, don't worry about anything, you'll get back parameters here.
as a pro you might uh recognize that this is giving out Rbis inference data, which implies that um we are currently in the migration to PyMC six and Rbus one.
Um okay, now instead let's
Try to use BayesFlow here and see what that would look like, right?
So if I use BayesFlow, then I need exactly one utility that is now sitting in the LAM Factory package.
And we may later even put it directly into BayesFlow.
So one utility transform BayesFlow to Onyx, right?
From there, again, you do simulation.
This is a utility that is coming from the SSM um simulators package.
So we're reusing the same simulator package.
We are just um targeting the kind of training data setup that BayesFlow needs.
um
debating deciding against more detours.
so here we are actually generating the training data alive now because we we do need to train the model that's not pre shipped with the ecosystem.
Um the process to productionise these um the kinds of likelihood estimators that are coming out of BayesFlow is currently ongoing.
So there are even some there are some discussions on how to approach this from an ecosystem perspective.
But um for now, these models you can use them, but they are not yet in the production set as far as HSSRM ecosystem is concerned.
So therefore we're training it live, um, which I did here.
You instantiate one of these, they're called ratio approximators in BayesFlow lingo, go through training process, right?
And then after, you take your BayesFlow network,
Transform it to Onyx and we'll save the we'll have a saved Onyx file in the in the path we of our choice.
Now if we return to HSSM, the only thing that changes here, right, is that I passed the log likelihood, which is allowed to be an Onyx path, on top of the parameters that I
passed.
So I still choose the same model.
I choose the log likelihood kind.
Which is in our lingo is approximate differentiable.
Neural networks fall under that class.
Um you could have a non-neural network approach that would be approximate differentiable, which would fall under that same lingo, but that's kind of immaterial right now.
Um and you can then choose to say what the log likelihood is supposed to be.
So log like kind and log like.
And I can literally just go and pass this onyx path from my BayesFlow network as
block likelihood, construct the model, and from there I can sample as before.
So like the interface is really designed to make it extremely easy to bring in bring custom networks and make use of all the rest of the handshake infrastructure to still
construct valid models.
And we can sample, right?
And now if we look at the outcomes, you can see the posteriors are so we'll have the LAN, LAN factory, posterior in blue, and the neural ratio estimator, which is one particular
neural network that is a um serves as a likelihood approximation out of BayesFlow.
And here they are not perfectly overlapping, but that's because I gave a small training budget to the NRE.
To make this run quickly here for testing.
While the LAN actually had a very large training budget because it's the production room.
But you can see even, you know, this entire notebook will run in a few minutes.
And you can see that even with that little effort, uh, you'll get very reasonable posteriors out of the neural ratio estimators.
So if you're willing to you know use a larger training budget and/or
optimize the hyperparameters.
we o we already know we can make them perfectly aligned.
And
If you're doing research on this kind of stuff, et cetera, right, there's also an option here to simply use HSSM as a spinal layer through which you're running different kinds of
numerical experiments, right?
That's a very convenient layer when you when you're proposing new new neural networks, etc., through which you then run numerical experiments with all kinds of different
formulations, post hoc formulations of particular Bayesian models.
including prior choice, et cetera, et cetera.
So okay, that's that's one of these notebooks.
Um yeah, really cool.
Super interesting.
And if these notebooks are um available online, of course, yeah, do you put them in the in the show notes afterwards because I'm pretty sure people will want to uh to reference uh
reference there.
Um yeah, But in principle those notebooks anyways, they are just stitching together stuff that
People can already find in the docs.
Ah, beautiful.
Okay.
Awesome.
Yeah.
So let's definitely uh put the links to uh to the docs in there.
any anything else you wanted to show us?
I think you you had uh you had at least one other topic that you wanted to to touch on since we're starting to like we're past the hour mark, so I'm gonna start you know, wind
us down.
But yeah, uh wanna make sure you had the time.
Okay.
um Let me so this is going to take a slightly different turn, but Yeah, and then we can go afterwards after that we'll go into uh you know uh what's the what does the future look
like for you and uh and then I will let you go because it's gonna be late for you.
No problem.
It's uh it's a pleasure.
Um okay, so two two small things that I wanted to show here.
The first one is kind of related to a topic we might touch upon at the very end.
Um so one of the things that um we are thinking about when thinking about the uh community and contributors, so there are different level layers, right, of contributions um in open
source and in academia and outside.
And one of the questions that comes up, right, is uh onboarding, right?
How to onboard people and especially now in the the times of um AI capabilities being more and more dominant.
Let's say the tool of choice when you're starting to contribute to anything is often to just throw cloth at it and uh see how much hardware you can make.
We as like more senior people, um, we come from a place of okay, I kind of know what I want to do, right?
And I'm I'm using using Claude.
But if you're a junior contributor, there's a there's a very weird um path to bootstrapping yourself to seniority now.
We've been thinking about how to approach like at least making it very convenient for people to like people become very used now to interactive ways of um gathering knowledge
and not like in a in a very dynamic way, not so static, right?
Um and nobody has patience anymore for reading books, right?
So one of the the approaches is to have a way here to generate uh short form books that help you learn things that you're curious about while
um infusing the learning path with uh our a priori expertise on what we think is relevant and the ecosystem, et cetera.
Right.
So where we landed on and um here I'll show the picture later.
um for existing or upcoming stalkers.
I'll show the picture uh Francesco Moya who is uh who has led the the effort on this which has legs uh way beyond uh the HSM ecosystem anyways.
Um
And the key idea here is that um we are trying to plug into the ecosystem, right?
So have like a um a structured database of things that we consider are important of um in the wider uh HSSM ecosystem and the the topic material in cognitive science, for example,
that is relevant, and allow people here in this web page um to design custom learning paths out of questions that they may have.
around HSSM, uh cognitive science foundations and related topics.
Right.
So in this on this web page you can say, so for example, it um let me go back here.
You can take one example question here or ask anything, right?
Um where what are the core you know likelihood free uh algorithms right that are used
Cognitive science.
And you can go and you can build a learning curriculum here.
Um here I'm just going for saved such um versions just to make it fast.
Right.
This is really constructing on the fly now a personalized learning curriculum that makes use of the kind of uh knowledge that that we would like to infuse it with.
Um and then you can get yourself like a little learning path, right?
For all of the things that you might care about.
And the idea here is that um
This should help people learn about the ecosystem.
It should help people on board also into contributing to the ecosystem.
And in the context also of workshops now, this should really be a take home tutor, right?
Where we have at least a little bit of control over how waiting across information sources is happening, right?
That's one that's one thing.
So that's really about um Yeah, that's right.
educational educational resources related to uh HSSM.
Yeah, that's really awesome.
So definitely check that out, folks.
the the link will be in the show notes.
I think it's it's yeah.
This is gonna be this gonna be hosted very soon.
Um hopefully by the time the podcast comes out, it's gonna be hosted.
Um and then a secondary uh thing here.
And it's really in the same vein.
So I'm I'm also trying to build the
conceptual bridge here is that part of the story here is around um verification, right?
So we can access like a ton of free flow information now and these AI tools are really extremely uh capable.
But there's often there's often sort of the case where you have a few experts that can actually go beyond if you give them their ability to help on specific topic domains,
right?
And just free flow usage of cloud can also be extremely misleading, right?
If you're you are the one driving the interaction while you are not yet a senior person, right?
And similar um in the sim in a similar vein.
Um let me show you this real quick.
So sorry about the infinite recursion, but I I need multiple tabs.
Um we are working on this little tool that we call it Bayesify.
Um
It's a tool where you can go and upload a research paper, also your own, obviously.
and get a structured report on strengths and weaknesses, stepwise on how you applied the um Bayesian workflow, right?
W including then hints at how to how to improve how to improve the paper.
If it's your own, this can serve as a review engine uh to drive you to a good point.
If it's uh other historical papers.
It serves as like a educational resource and later downstream possibly as a as a well of um data through which we can do some historical forensic analysis also on let's say the
quality right of Bayesian analysis over time.
Um so I could go here, right, load a paper.
It will take like a a minute or two to to run this.
So I can just upload the paper here, right?
But I also have it preloaded on this on this um tab.
So this is like my own paper now, right?
Um Like Root Approximation Networks for Fast Inference of Simulation Models and Cognitive Neuroscience.
So what we'll get is this little report here.
So it what is the report about?
Right.
Um first it has a relevance gate, whether the the the paper is even you know in the realm of what makes sense to analyze through Bayesian workflows.
Then we have a rubric here.
So we have um we allow anchoring the analysis on very specific workflow papers.
There's for example a famous paper from Gellmann at al, a Bayesian workflow, right?
And we have a few options, and then we also have our we call it now the gold standard synthesis, which is a synthesis of um a few such candidate papers, right, that is then
projected into a synthesis version of
the stepwise approach and we'll always link to um any particular step to the sources that were used to define the step.
Right.
We'll see in one second.
And it it classifies the paper type.
So in my case it's um method development.
So this is not just an empirical um data analysis.
So it will then check which steps of the workflow are useful.
Nine remain of an original eleven.
And it will give you a score on how many of them sorry, in this case it's seven out of nine.
Sorry.
Um and it will give you an overarching Bayesify score, right?
So far it's been harsh on any of us who tried it.
Uh have not seen a very high score.
Um did you try it on the actual Bayesian workflow paper by Gelman?
It's a good idea.
Well the vibration will be defined.
That paper will be defined as method development.
So we should you're right, we should do we should do that.
Um but since it's defining the framework, you couldn't necessarily ask it to follow itself the framework.
That's uh that's a recursive approach.
Um but yeah, we are we are still working in generally speaking.
Um the goal here is to have um a what we'll call a gold set of like human rated papers.
And um
test the engine for how well it's calibrated to real expert judgment of papers.
So we're currently in the process of collecting such a gold set so that we can check the scoring mechanism here against what experts in the field would have actually said.
So you can get you'll get yourself this report, right?
Many many different steps here, how well they are applied, um, whether or not you were adequate.
You can see I was adequate only on a few of them.
Um, I have a couple of them missing in my paper.
Uh in hindsight, lucky that I got it published.
And then you'll you'll get suggested fixes.
Um it gives you the sources of evidence that were used to judge the step, right?
And it will give you the the standards that were
Pull this in.
The standards that were applied.
Sorry, where is it?
Yeah.
So this particular step, right?
Posterior predictive, retradictive checks, um, is based around the scoring mechanism is based around these two papers.
The visualization invasion workflow um and the Gelman Bayesian workflow paper.
So that's a tool we're we're working on.
Um it's very close to it's actually already reachable now.
We have Bayesify.org.
Um but we're still making refinements to this to make sure that um eventually we can um properly claim that it's uh solid and it also tracks um in a calibration sense real expert
judgments.
Hmm.
Yeah.
Yeah, I I have it here on the screen, Bayesify.org.
So I will add it to the show notes right now.
and that's amazing.
Thanks
Thanks Alex and and and congratulations to all the people who worked on that because I know it's always a collective effort.
Yeah, I want to uh yeah mention here that this is this is actually now joint work with Stefan Radev uh his postdoc Jerry Huang.
I think I know Stefan Radev, like I I'm not sure, but I think the name rings a bell.
You know, I'm I'm not good at memory, but I think so.
Uh yeah, that's awesome.
And I and obviously love these kind of project which is a mix of uh cutting edge research, applied research and educational outreach to make sure that cutting edge research is
actually applied.
So um yeah, it's obviously dear to my heart.
So so very happy to to see that uh you guys are doing that.
Well done.
and yeah, so actually I think it's a good
Time now to start closing us out and and talk about what's next for you.
Uh I know you wanted to talk about that a bit, like what are you looking forward to learn and work on in in the coming months, what are you focusing on, things like that.
Yeah, so um to like stick to the kinds of developments that are going on in the um HSSM ecosystem, I really want to um I really apologize to all the people that are um just
listening.
Um this is uh under optimized for that.
I apologize.
Um I would like to highlight the the kinds of generalizations that are happening here.
So we have we have focused on on one of them, which is that we are really we are looking at HSSM as like a
node in a broad like network graph, right?
It's not some sort of it's not s quote unquote self contained, it's like open to uh infusion, right?
And um the the BayesFlow and SBI integration, et cetera, to have different kinds of networks coming in is one such front.
And the other front is um other types of simulation libraries.
Um there there's one that is used
De facto itself considers itself also an inference library, but um it it has a different approach to simulating from these models.
IDDM is the name of that one.
And then we are also directly working on bringing different classes of um models, making different classes of models um natively supported.
So this is this is some work here of um Ceng Liu and Andrew Zhang, um, who are working on uh algorithms to
De facto that's a throwback now to my Caltech times.
These are these are models that make use of the attention trajectory within a given trial.
So what you're actually collecting are saccade times and uh fixation positions on a screen, and you're trying to feed those in live into in this particular case the drift,
the underlying drift of these of this diffusion model.
Um so Si Cheng Liu
worked out some particular algorithms for this and we are trying to generalize to incorporate those uh models into the toolbox.
Work of um Krishna, supported by Paul Shu and um Carlos Penagua.
ah I apologize to you, Carlos.
Um
This is this combines uh reinforcement learning uh as a backend on uh any parameter of these diff diffusion models into the toolbox.
There's a new class coming here that is um directly modularly combining uh reinforcement learning backends into uh any of the parameters of um any of the diffusion models we're
dealing with.
Um so reinforcement learning without um
Going into like severe detail on into this.
It's just a it's a very canonical paradigm with many, many different variations in its own right to describe a process of learning across trials.
So learning from experience across trials.
And it's there there are many different kinds of experimental paradigms that um in particular then try to tease out the effect.
of for example certain types of mental health conditions on learning trajectories which projects into these into the parameters of these reinforcement learning processes.
and so here's um very quick quickly up here this is work of Krishna himself.
So he he basically works on such patient population data.
Um so HC here is for healthy control and S S S C D is for um schizophrenia.
And he projects basically these um mental health conditions into the parameter space on a rather complicated such reinforcement learning model.
So you can tease apart um effects in parameter space to give a low dimensional interpretational uh interpretable um
How you say lens into what might be impaired uh for people with given diagnosis?
Yeah, yeah.
And then here's the final picture for Francesco Moya.
Francesco is working on bringing in hidden Markov model backends, so also another process backend into um on the parameter side so that uh you can check for regime switches in for
example if people are lapsing in attention or um for any other induced uh reason, switch parameters across time.
So one of the canonical uh reasons to have a regime change is that people's attention fluctuates and they're sometimes switched on, sometimes switched off, for example.
So there's more coming on this front.
These two we have shown now.
So this is the the mentor and Bayesify and here we have Radev and his um postdoc student Jerry One.
So these other things going on.
Uh just to at least put the pictures there.
There are more people to thank.
Yeah, yeah.
that's amazing.
Dave, thanks a lot, uh Alex and and yeah, so many exciting things in in the coming month.
are you also
How a a question I I start asking now also to to guest is uh how do you uh how are you starting to integrate AI tools and LLMs into your own work uh and maybe if you have some
recommendations for the listeners into how to do that in a verifiable and efficient way.
Yeah, so uh let me r I'll raise two things.
so they they are connected by desire.
Um so definitely uh AI is infusing at this point everything that is going on also in the HSM ecosystem.
So like uh co authorship of PRs with AI is at this point totally common, automatic PR reviews and uh then processes to make sure that
Anyone who proposes a PR gets first automatic reviews until it's in a state where human review makes sense, for example, those kind of things.
They're they're happening in the ecosystem.
Um But two particular things that I care about and that we are working towards, uh partly as a result of me caring about it.
Um is that so the development process has changed a little bit from uh working on individual packages.
To having something that is close to a monorepo with some division of um labor so that um you have an easier time context managing.
So now we have this thing which is called HSSM Spine, which is like a a meta design of a repo in which you instantiate all the repos.
So you have like a wholesale way of instantiating the entire ecosystem um in your workspace.
And at least in my experience so far.
um making sure that you always have all the relevant uh packages directly in your workspace really helps with working in a way that maintains glue between something that is
supposed to be an ecosystem just uh spread into packages, right?
Um that's one thing.
And the other thing that um I'm still working through, right?
But um the the spine is one vehicle there is to establish
a nucleus of a common um setup across the depths.
Like maintain a certain type of uh common culture um and a shared toolkit amongst the depths, right?
That is something that uh from what I have seen so far is is an attempt to counteract the extreme uh local specialization of setups that is happening via AI tooling now.
Um
And especially again for uh people that are onboarding and if you have people that are not yet experts into your into your team, um having something that is like a recommended set
of tools that they should install that comes automatically with the spine um to get them sort of to to a baseline level of contribution quality, that doesn't feel hopeless.
is something that as far as I'm concerned is uh is becoming important, right?
Um before we had different means of establishing common culture, but these these AI tools have really led to a total proliferation of extremely custom uh setups that are also all
originating just in the last few months.
Yeah.
Yeah.
Okay.
Yeah, super fun.
Uh it and again, if you have any resources to share on that, uh in the show notes, please do that.
Um so I think you can use the agent uh skill coming.
Actually failed just failed to uh open the PR before uh before the podcast.
Hmm, okay.
So that's actually new for me too.
I didn't know you were doing that.
So that's a great uh great surprise.
So for
These Bayesian skills could become a good central point where you can point towards like all kinds of relevant toolkits uh in the ecosystem and make smart use of it instead of
naive.
Yeah.
Yeah, exactly.
That's the idea um of the the whole repo.
So for people who might have forgotten about it, Bayesian skills um is the repo that I created on the learning Bayesian statistics GitHub.
We already have
uh handful of skills over there, Bayesian workflow, causal inference and amortized inference, which was heavily uh written and influenced by Stefan Radev.
and yeah, like the good thing also is that then the skills can, you know, mention each other and reference each other.
And so basically being a bit a bit smarter when you're doing Bayesian modeling and causal modeling with uh with AI agents.
Um
So yeah, it's all open source.
Yeah, yeah, exactly.
So so yeah.
it's all open source, so feel free to check it out, folks.
download it with your own uh engines, use them.
If you see issues, well, open an issue that that's awesome.
Even better is opening a PR to fix the issue for everybody.
Um, and if you want to contribute new skills, as Alex is gonna do it very soon, well then uh feel free to do so.
and
And then maybe if the skill is is uh is cool and and really interesting, you you can come on the show and present it and explain why you did it.
That that'd be a a fun episode.
Or let it present itself eventually.
Yeah.
It comes pretty close to that.
Yeah, yeah.
Actually do you wanna do you wanna give us the like just a teaser about the what the s the skill is gonna be about?
Yeah, so it's basically an HSSM workflow skill.
So um it's kind of heavily influenced by you know, d there are some tutorials in the docs that showcase uh what a scientific workflow with HSSM would look like.
Starting from a data set, doing some branches of investigations, really in particular on um proposing different models for the same data and having a uh I say a structured process
to evaluate.
what's good, what's bad, and which direction to take.
Um so one of the one of the notebooks that I had prepared here would have been kind of around this just for for timing reasons.
I uh I left it out for now.
Um and then basically go through model comparison, um different kinds of regression functions that you might want to test, etc.
etc.
And eventually land at like a proposed model for for your data set.
So this is like how to use HSSM for that purpose.
Obviously right now it's tailored towards these as we had discussed the cognitive process models.
Um eventually will hopefully branch out to include different model classes.
Amazing.
Yeah.
I think you can stop sharing your screen by the way now.
but yeah, okay, awesome.
Well actually I think by the time I release the episode, probably the skill will be out.
So definitely add it to the show notes uh in the in the episode.
if I see it's out it's out I'll probably give it a shout out in the intro for for this episode.
And and and yeah, like uh any link that you have to the to actual usage and and tutorials, that's gonna be very interesting to people.
Fantastic.
Well I forgot to send you my headshot, so I used the occasion to send you a a massive package of headshot and links.
Yeah, exactly.
Exactly.
Yeah, fantastic.
Well I think we can we can call it a show now.
Uh thank you so much, Alex, for for coming on the show.
Of course, I'm gonna ask you the last two questions ask every guest at the end of the show.
So first one, if you had unlimited time and resources, we'd pop sorry, gonna do that again.
It's it's still the beginning of the day here.
So if you had unlimited time and resources, which problem would you try to solve?
Yeah.
Um I think today this is the first of two questions, so I won't necessarily end on a negative note.
Um I'm hoping to bring read it back in for the second question.
Um I think today one of the core questions that interest me that but that at least for personal limited resources and time I cannot really address in any uh satisfactory way is
um how to
incentivize across generations uh society now to maintain the ability to um you know develop some of this the similar like let to to not let go of the cognitive skills of the
last generation too quickly before we understand whether that's a good idea and how to incentivize people uh from a societal viewpoint to to do this um because I think every one
of us
feels the the ease with which things can now be done and how personally disincentivized one is to do things that are slow but hard now and how much it means to go against the
grain.
if you if you're attempting this.
And I'm that's just speaking as someone that is now to uh reveal my age, thirty seven, right?
But what this really means for
The new generation is a very it's very unclear and certainly it doesn't seem like the incentives are pointing in the direction of preservation.
Um and I'm not yet sure if that's the right approach.
And um if I had the means I would I would try to investigate how one can how can we there's maybe an intersection here with economics on some fundamental way.
How one can work on incentivization um to maintain this ability to another framing would be the ability, the tenacity to generate out of distribution data personally.
But you need security somehow from society and they need to be framed the right way.
Otherwise I see a collective disincentivization to do so at this point.
Okay.
Not the other question.
I don't know if this is too dark.
Yeah.
No, that's that that's great.
that's a very very interesting answer.
First time somebody answers that, so yeah.
Um really, really resonate with it and and appreciate it so uh love it, yeah.
And and second question, if you could have dinner with any great scientific mind, dead, alive or fictional, who would it be?
Yeah.
The guy that
So apart from like really historical figures, right, like Gauss or something, um which you know
may have had a certain type of freakish intelligence that would be uh interesting to observe during a dinner.
Like a more contemporary example for me, someone that is uh I think just became prof Professor Emeritus or just uh resigned.
There's a guy that is called uh Michael I.
Jordan, who is also famous as the Michael Jordan of computer science.
I think a New York Times article or something.
really from a distance in a way, the guy was somehow uh a great inspiration for me and how with how much um I say courage he actually traversed uh academia and managed to make
massive contributions to so many different kinds of um early emerging fields in statistics and machine learning.
So he's he's someone that has worked on the original expectation maximization
has worked on uh recurrent neural networks early uh in nineteen ninety or something.
He wrote the first, I think, right, the first major textbooks on graphical models.
He was behind the popularization of variational inference.
he was uh behind Bayesian non per very famous Bayesian nonparametric methods uh yeah they're there are two canonical ones Chinese restaurant process and uh Indian buffet
process.
Um he is
A guy on the latent Dirichlet allocation paper.
Um it's just a a really immense breadth of contributions.
And um on top of that, he kind of as far as I know, uh he speaks eight languages.
Um he passed through Bucconi when I was in Italy, and I learned that actually he ended up hanging out there for half a year, speaks perfect Italian, uh is on top of everything also
completely humble.
Um
the but I never had the occasion to actually exchange like any type of serious discourse with them.
Um but really from a sort of academic hero perspective uh his is uh impossible to achieve but he has like four hundred thousand citations or something at this point.
But his is sort of the the academic life that uh I treat as some fundamental ideal.
Like going fully with your curiosity, showing a lot of courage and leaving things behind, working on the next thing.
Um and still having the flexibility and breath in life to keep learning totally orthogonal things.
Yeah.
Yeah.
Yeah, very impressive.
Uh would love to join these dinner for sure.
Uh Damn.
Fantastic.
Well I really recommend check the guy out.
Yeah, yeah, yeah, yeah.
Yeah, definitely.
So folks uh check out the the show notes for this episode because they will be dense.
and and remember that we have now on each episode's page a blog post that basically summarizes uh what we talked about during the our our conversation.
We have related episodes, we have key takeaways, the transcript, so um a lot of uh different media that you can use to
uh to uh to actually uh remember and understand what we talked about.
Also now we put all the episodes on an oddbook notebook LM link.
So you can use that also to generate quiz, uh QAs, ask uh Gemini directly, generate um video explainers, audio explainers, lots of things.
So really taking advantage of all the multimodal goodies coming our way.
now with L L so definitely check that out folks.
A lot of uh things have changed on the website in the last in the last few months and I'm I'm very happy with that.
Hopefully it helps you learn even better, if not faster, to come back to your point, Alex, but at least better in a more uh hands on way, which is the goal because learning is
always taking time, you know, it's
repetitions and so yeah to repeat you lead time that's all.
Uh fantastic.
Well Alex thank you so much for uh taking so much time and coming on this show.
Yeah.
Thank you so much and see you soon.
Hopefully sometime in San Francisco.
Of course yeah whenever whenever you're around let me know.
Thank you very much for hosting me
This has been another episode of Learning Bayesian Statistics.
Be sure to rate, review, and follow the show on your favorite podcatcher and visit learnbasetats.com for more resources about today's topics, as well as access to more
episodes to help you reach true Bayesian state of mind.
That's LearnBaseTats.com.
Our theme music is GoodBayan by Beba Brinkman.
Fit MCLas and Megaran.
Check out his awesome work at bebabrinkman.com.
I'm your host.
Alexandora.
You can follow me on Twitter at Alex underscore Andora like the country.
You can support the show and unlock exclusive benefits by visiting patreon.com slash learnbase tax.
Thank you so much for listening and for your support.
You're truly a good baby and change your predictions after taking information.
And if you're thinking I'll be less than amazing, less than just those expectations.
Let me show you how to be a good daisy.
Change calculations after taking fresh data.
Those predictions that your brain is making.
Let's get them on a solid foundation.
HSSM stands for hierarchical sequential sampling models, a generalization of HDDM (hierarchical drift diffusion models), the older toolbox for the same class of decision-making models, but HSSM is built from the ground up on simulation-based inference. That's what lets it handle any variation of the underlying process model, not just the ones with a tractable closed-form likelihood.
The drift diffusion model treats a decision as a random walk that accumulates evidence until it crosses one of two boundaries, with parameters controlling boundary separation, starting bias, and drift rate. It's been used in thousands of published papers largely because it has a closed-form likelihood, which makes standard Bayesian and maximum-likelihood inference fast. Small variations on the model are often just as scientifically motivated, but if their likelihoods aren't analytically convenient, the literature using them stays sparse.
A LAN is a neural network trained to take in a process's parameters and a trial's outcome and output how likely that outcome was, learned purely from repeated simulation rather than derived analytically. Once trained, it functions as a fast, reusable likelihood you plug directly into Bayes' rule, in place of a closed-form solution that may not exist for the model you actually want to fit.
Amortizing the likelihood, HSSM's approach, means training a network once to approximate the likelihood, then reusing that same network across arbitrarily many downstream models: different priors, hierarchical structures, or regression backends, with no retraining. Amortizing the posterior directly, the approach tools like BayesFlow take, gives near-instant inference once trained, but locks the network into the specific scenario it was trained for. Neither is strictly better; it's a trade based on how much reuse you expect.
The deciding factor is reuse. If you plan to run inference over and over, to check whether a new model is broadly identifiable across its parameter space, to run simulation-based calibration, or because many researchers will apply the same model to different datasets, amortizing pays for itself because the trained network is reused indefinitely at near-zero marginal cost per run.
bayesify.org, is built with friend-of-the-show Stefan Radev and his postdoc Jerry Huang. It ingests a research paper and grades it step by step against a Bayesian-workflow rubric anchored on papers like Gelman et al.'s "Bayesian Workflow". It flags whether the paper is even in scope, scores how many of the applicable steps it followed, links each judgment to its source, and suggests fixes.
AI-assisted PR co-authorship and automatic first-pass PR review are already routine in the ecosystem. But Alex is more focused on a side effect he's watching closely: AI coding tools have made it easy for every contributor's local setup to drift into something completely custom, which is especially bad for onboarding people who aren't senior enough yet to tell a good AI suggestion from a bad one. The HSSM spine monorepo and a push toward a shared developer toolkit are both attempts to counteract that fragmentation.
That AI is removing the friction that used to force people to build hard-won cognitive skills, and that society doesn't yet have the incentives in place to make anyone, himself included, choose the slow, difficult path when a fast one is available. He frames it as an open problem at the intersection of psychology and economics: how do you incentivize people, across generations, to keep the tenacity to work things out for themselves, before we even know whether losing that ability matters.

#107 Amortized Bayesian Inference with Deep Neural Networks, with Marvin Schmitt
Listen →
#151 Diffusion Models for SBI in Python, with Jonas Arruda
Listen →
#157 Amortized Inference & BayesFlow in Practice, with Stefan Radev
Listen →
#158 Bayesian Workflows & Foundation Models, with Stefan Radev
Listen →
#112 Advanced Bayesian Regression, with Tomi Capretto
Listen →
#123 BART & The Future of Bayesian Tools, with Osvaldo Martin
Listen →
#142 Bayesian Trees & Deep Learning for Optimization & Big Data, with Gabriel Stechschulte
Listen →